A Radar Rainfall Detection Method and System Based on ResNet50-PCA-XGBoost
Through the ResNet50-PCA-XGBoost method, combined with deep learning and gradient enhancement classifier, the problem of insufficient accuracy and robustness in the existing radar rainfall detection methods is solved, and efficient and accurate radar rainfall detection is achieved.
Patent Information
- Application Number
- CN202510466304.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-15
AI Technical Summary
Existing radar rainfall detection methods rely on a single image processing technology or traditional machine learning models, making it difficult to accurately identify rainfall noise interference in radar echo signals, resulting in limited detection accuracy and robustness.
The method based on ResNet50-PCA-XGBoost is adopted to extract the depth characteristics of the radar echo signal through the improved ResNet50 algorithm, and dimensionality reduction is achieved by combining principal component analysis. The XGBoost classifier is used to determine whether the radar echo is affected by rainfall, and local and global feature information is integrated.
It improves the accuracy and robustness of radar rainfall detection, reduces computing resource consumption, and realizes efficient rainfall detection, which is suitable for real-time monitoring and wind field inversion of X-band navigation radar.
Smart Images

Figure CN120009892B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of radar rainfall detection. More specifically, it relates to a radar rainfall detection method and system based on ResNet50-PCA-XGBoost. Background Art
[0002] This study aims at the problem of rainfall noise interference when inverting sea surface wind field parameters by X-band navigation radar, and proposes an efficient rainfall detection method. Rainfall will seriously affect radar echo signals and reduce the accuracy of wind field information extraction. Therefore, accurate identification of rainfall contamination is crucial for subsequent rain-noise correction and wind field inversion. Traditional maritime rainfall detection usually relies on shipborne rain gauges, while the rainfall detection method based on X-band radar images proposed by Lund et al. uses the proportion of the number of pixels with gray values less than 5 as the judgment criterion. However, this method is highly dependent on radar system parameters and it is difficult to obtain accurate and stable rainfall information. In addition, many existing methods use single image processing techniques or traditional machine learning models, such as manual extraction based on shallow features or support vector machine (SVM) classification, which are difficult to capture complex texture and intensity patterns in radar images, resulting in limited detection accuracy and robustness. Typical existing solutions for rainfall detection based on radar include:
[0003] (1) Clustering recognition: The Chinese invention patent with the publication number CN113450308A discloses a radar rainfall detection method, device, computer device and storage medium, including obtaining a radar image and performing preprocessing; extracting features from the radar image; calculating the distance between the features of the radar image and several preset clustering centroids, where the clustering centroids include the clustering centroids of rainy radar images and the clustering centroids of rainless radar images; determining whether the clustering centroid with the smallest distance between the features of the radar image and the clustering centroids is the clustering centroid of a rainy radar image; if so, the obtained radar image is a rainy radar image, otherwise it is a rainless radar image. This method is actually a typical K-means algorithm, and its result highly depends on the selection of the initial centroids. In the case of randomly selecting initial centroids, different initial centroids may converge to local optima, resulting in fluctuations in the classification results. A solution similar to this one is the Chinese invention patent with the publication number CN113920422A, a method for identifying rainfall radar images, which has the same problem.
[0004] (2)Threshold recognition: A Chinese invention patent with the publication number CN110208806A discloses a method for recognizing rainfall in marine radar images. The original radar image is obtained by reading the radar file, and the co-frequency interference is removed. Within the Cartesian box area of the radar image to be detected, the echo difference value of a point is calculated using pixel points that are half a main wavelength distance apart from the point to be calculated in space, and the average value of the echo difference values in the Cartesian box area of the radar image is obtained. By comparing the average value of the echo difference values with the detection threshold, rainfall radar images and non-rainfall radar images are recognized. This solution relies on the average echo difference and is easily interfered by outliers (such as local strong reflection targets). Ignoring the spatial distribution characteristics may lead to misjudging uniform non-precipitation clutter as rainfall.
[0005] Therefore, there is still a need for a method that can accurately identify radar echo signals affected by rainfall. Summary of the Invention
[0006] To solve the above problems, the technical solution adopted in this application is a radar rainfall detection method based on ResNet50-PCA-XGBoost, which includes storing the radar echo signal as a two-dimensional data matrix, and also includes:
[0007] ResNet50-PCA feature extraction: The two-dimensional data matrix is preprocessed to obtain a tensor as the input of ResNet50, and feature extraction is performed through the improved ResNet50 algorithm, and then dimensionality reduction is performed through principal component analysis to obtain the ResNet50-PCA feature vector after dimensionality reduction;
[0008] Histogram feature extraction: The occurrence frequency of different pixel intensity values in the two-dimensional data matrix is counted to obtain an intensity histogram, and the feature vector composed of the bin values of the intensity histogram is extracted and normalized to obtain a normalized histogram feature vector;
[0009] Feature concatenation: The deep feature vector and the histogram feature vector are concatenated to form an input vector for XGBoost classification, and it is judged whether the radar echo is affected by rainfall based on XGBoost classification.
[0010] Optionally, the data preprocessing includes: passing the two-dimensional data matrix through median filtering and smoothing processing, then generating a single-channel grayscale image, performing dimensionality reduction using multiple rounds of max pooling, adjusting it to the standard input size through bilinear interpolation, and then converting the dimensionality-reduced single-channel grayscale image to a three-channel pseudo-RGB format and normalizing it using ImageNet.
[0011] Optionally, the improved ResNet50 algorithm removes the fully connected layer and softmax layer of ResNet50, only retains the convolutional layer and average pooling layer, directly outputs the feature vector, and the process of feature extraction by the improved ResNet50 is carried out in the evaluation mode while disabling gradient calculation.
[0012] Optionally, feature concatenation includes performing Z-score normalization on the ResNet50-PCA features and histogram features respectively, and directly concatenating the normalized feature vectors.
[0013] Optionally, feature concatenation includes performing Z-score normalization on the ResNet50-PCA features and histogram features respectively, evaluating the feature performance using the XGBoost classifier according to the optimized wind speed bins based on the normalized data, and for each wind speed interval obtained after the optimized wind speed binning, assigning weights to the ResNet50-PCA features and histogram features according to the accuracy rates of the ResNet50-PCA features and histogram features, and then performing weighted concatenation.
[0014] Optionally, the optimized wind speed binning includes:
[0015] Initial binning: The wind speed is divided into k intervals according to equal-frequency binning. For the i-th interval, if its sample size is less than 5% of the total sample size, then the i-th interval and the i+1-th interval are merged;
[0016] Accuracy rate merging: Calculate the accuracy rate difference between adjacent intervals according to the following formula, and merge the intervals with an accuracy rate difference less than 0.5%:
[0017] ;
[0018] In the formula, represents the interval accuracy rate, represents the accuracy rate of classifying the ResNet50-PCA features using the XGBoost classifier in the i-th interval, represents the accuracy rate of classifying the ResNet50-PCA features using the XGBoost classifier in the i+1-th interval, represents the accuracy rate of classifying the histogram features using the XGBoost classifier in the i-th interval, represents the accuracy rate of classifying the histogram features using the XGBoost classifier in the i+1-th interval; and then the wind speed intervals of the optimized wind speed binning are obtained.
[0019] Optionally, the weight assignment is carried out according to the following formula:
[0020] ; ;
[0021] Wherein, α is the ResNet50-PCA feature weight in the i interval, and β is the histogram feature weight in the i interval.
[0022] Optionally, dimensionality reduction by principal component analysis includes: standardizing the feature vectors obtained by feature extraction of the improved ResNet50 to obtain a standardization matrix, constructing a covariance matrix based on the standardization matrix, performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors, selecting the first k principal components with the largest eigenvalues to make the variance contribution rate reach 95%, and multiplying the standardization matrix by the matrix composed of the first k eigenvectors to obtain the dimensionality-reduced data matrix.
[0023] Optionally, XGBoost classification includes: inputting feature vectors, traversing all CART trees for prediction, accumulating the prediction values of all CART trees, mapping the original prediction values through the Sigmoid function, and judging the output category according to the set threshold.
[0024] This application also provides a radar rainfall detection system based on ResNet50-PCA-XGBoost, including:
[0025] Data input module: used to receive radar echo signals and store them as two-dimensional data matrices;
[0026] Data preprocessing module: performing data preprocessing on the two-dimensional data matrix to obtain a tensor as the input of ResNet50;
[0027] ResNet50 feature extraction module: used to perform feature extraction through the improved ResNet50 algorithm;
[0028] PCA dimensionality reduction module: used to perform dimensionality reduction by principal component analysis to obtain the dimensionality-reduced ResNet50-PCA feature vectors;
[0029] Histogram feature extraction module: used to count the occurrence frequencies of different pixel intensity values in the two-dimensional data matrix to obtain an intensity histogram, extract the feature vectors composed of the intensity histogram bin values for normalization, and obtain the normalized histogram feature vectors;
[0030] Feature splicing module: used to splice the deep feature vectors and histogram feature vectors to form the input vector for XGBoost classification;
[0031] XGBoost classification module: judging whether the radar echo is affected by rainfall based on XGBoost classification.
[0032] The beneficial effects of a radar rainfall detection method and system based on ResNet50-PCA-XGBoost provided by this application are as follows:
[0033] In this application, the improved ResNet50 algorithm is used to extract the deep features of radar echo signals. Dimensionality reduction is performed through principal component analysis to improve the local perception ability of feature extraction. The spliced histogram features reflect the global intensity distribution. Feature extraction of radar data is carried out from two different perspectives, namely local and global, enabling the model to utilize both local details and global statistical information simultaneously, and comprehensively understand the rainfall features in radar images, thereby improving the accuracy and robustness of rainfall detection. Principal component analysis is optimized based on the deep features extracted by ResNet50, so that the features after dimensionality reduction can still retain most of the useful information while reducing the consumption of computing resources. The XGBoost gradient boosting decision tree classifier is adopted to further optimize the rainfall detection results. XGBoost can make full use of the features after dimensionality reduction by principal component analysis while maintaining computational efficiency, and improve the generalization ability of classification. In addition, the decision tree structure of XGBoost can capture the non-linear features of data, making the rainfall detection results more robust. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art.
[0035] Figure 1 It is a comparison chart of radar image intensity histograms under similar wind speeds;
[0036] Figure 2 It is the overall technical roadmap provided by the embodiments of the present invention;
[0037] Figure 3 It is a schematic diagram of the complete residual network structure;
[0038] Figure 4 It is a comparison chart of the accuracy rates of each experimental group in the embodiments of the present invention;
[0039] Figure 5 It is a schematic diagram of the module segmentation of the radar rainfall detection system based on ResNet50-PCA-XGBoost provided by the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] In order to make the technical problems, technical solutions and beneficial effects to be solved by this application more clearly understood, the following further details this application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0041] Embodiment 1
[0042] As Figures 1-3 shown, a radar rainfall detection method based on ResNet50-PCA-XGBoost includes storing radar echo signals as a two-dimensional data matrix, and also includes:
[0043] ResNet50-PCA feature extraction: performing data preprocessing on the two-dimensional data matrix to obtain a tensor as the input of ResNet50, performing feature extraction through the improved ResNet50 algorithm, and then performing dimensionality reduction through principal component analysis to obtain the dimensionality-reduced ResNet50-PCA feature vector;
[0044] Data preprocessing includes: passing the two-dimensional data matrix through median filtering and smoothing processing, then generating a single-channel grayscale image, performing dimensionality reduction using multiple rounds of max pooling, adjusting it to the standard input size through bilinear interpolation, and then converting the dimensionality-reduced single-channel grayscale image to a three-channel pseudo-RGB format and normalizing it using ImageNet.
[0045] The X-band navigation radar generates a two-dimensional data matrix by transmitting electromagnetic wave pulses and receiving the echo signals backscattered from the sea surface for wind field inversion and rainfall detection. Each matrix element corresponds to a specific distance and azimuth angle and is stored as a grayscale value from 0 to 255, directly reflecting the sea surface scattering characteristics and the intensity distribution of rainfall interference. Generally, the rainfall area appears as a higher grayscale value in the echo image, while the non-rainfall area shows a lower grayscale value. However, due to the large size of the original data matrix and the presence of noise, directly inputting it into a deep neural network may not only lead to information loss but also significantly increase the computational complexity. Therefore, the present invention provides an efficient data preprocessing process to make the radar echo data adapt to the input requirements of ResNet50 (RGB image of 224×224×3) while retaining the key information related to rainfall.
[0046] In data preprocessing, not only the feature extraction of a single echo is concerned, but also the spatial distribution characteristics of rainfall echoes are fully considered, and the original two-dimensional data is cleaned and dimensionally reduced to optimize the model input. First, regarding the noise problem of radar signals, an adaptive median filtering method is introduced to remove isolated noise points while enhancing the coherence of the rainfall area, thereby improving the data distinguishability. Second, since traditional cropping methods may cause information loss, in order to retain the main features of the rainfall area while reducing the data size, the max pooling method is used, so that the dimensionally reduced data can still fully express the spatial distribution characteristics of rainfall and provide more representative input data for subsequent deep learning models. This preprocessing scheme not only improves the data quality but also ensures that the radar echo information can still be effectively used for rainfall detection and wind field inversion after dimensionality reduction.
[0047] The preprocessing process is divided into the following steps:
[0048] Data cleaning: Apply median filtering to the original two-dimensional matrix to remove isolated noise points (such as outliers caused by equipment interference or random scattering), and smooth the data to highlight the coherent characteristics of the rainfall area. Median filtering replaces the pixel value by taking the median of the local neighborhood, effectively retaining edge information and avoiding excessive blurring of the intensity distribution of the rainfall area.
[0049] Matrix dimensionality reduction: Considering that the size of the original matrix far exceeds 224×224, directly cropping may lose rainfall information in the edge area. This method uses MaxPooling for dimensionality reduction. MaxPooling takes the maximum value of the local area by setting an appropriate pooling kernel size and stride, gradually reducing the matrix size to 224×224. The advantage of MaxPooling is to highlight the high-intensity echoes in the rainfall area while covering the entire matrix to avoid information loss. Subsequently, the matrix is fine-tuned to the exact 224×224 size through bilinear interpolation to ensure smooth transition of the image content and avoid detail distortion.
[0050] Generate a pseudo RGB image: The dimensionality-reduced matrix is still a single-channel grayscale image. To meet the three-channel input requirements of ResNet50, the grayscale value is copied to the three channels of R, G, and B (i.e., R = G = B = grayscale value) to generate a pseudo RGB image. Although this conversion does not introduce color information, the rainfall characteristics are mainly reflected in the intensity distribution, and copying the channels will not affect subsequent feature extraction. Then, apply ImageNet normalization (subtracting the mean and dividing by the standard deviation) to the pseudo RGB image to generate a tensor with the shape of [1, 3, 224, 224] as the input of ResNet50.
[0051] The improved ResNet50 removes the fully connected layer and the softmax layer of ResNet50, only retains the convolutional layer and the average pooling layer, and directly outputs the feature vector. The process of feature extraction through the improved ResNet50 is carried out in the evaluation mode, and the gradient calculation is disabled at the same time.
[0052] The preprocessed tensor is input into ResNet50 for deep feature extraction. ResNet50 is a deep residual network, and its core lies in the residual connection, which can effectively alleviate the problem of gradient disappearance and improve the stability and discriminability of feature extraction. Different from traditional handcrafted feature extraction or shallow machine learning methods, ResNet50 can still maintain gradient stability through the residual connection structure, thus ensuring the efficient extraction of rainfall characteristics. In this process, the data is not only transformed through the network, but also undergoes multi-layer feature mapping, enabling the initial low-level features (such as edges and textures) to gradually evolve into high-level features (such as the global distribution pattern of the rainfall area).
[0053] In this application, the specific processing flow of the pseudo RGB image entering ResNet50 is as follows:
[0054] First, the preprocessed tensor (shape [1, 3, 224, 224]) passes through a 7×7 convolutional layer with a stride of 2 for initial feature extraction to generate low-level features, and further reduces the spatial resolution and computational amount through a 3×3 max pooling layer (stride = 2). Subsequently, the tensor enters four residual block groups (Conv2_x to Conv5_x). Each residual block adds the input directly to the output through a skip connection to alleviate the vanishing gradient problem. The skip connection mechanism ensures that even as the network deepens, the gradient can still be effectively backpropagated, thereby improving training stability and feature extraction ability.
[0055] After multiple convolutional operations, the early layers capture low-level features, and the deep convolutional layers extract high-level features. At the final stage of the network, the feature map is processed through a 7×7 global average pooling layer to compress the spatial dimension into a single value, generating a 2048-dimensional feature vector (shape [1, 2048]). To adapt to the rainfall detection task, this method removes the last fully connected layer and softmax layer of ResNet50, and only retains the convolutional layer and average pooling layer to directly output the feature vector. The feature extraction process is carried out in evaluation mode, disabling gradient calculation to reduce memory consumption and accelerate inference.
[0056] Deep features are gradually extracted through multiple residual blocks. Combining the rainfall intensity information retained in the preprocessing stage, the 2048-dimensional feature vector output by ResNet50 contains rich semantic information related to rainfall in the radar image, providing a strong feature representation for subsequent dimensionality reduction and classification.
[0057] Based on the above-mentioned partial structure, Figure 3 What is shown is the complete ResNet50 structure, including the entire computational process from the input image to the final output classification result. The network structure adopted in the present invention is composed of multiple convolutional layers, residual blocks, and pooling layers, removing the fully connected layer in the figure to achieve efficient feature extraction and classification tasks.
[0058] Dimensionality reduction through principal component analysis includes: normalizing the feature vector obtained by feature extraction of the improved ResNet50 to obtain a normalized matrix, constructing a covariance matrix based on the normalized matrix, performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors, selecting the first k principal components with the largest eigenvalues to make the variance contribution rate reach 95%, and multiplying the normalized matrix by the matrix composed of the first k eigenvectors to obtain the dimensionality-reduced data matrix.
[0059] The feature vectors extracted by ResNet50 have a relatively high dimension. Directly inputting them into the classifier may lead to a too high computational complexity and a large amount of redundant information. Therefore, in the dimensionality reduction stage, the principal component analysis (PCA) method is introduced. By utilizing its linear transformation characteristics, the high-dimensional feature vectors are projected onto a low-dimensional space that can best represent the data variation. In the present invention, PCA is optimized based on the deep features extracted by ResNet50, so that the features after dimensionality reduction can still retain most of the useful information while reducing the consumption of computing resources.
[0060] Suppose there is a dataset containing (n) samples and (m) features, denoted as the matrix , where each row is a sample and each column is a feature. The goal of PCA is to find a new coordinate system (principal components) such that: the variables in the new coordinate system are orthogonal to each other (i.e., uncorrelated), the first principal component explains the largest variance in the data, the second principal component explains the largest part of the remaining variance, and so on. Dimensionality reduction is achieved by selecting the first k principal components (k < m), while minimizing the information loss.
[0061] Before performing PCA, it is usually necessary to standardize the data to eliminate the influence of different feature dimensions. The standardized data is calculated by the formula:
[0062] ;
[0063] where is the element in the i-th row and j-th column of the original matrix, is the mean of the j-th feature, is the standard deviation of the j-th feature. After standardization, the mean of each feature is zero and the variance is one.
[0064] PCA determines the directions of the principal components by analyzing the covariance structure of the data. The covariance matrix of the standardized data is defined as:
[0065] ;
[0066] In the formula, is the feature matrix, represents the standardized data matrix. The covariance matrix captures the linear correlations between the features. The goal of PCA is to find a transformation matrix to make the covariance matrix of the new variables a diagonal matrix.
[0067] Since C is a symmetric matrix, it can be eigen-decomposed:
[0068] ;
[0069] In the formula, is the eigenvector matrix, the column vectors are the eigenvectors of C and are orthogonal to each other , is a diagonal matrix, and the elements on the diagonal are the eigenvalues of C. The eigenvalue represents the variance magnitude along the direction of the i-th eigenvector. The eigenvector defines the new coordinate direction, that is, the principal component direction.
[0070] For dimensionality reduction, the first k principal components with the largest eigenvalues are selected. These principal components retain the largest part of the variance in the data. The proportion of retained variance is calculated by the following formula:
[0071] ;
[0072] Usually, k is selected so that the variance contribution rate reaches a certain threshold. Let be the matrix composed of the first k eigenvectors. Then the data after dimensionality reduction is:
[0073] ;
[0074] where is the data after dimensionality reduction, and each column represents the projection of the data in the direction of the i-th principal component.
[0075] The solution of the components can be regarded as an optimization problem. Suppose we are looking for the direction of the first principal component (unit vector, ), the goal is to maximize the variance after projection:
[0076] ;
[0077] This is an optimization problem with constraints. The constraint condition is , and the Lagrange multiplier method is used:
[0078] ;
[0079] Take the derivative of and set it equal to zero:
[0080] ;
[0081] This indicates that is the eigenvector of C, is the corresponding eigenvalue. Substitute into it, and it can be known that when the variance is the largest, takes the largest eigenvalue , . Similarly, the second principal component is orthogonal to Maximize the variance under orthogonal conditions, and so on.
[0082] If it is necessary to reconstruct the original data from the dimension-reduced data (Z), the following formula can be used for approximate restoration:
[0083] ;
[0084] Since (unless ), there will be information loss in the reconstructed standardized data matrix and the original data is approximated as:
[0085] ;
[0086] In the formula, represents the reconstructed original data matrix.
[0087] The eigenvectors after PCA dimension reduction significantly reduce the computational complexity, making the training and prediction of XGBoost more efficient. The eigenvectors are reduced from 2048 dimensions to 100 dimensions, and the amount of calculation is reduced by about 20 times. The training time is also significantly reduced. At the same time, noise interference is reduced, and the sensitivity of the model to rainfall patterns is improved. The dimension-reduced features still retain 95% of the original information under the setting of k = 100, ensuring the accuracy of rainfall detection. For example, the features of strong echo regions are highlighted, and the texture patterns of weak rainfall can also be effectively captured. In the rainfall detection of X-band radar images, PCA with a fixed k of 100 also enhances the robustness of the model, adapts to image changes under different weather conditions (such as different rainfall intensities and background noise changes), and avoids overfitting problems caused by high-dimensional features.
[0088] Histogram feature extraction: The frequency of different pixel intensity values in the two-dimensional data matrix is counted to obtain the intensity histogram, and the eigenvector composed of the bin values of the intensity histogram is extracted and normalized to obtain the normalized histogram eigenvector;
[0089] Feature concatenation: The deep eigenvector and the histogram eigenvector are concatenated to form the input vector for XGBoost classification, and it is judged whether the radar echo is affected by rainfall based on XGBoost classification.
[0090] Feature concatenation includes performing Z-score normalization on the ResNet50-PCA features and the histogram features respectively, and directly concatenating the normalized eigenvectors.
[0091] This method uses the ResNet50-PCA features (100 dimensions) and the histogram features (64 dimensions) for concatenation to form the final 164-dimensional eigenvector. The concatenation method is as follows:
[0092] First, the two sets of features are mean-variance normalized (Z-score normalization), that is:
[0093] ;
[0094] in is the feature vector, and Its mean and standard deviation ensure that features of different scales have similar distribution.
[0095] The normalized feature vectors are directly concatenated, namely:
[0096] ;
[0097] in is the 100-dimensional PCA dimension reduction feature, The 64-dimensional histogram features are concatenated to obtain a 164-dimensional feature vector. In the overall vector, ResNet50-PCA features account for 61% (100 / 164) and histogram features account for 39% (64 / 164). This ratio ensures the leading role of deep features and improves the distinguishing ability of rainfall detection by combining statistical features.
[0098] The concatenated 164-dimensional feature vector is finally input into the XGBoost classifier and classified using the gradient boosting tree algorithm. XGBoost prevents overfitting by optimizing the objective function, supports parallel computing and column sampling to improve training efficiency, and ultimately achieves high-precision rainfall detection. Since XGBoost can capture the nonlinear characteristics of data, this method can effectively distinguish between rainfall and rain-free states under the same wind speed conditions. Even under strong wind conditions, it can still correctly identify changes in feature distribution caused by rainfall, thereby improving the robustness of the model. Compared with traditional methods, this method combines the advantages of deep learning and machine learning, while reducing the demand for computing resources, achieving higher detection accuracy and robustness, and providing innovative technical support for rainfall monitoring and wind field inversion of X-band navigation radar.
[0099] Based on the feature vector after PCA dimension reduction, this method further introduces the pixel intensity histogram feature to improve the accuracy and robustness of rainfall detection. The pixel intensity histogram can intuitively show the intensity distribution characteristics of the X-band radar two-dimensional data matrix, and this distribution is affected by both wind speed and rainfall. The radar scattering cross-section depends on the sea surface wind speed and the angle between the radar line of sight and the wind direction, which can be described by the second-order harmonic empirical expression. However, under rainfall conditions, the wind-generated spectrum will be disturbed by the annular spectrum generated by raindrop impact, causing the original scattering model to fail, thereby affecting the echo characteristics of the radar signal and changing the distribution of the pixel intensity histogram.
[0100] Specifically, the intensity histogram presents its distribution in the form of a probability density function by counting the occurrence frequencies of different pixel intensity values in the two-dimensional data matrix. Under the condition of the same wind speed, in the case of no rain, the pixel intensity mainly concentrates in the low gray value region, showing a relatively concentrated histogram shape; while under rainfall conditions, the proportion of medium and high intensity pixels increases significantly, and the histogram distribution becomes more dispersed. Although the change in wind speed will also affect the intensity of the sea surface echo, a higher wind speed will enhance the sea surface echo, resulting in an overall increase in intensity, but the impact of wind speed on the histogram is much smaller than that of rainfall. Under rainfall conditions, the change in the histogram is more significant, and it can clearly distinguish between rainfall and no-rain states. Although the histograms under different wind speed conditions vary, under rainfall conditions, the impact of rainfall on the histogram is more prominent. Therefore, the change in wind speed will not cause significant interference to rainfall classification. In this way, the intensity histograms of rainfall and no-rain states can still maintain a significant difference under different wind speed conditions, providing strong support for rainfall classification.
[0101] To ensure the consistency of histogram features under different radar coverage ranges and image sizes, this method first calculates the intensity histogram of the two-dimensional data matrix and normalizes the obtained bin values to represent the discrete probability density function of pixel intensity. Subsequently, the normalized histogram feature vector is concatenated with the ResNet50 depth features after PCA dimensionality reduction to form a complete input vector. Finally, this comprehensive feature vector is input into the XGBoost classifier to utilize the advantages of its gradient boosting decision trees to achieve efficient and robust rainfall detection. XGBoost can not only make full use of the depth features after PCA dimensionality reduction but also combine the statistical information of histogram features to improve the sensitivity to rainfall patterns, ensuring that the detection method has high classification accuracy and real-time performance under different weather conditions and wind speeds.
[0102] In addition, the 100-dimensional ResNet50-PCA feature vector is more suitable for the gradient boosting tree training of XGBoost, and the tree splitting and optimization process is more efficient. Eventually, XGBoost can quickly output the results of "rainy" or "rain-free", meeting the actual needs of real-time rainfall detection, such as applications in meteorological monitoring stations or automatic warning systems. The PCA dimensionality reduction scheme with a fixed k = 100 also reduces the hardware requirements, enabling this method to run on resource-constrained embedded devices.
[0103] XGBoost classification includes: inputting the feature vector, traversing all CART trees for prediction, accumulating the prediction values of all CART trees, mapping the original prediction values through the Sigmoid function, and judging the output category according to the set threshold.
[0104] Finally, in the classification stage, we adopt the XGBoost gradient boosting decision tree classifier to further optimize the rainfall detection results. Different from traditional classification methods such as SVM, XGBoost can make full use of the features after PCA dimensionality reduction while maintaining computational efficiency, improving the generalization ability of classification. In addition, the decision tree structure of XGBoost can capture the non-linear features of the data, making the rainfall detection results more robust.
[0105] In the ResNet50-PCA-XGBoost process, the XGBoost classifier, as the core classification module, is responsible for achieving accurate rainfall classification based on the 100-dimensional feature vector after PCA dimensionality reduction. In the rainfall detection task of X-band radar images, the features after PCA processing have removed redundant information, but an efficient classifier is still needed to distinguish "rainy" and "rainless". Traditional classifiers such as support vector machines (SVM) or random forests may be inefficient or have insufficient generalization ability when dealing with high-dimensional features, while XGBoost uses the ensemble learning method of gradient boosting trees, gradually constructing multiple decision trees through iterative optimization, significantly improving the classification accuracy and robustness. XGBoost combines L1 / L2 regularization and subsampling techniques to prevent overfitting and ensure that the model performs stably in different rainfall scenarios (such as heavy rainfall and light rainfall), which is particularly suitable for the actual needs of real-time rainfall detection.
[0106] In the training process of XGBoost, a CART (Classification and Regression Tree) is generated in each iteration, and finally an ensemble model of multiple CART trees is formed. This method approaches the target step by step, which can more effectively avoid overfitting compared with the way of approaching quickly in large steps. XGBoost does not completely rely on a single residual tree. Instead, it believes that each tree only grasps part of the truth, and only uses a part of the prediction results of each tree when accumulating, making up for the deficiencies through the combination of multiple trees. The final predicted value is the aggregation of the prediction results of all trees.
[0107] The XGBoost classifier can accurately capture the characteristics of rainfall images and judge whether the radar echo is affected by rainfall. Its gradient boosting mechanism reduces the prediction deviation through iterative optimization, and at the same time combines regularization techniques to prevent overfitting, ensuring the reliability of the model under complex weather conditions (such as a mixture of rainfall and noise). In addition, through PCA dimensionality reduction optimization, the training time is shortened to meet the requirements of real-time monitoring.
[0108] Through this overall framework, we have established an end-to-end rainfall detection method, realizing the efficient processing from the original radar data to the rainfall classification results. Different technical modules cooperate with each other and are gradually optimized to ensure that while maintaining the detection accuracy, the computational complexity is reduced, providing stable and reliable technical support for the wind field inversion of X-band navigation radars.
[0109] This study proposes a rainfall detection method based on ResNet50-PCA-XGBoost, which integrates deep learning feature extraction, dimensionality reduction technology, and gradient boosting classification model, making full use of the deep features and dynamic patterns of radar data to improve the accuracy and real-time performance of rainfall classification. Specifically, this method first preprocesses the radar echo data and uses max pooling for dimensionality reduction to reduce information loss; then it extracts deep features using ResNet50 and reduces the computational complexity through principal component analysis (PCA); meanwhile, it combines the pixel intensity histogram features to quantify the impact of rainfall on the radar echo distribution and enhance the classifier's ability to identify rainfall areas; finally, it uses the XGBoost classifier for binary classification prediction. This method can effectively distinguish rainfall from non-rainfall under the same wind speed conditions and achieve efficient and robust rainfall detection. The overall technical route is as Figure 2 shown.
[0110] This study proposes a low-complexity rainfall detection method based on ResNet50-PCA-XGBoost, which combines deep learning feature extraction, dimensionality reduction technology, and gradient boosting classification algorithm to fully explore the deep features and dynamic patterns in radar data to improve the accuracy and real-time performance of detection. This method models X-band radar rainfall detection as a binary classification task (rainy / no-rainy), and realizes efficient rainfall identification through steps such as data matrix preprocessing, deep feature extraction, dimensionality reduction, and classification. The radar echo signal is stored as a two-dimensional data matrix, and each matrix element corresponds to the echo intensity at a specific distance and azimuth angle, with a value range of 0-255, which directly reflects the sea surface scattering characteristics and the impact of rainfall. Rainfall areas usually show high gray values, while non-rainfall areas are mainly low gray values.
[0111] To adapt to the input of the ResNet50 network, this method first preprocesses the data matrix, including median filtering to remove isolated noise points and smoothing the data to enhance the overall features of the rainfall area. Since the original matrix already has gray values of 0-255, additional normalization steps are avoided, and a single-channel gray image can be directly generated. To reduce the data size and retain key information, this method uses max pooling instead of the traditional central cropping strategy. Max pooling can retain the maximum value in the local area while reducing the resolution, highlighting the strong echo signal in the rainfall area, covering the entire matrix at the same time, and avoiding the loss of edge information caused by cropping. After multiple rounds of pooling, the matrix size is reduced to close to 224×224, and then it is adjusted to the standard input size through bilinear interpolation to ensure smooth transition of the image content and reduce detail distortion.
[0112] The single-channel grayscale image after dimensionality reduction needs to be converted to a three-channel pseudo-RGB format to meet the input requirements of ResNet50. The conversion method is to copy the grayscale value to the R, G, and B channels (i.e., R = G = B = grayscale value) to generate a pseudo-RGB image. Although this method does not introduce additional color information, it can ensure that ResNet50 processes the image correctly while retaining the intensity characteristics of the rainfall area. Subsequently, the image is normalized to a tensor and normalized using the ImageNet mean and variance to match the pre-trained weights of ResNet50. The final tensor shape is [1, 3, 224, 224], meeting the batch input requirements.
[0113] In the feature extraction stage, the convolutional layers and average pooling layers of ResNet50 are retained, while the fully connected layers are removed to ensure the output of high-dimensional depth features. After the image undergoes multiple convolutional processes, the low layers extract edge and texture information, and the high layers capture the complex distribution patterns of the rainfall area. Finally, through the global average pooling layer, a 2048-dimensional feature vector ([1, 2048]) is output for subsequent dimensionality reduction and classification. To reduce the computational complexity and reduce redundant information, this method uses principal component analysis (PCA) to reduce the feature vector to 100 dimensions while retaining approximately 95% of the variance information. Compared with LDA or Autoencoder, PCA has a lower computational cost and is suitable for real-time detection.
[0114] To further improve the classification effect, this method introduces pixel intensity histogram features to reveal the probability distribution characteristics of X-band radar echo signals. At the same wind speed, the intensity distribution of radar echoes shows a systematic shift between rainy and rainless conditions, and this shift can be quantified and utilized through histogram analysis. When there is no rain, the pixel intensity mainly concentrates in the low grayscale area, while when it is raining, the proportion of medium and high intensity values increases and the overall distribution is more dispersed. In addition, under higher wind speed conditions, although the overall radar echo signal in the rainless area is enhanced, it still shows a relatively concentrated distribution pattern, while the echo signal in the rainfall area shows a more complex scattering characteristic. For example, under the condition of a wind speed of 13.6 m / s, although the echo signal in the rainless area is enhanced, the pixel intensity histogram in the rainfall area still shows greater discreteness and higher intensity peaks compared to the rainless situation. In addition, the intensity histograms of rainy and rainless states can still maintain significant differences under different wind speed conditions, providing strong support for rainfall classification. Therefore, this method directly extracts the intensity histogram from the original two-dimensional echo intensity matrix and normalizes the bin values to adapt to the size changes brought about by different radar coverage ranges to ensure the robustness of the features. Finally, the normalized histogram feature vector is concatenated with the ResNet50 depth features after PCA dimensionality reduction to form a complete input vector, providing a richer feature representation for the XGBoost classifier and improving the accuracy and robustness of rainfall detection.
[0115] Model simulation and correctness verification
[0116] Feature vectors are extracted from the pseudo-RGB images of 7775 radar two-dimensional matrices, including 5906 rain-free pseudo-RGB images and 1869 rain-and-pollution pseudo-RGB images. Using the feature extraction method proposed above, in the classification and recognition stage, XGBoost is used as the classifier to achieve accurate classification of rain-and-pollution images. In the ablation experiment, the dataset is divided into a training set and a test set according to a ratio of 8:2. The following experimental groups are set:
[0117] Experimental group 1: Only use ResNet50 to extract features, without PCA dimensionality reduction and XGBoost classification, and directly use the extracted features for nearest neighbor classification.
[0118] Experimental group 2: After using ResNet50 to extract features, perform PCA dimensionality reduction and use the XGBoost classifier for classification.
[0119] Experimental group 3: Use the gray histogram to extract features, and then use the XGBoost classifier for classification.
[0120] Experimental group 4: The complete method, that is, ResNet50 extracts features, PCA dimensionality reduction, and XGBoost classification, and at the same time introduces the gray histogram method to extract feature vectors.
[0121] To measure the performance of the model in radar image classification, this paper selects the accuracy rate as an index for evaluation. The accuracy rate refers to the percentage of the number of samples correctly classified by the model in the total number of samples, as shown in the following formula.
[0122] ;
[0123] In the formula, Accuracy represents the accuracy rate, represents the true positive, represents the false positive, represents the false negative, represents the true negative, and the specific explanations are shown in Table 1.
[0124] Table 1 Specific explanations of TP, FP, FN, and TN
[0125]
[0126] Such as Figure 4As shown, the experimental results indicate that the proposed complete method (Experimental Group 4) outperforms other experimental groups in all performance metrics, fully verifying the rationality and synergy among various modules. In terms of accuracy, Experimental Group 4 performed the best, reaching 98.3%, significantly higher than Experimental Group 1 (94.5%), Experimental Group 2 (95%), and Experimental Group 3 (94.7%). This result demonstrates the advantage of the model proposed in this paper in improving classification accuracy.
[0127] In Experimental Group 1, only ResNet50 was used to extract features and classified by a nearest neighbor classifier. The results show that the features extracted by ResNet50 alone cannot effectively improve classification performance, with an accuracy of only 94.5%. In Experimental Group 2, after feature extraction by ResNet50, PCA dimensionality reduction was performed and an XGBoost classifier was used for classification, with an accuracy of 95%. This experimental result shows that while PCA dimensionality reduction can remove redundant information, it can effectively improve classification performance, but there is still room for improvement compared to the complete method.
[0128] In Experimental Group 3, only gray - level histogram features were used and classified by an XGBoost classifier, with an accuracy of 94.7%. Although gray - level histogram features can provide useful information in some cases, compared with the results of combining with the deep features extracted by ResNet50, the accuracy gap is obvious, indicating that the deep features of ResNet50 are crucial for the classification task.
[0129] In contrast, the accuracy of Experimental Group 4 is 98.3%, the highest among all experimental groups, and it also performs excellently in terms of computational efficiency. Through PCA dimensionality reduction, the dimension of the feature vector is greatly reduced, and the training and inference speeds of XGBoost are significantly faster than those of Experimental Group 3, making the model more practical in resource - constrained environments. In addition, while maintaining classification performance, Experimental Group 4 can significantly improve computational efficiency, further demonstrating the advantage of combining PCA dimensionality reduction with deep feature extraction.
[0130] In summary, the experimental results show that the complete method (Experimental Group 4) not only improves the recognition accuracy but also achieves a significant improvement in computational efficiency, becoming the optimal solution.
[0131] The optimal parameters of XGBOOST obtained through grid search are shown in Table 2.
[0132] Table 2 Optimal hyperparameter settings of the XGBOOST network
[0133]
[0134] Example 2
[0135] The difference from Example 1 is that the feature splicing in this example includes performing Z-score standardization on the ResNet50-PCA features and histogram features respectively. Based on the standardized data, the XGBoost classifier is used to evaluate the feature performance according to the optimized wind speed binning. For each wind speed interval obtained after the optimized wind speed binning, the weights of the ResNet50-PCA features and histogram features are assigned according to the accuracies of the ResNet50-PCA features and histogram features, and then weighted splicing is performed.
[0136] The optimized wind speed binning includes:
[0137] Initial binning: The wind speed is divided into k intervals according to equal-frequency binning. For the i-th interval, if its sample size is less than 5% of the total sample size, then the i-th interval and the i + 1-th interval are merged;
[0138] Accuracy merging: Calculate the accuracy difference between adjacent intervals according to the following formula, and merge the intervals with an accuracy difference less than 0.5%:
[0139] ;
[0140] In the formula, represents the interval accuracy, represents the accuracy of the ResNet50-PCA features classified by the XGBoost classifier in the i-th interval, represents the accuracy of the ResNet50-PCA features classified by the XGBoost classifier in the i + 1-th interval, represents the accuracy of the histogram features classified by the XGBoost classifier in the i-th interval, represents the accuracy of the histogram features classified by the XGBoost classifier in the i + 1-th interval; then the wind speed intervals of the optimized wind speed binning are obtained. The optimized wind speed binning can better capture the feature performance under different wind speed conditions, thereby improving the generalization ability of the model under different wind speed conditions. This enables the model to maintain good performance under new wind speed conditions.
[0141] The weight assignment is performed according to the following formula:
[0142] ; ;
[0143] In the formula, α is the weight of the ResNet50-PCA features in the i-th interval, and β is the weight of the histogram features in the i-th interval. The weight assignment can balance the contributions of the two features and avoid a single feature dominating the model's decision. This helps to improve the robustness of the model under complex weather conditions and enables it to maintain good performance under different wind speed conditions.
[0144] Example 3
[0145] As Figure 5 shown, the present application also provides a radar rainfall detection system based on ResNet50-PCA-XGBoost, including:
[0146] Data input module: used to receive radar echo signals and store them as two-dimensional data matrices; used to uniformly receive the original data of different radar devices, shield hardware differences, and separate module segmentation is beneficial for radar device adaptation.
[0147] Data preprocessing module: perform data preprocessing on the two-dimensional data matrix to obtain a tensor as the input of ResNet50;
[0148] ResNet50 feature extraction module: used to perform feature extraction through the improved ResNet50 algorithm;
[0149] PCA dimensionality reduction module: used to perform dimensionality reduction through principal component analysis to obtain the ResNet50-PCA feature vector after dimensionality reduction;
[0150] Histogram feature extraction module: used to count the occurrence frequency of different pixel intensity values in the two-dimensional data matrix to obtain an intensity histogram, extract the feature vector composed of the intensity histogram bin values for normalization, and obtain the normalized histogram feature vector;
[0151] Feature splicing module: used to splice the deep feature vector and the histogram feature vector to form the input vector for XGBoost classification; independent segmentation of this module is beneficial for adopting different feature splicing methods.
[0152] XGBoost classification module: based on XGBoost classification to judge whether the radar echo is affected by rainfall.
[0153] Performing the above module segmentation enables each module to be improved separately without reconstructing the entire system.
[0154] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A radar rainfall detection method based on ResNet50-PCA-XGBoost, characterized by: This includes storing radar echo signals as a two-dimensional data matrix, and also includes: ResNet50-PCA feature extraction: Preprocess the two-dimensional data matrix to obtain a tensor as the input of ResNet50, extract features using the improved ResNet50 algorithm, and then reduce the dimension through principal component analysis to obtain the reduced-dimensional ResNet50-PCA feature vector; Histogram feature extraction: Count the frequency of occurrence of different pixel intensity values in the two-dimensional data matrix to obtain an intensity histogram, extract the feature vector composed of the intensity histogram bin values, normalize them, and obtain a normalized histogram feature vector; Feature splicing: The depth feature vector and the histogram feature vector are spliced to form the input vector of the XGBoost classification. Based on the XGBoost classification, it is determined whether the radar echo is affected by rainfall.
2. The radar rainfall detection method based on ResNet50-PCA-XGBoost according to claim 1, characterized in that: The data preprocessing includes: median filtering and smoothing the two-dimensional data matrix to generate a single-channel grayscale image, using multiple rounds of maximum pooling to reduce the dimension, adjusting it to the standard input size through bilinear interpolation, and then converting the reduced single-channel grayscale image into a three-channel pseudo RGB format and standardizing it using ImageNet.
3. The radar rainfall detection method based on ResNet50-PCA-XGBoost according to claim 1, characterized in that: The improved ResNet50 algorithm removes the fully connected layer and softmax layer of ResNet50, retains only the convolution layer and the average pooling layer, directly outputs the feature vector, and the process of feature extraction through the improved ResNet50 is performed in the evaluation mode, while disabling gradient calculation.
4. The radar rainfall detection method based on ResNet50-PCA-XGBoost according to claim 1, characterized in that: The feature splicing includes performing Z-score normalization on the ResNet50-PCA features and the histogram features respectively, and directly splicing the normalized feature vectors.
5. The radar rainfall detection method based on ResNet50-PCA-XGBoost according to claim 1, characterized in that: The feature splicing includes Z-score standardization of ResNet50-PCA features and histogram features, and performing feature performance evaluation using an XGBoost classifier based on the standardized data according to optimized wind speed binning. For each wind speed interval obtained after optimizing the wind speed binning, weights are assigned to the ResNet50-PCA features and the histogram features according to their accuracy, and then weighted splicing is performed.
6. The radar rainfall detection method based on ResNet50-PCA-XGBoost according to claim 5, characterized in that: The optimized wind speed binning includes: Initial binning: Divide the wind speed into k intervals according to equal frequency binning. For the i-th interval, if its sample size is less than 5% of the total sample size, merge the i-th interval and the i+1-th interval; Accuracy merging: The accuracy difference between adjacent intervals is calculated according to the following formula, and the intervals with accuracy difference less than 0.5% are merged: ; In the formula, represents the interval accuracy, Indicates the accuracy of classification of ResNet50-PCA features using XGBoost classifier in interval i, Indicates the accuracy of classification of ResNet50-PCA features using XGBoost classifier in interval i+1, Indicates the accuracy of the histogram feature classification using the XGBoost classifier in interval i, It indicates the accuracy of the classification of the histogram features using the XGBoost classifier in interval i+1; and then the wind speed interval of the optimized wind speed bin is obtained.
7. The radar rainfall detection method based on ResNet50-PCA-XGBoost according to claim 6, characterized in that: The weight distribution is performed according to the following formula: ; ; Where α is the ResNet50-PCA feature weight in interval i, and β is the histogram feature weight in interval i.
8. The radar rainfall detection method based on ResNet50-PCA-XGBoost according to claim 1, characterized in that: The dimensionality reduction through principal component analysis includes: standardizing the eigenvectors obtained by feature extraction of the improved ResNet50 to obtain a standardized matrix, constructing a covariance matrix based on the standardized matrix, performing eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors, selecting the principal components with the largest first k eigenvalues to make the variance contribution rate reach 95%, and multiplying the standardized matrix by the matrix composed of the first k eigenvectors to obtain a data matrix after dimensionality reduction.
9. The radar rainfall detection method based on ResNet50-PCA-XGBoost according to claim 1, characterized in that: The XGBoost classification includes: inputting a feature vector, traversing all CART trees for prediction, accumulating the prediction values of all CART trees, mapping the original prediction values through the Sigmoid function, and determining the output category according to a set threshold.
10. A radar rainfall detection system based on ResNet50-PCA-XGBoost, characterized in that: include: Data input module: used to receive radar echo signals and store them as a two-dimensional data matrix; Data preprocessing module: preprocess the two-dimensional data matrix to obtain tensors as input for ResNet50; ResNet50 feature extraction module: used to extract features using the improved ResNet50 algorithm; PCA dimensionality reduction module: used to reduce the dimension through principal component analysis to obtain the reduced ResNet50-PCA feature vector; Histogram feature extraction module: used to count the occurrence frequency of different pixel intensity values in the two-dimensional data matrix to obtain an intensity histogram, extract the feature vector composed of the intensity histogram bin value, normalize it, and obtain a normalized histogram feature vector; Feature splicing module: used to perform feature splicing on the deep feature vector and the histogram feature vector to form the input vector of XGBoost classification; XGBoost classification module: Determine whether the radar echo is affected by rainfall based on XGBoost classification.
Citation Information
Patent Citations
Navigation radar image rainfall recognition method
CN110208806A
Radar rainfall detection method and device, computer equipment and storage medium
CN113450308A
Rainfall radar image identification method
CN113920422A
Artificial intelligence selection and configuration
CN115413346A
Sea wave parameter inversion method combining sea wave spectrum analysis method and radar image geometric shadow statistical method based on Canny operator
CN116908854A