SAR (Synthetic Aperture Radar) image ship fine-grained identification method based on feature fusion
By designing a learning framework for multi-level fusion of multi-physical features and deep features, the problem of existing feature fusion methods being unsatisfactory when distinguishing similar targets is solved, and a higher SAR image ship recognition accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510163182.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
The existing feature fusion method is not effective in distinguishing similar targets, and fails to fully combine expert knowledge and human cognitive perspectives, resulting in insufficient recognition accuracy.
A learning framework based on multi-level fusion of multi-physical features and deep features is designed. By constructing a contour feature extraction module, a key component scattering topological feature extraction module and a length category matching module, combining a backbone network for feature fusion, to achieve fine-grained identification of ships.
Through multi-level feature fusion, the model can more effectively capture the detailed characteristics of the target, improve the ability to distinguish similar targets, and significantly improve the accuracy and reliability of SAR image ship recognition.
Smart Images

Figure CN120107901A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing information processing and application, and specifically relates to a SAR image ship fine-grained recognition method based on feature fusion. Background Art
[0002] There are currently two main solutions for ship recognition in SAR images, one is the transfer learning method and the other is the feature fusion method. In order to cope with the challenges of SAR image modality differences brought about by complex observation conditions, researchers have designed a transfer learning method suitable for SAR images. By migrating the deep learning method based on convolutional neural networks suitable for optical images to SAR images, Kang et al. proposed a SAR target recognition method based on multi-scale feature fusion of R-CNN, which improved the recognition accuracy of targets by fusing features of different scales. Duran et al. alternately trained the Fast R-CNN algorithm and the RPN network to further improve the detection speed of targets in SAR images. Liu Fangjian et al. proposed a NanoDet target detection method for SAR remote sensing images based on visual saliency, which enhanced the recognition ability of targets by introducing a visual saliency mechanism. However, this fine-tuning has little effect on the target recognition effect of SAR images, because the above method still relies solely on the features of the image domain as the basis for recognition, ignoring the unique scattering domain features of SAR images.
[0003] Fusion of scattering domain physical features and deep features is an important development trend in SAR image target recognition research. This method combines knowledge-driven physical models with data-driven deep learning models, which can not only improve the efficiency of target feature extraction and recognition, but also accelerate the convergence speed of the network, so that the use of few-sample learning can also achieve good results, thus providing a solution to the problems of difficulty in obtaining SAR image samples and imbalance between classes; it can also enable the deep network model to give more physically interpretable results during recognition. Compared with the traditional optical image learning network, the unique physical characteristics of SAR images are integrated into the learning network, which is more targeted. At present, some researchers have proposed some fusion methods. Liu Junga et al. fused the learnable scattering features with deep features. This method regards the attribute scattering center as a collection data and pays more attention to the global relationship of the scattering features. Zhang Jinsong et al. also reconstructed the scattering features and deep features into feature maps from the perspective of fusion of scattering features and deep features, realized the fusion of feature maps, and retained the spatial information of the features. Li Chen et al. proposed a method of combining scattering features with graph convolutional networks, converted scattering features into image data, and used GCN to better utilize scattering features.
[0004] The limitations of existing transfer learning methods are mainly reflected in the large domain differences between the source domain (visible light images) and the target domain (SAR images), such as imaging mechanisms, data distribution, and feature characteristics. Visible light images are formed by reflecting light and have rich colors and details, which are mainly reflected in image domain features such as edges and shapes. SAR images are obtained by transmitting microwave signals. Affected by the surface structure and material properties of the object, they show unique speckle noise and scattering effects, which are mainly reflected in scattering domain features such as scattering features and texture information. This difference in imaging mechanism makes it difficult to directly apply the features learned in the source domain to the target domain. Existing transfer learning methods usually focus on the transfer of image domain features and ignore the scattering domain features, resulting in the model being unable to capture effective information. Therefore, when processing SAR images, the model needs to specifically learn and recognize these features, and requires a deeper adaptation mechanism, such as domain adaptation and generative adversarial networks, to effectively improve the recognition ability of SAR images. The existing feature fusion method only performs preliminary structured processing on the extracted scattering features and then fuses them with the deep features at a certain level. It does not fully consider the characteristics of the target in combination with expert knowledge. It analyzes the ship recognition process from the perspective of human cognition and guides network learning accordingly. Although certain results have been achieved, the effect is still not ideal for distinguishing similar targets. Summary of the invention
[0005] (I) Purpose of the invention
[0006] The purpose of the present invention is to provide a SAR image ship fine-grained recognition method based on feature fusion, so as to solve the technical problem that the existing feature fusion method is not ideal for distinguishing similar targets.
[0007] (II) Technical solution
[0008] In order to achieve the above objectives and solve the above technical problems, the technical solutions of the present invention are as follows: 1. A SAR image ship fine-grained recognition method based on feature fusion, comprising the following steps:
[0009] Step 1: Construct a contour feature extraction module based on elliptical Fourier descriptor
[0010] 1.1 Contour feature extraction
[0011] The Canny operator is used to extract the contour preliminarily, and then the closed contour set is obtained through morphological opening, expansion, and erosion. The area of each closed contour is calculated and the contour with the largest area is taken as the target contour. The curvature analysis method is used to dynamically determine the number of points describing the target contour.
[0012] 1.2 Elliptical Fourier Descriptor of Contours
[0013] First, the points (xi ,y i ) is normalized, where i = 1, 2, ..., N;
[0014] Calculate the centroid (Cx, Cy) of the target contour, as shown in Equation 1:
[0015]
[0016] The points of the target contour are translated to the centroid (x′) by formula 2. i , y′ i ):
[0017] x′ i =x i -C x ,y′ i =y i -C y (2)
[0018] The size of the target contour is calculated using formula 3, and scaled using formula 4 to complete the normalization process;
[0019]
[0020] Calculate the parameterized target contour and represent the points of the target contour as complex numbers Z i =x″ i +jy″ i , and perform parameterization; define parameter t as the arc length of the target contour, and use uniformly distributed parameters to represent
[0021] Then calculate the Fourier series A k and B k :
[0022]
[0023] Construct an elliptic Fourier descriptor from the Fourier series:
[0024] D k =A k +jB k (6)
[0025] Select the first m descriptors as the feature representation of the target contour:
[0026] D={D 1 ,D 2 ,...,D m} (7)
[0027] Step 2: Construct a key component scattering topology feature extraction module based on persistence image
[0028] 2.1 Scattering feature extraction of key components
[0029] Extract the attribute scattering center and obtain the scattering situation of the target key components after processing;
[0030] The following simplified ASC model is used to accurately parameterize the target scattering characteristics:
[0031]
[0032] in represents the scattered field of the i-th scattering center, f is the operating frequency, φ is the function of the aspect angle, θ i For parameter sets Together, A i is the amplitude of the ith scattered field, [x i ,y i ] represents its geometric coordinates, L i and Indicates length and direction angle;
[0033] According to the simplified ASC model, the ORB method, which is a rotationally invariant improvement of the original BRIEF algorithm, is used to extract the scattering center, obtain the feature points and the corresponding descriptors, and then the K-means clustering is used to cluster similar feature points into a "cluster", and the mean of all the points in each cluster is used to represent all the points in the cluster.
[0034] 2.2 Scattering characteristics of key components topology
[0035] In order to make the clusters obtained after clustering better reflect the scattering characteristics of key components, and at the same time retain the main scattering information of key components while taking into account a reasonable amount of calculation, in the specific scenario of ship targets, the number range of key points is set according to experience and actual needs, and then the number of "clusters" of key components reflecting the target is determined based on importance sorting, setting persistence thresholds, and clustering effect comparison;
[0036] The distribution characteristics of the above "clusters" are then represented by a persistence graph, thereby obtaining the scattering topological structure of the key components. In order to facilitate the input of the deep learning framework, the persistence image needs to be quantized, and the target data is topologically encoded using Alpha complex filtering to convert complex high-dimensional data into low-dimensional vectors or tensors with invariant properties;
[0037] Step 3, construct a length category matching module based on Gaussian kernel density estimation;
[0038] 3.1 Gaussian Kernel Density Estimation
[0039] The Gaussian kernel density function is used to estimate the probability of the length corresponding to the category. The value of the bandwidth parameter h in the Gaussian kernel density function is dynamically adjusted according to the characteristics of the target data. By calculating the standard deviation of each category of data, the bandwidth parameter h is proportional to the standard deviation of the target data. When the standard deviation of a certain category of target data is large, a larger bandwidth parameter h is used to obtain a smoother probability density function; conversely, when the standard deviation is small, a smaller bandwidth parameter h is used to obtain a more refined probability density function, thereby better capturing the characteristics of each category of target data;
[0040] 3.2 Constructing the matching relationship between length and category
[0041] A self-made dataset is used. When constructing the sample set, the actual resolution of each sample is added to construct the correspondence between the actual length of the target label and the category. Then, a certain amount of processing is performed to form a length-category dataset. For each image, the contour feature extraction module in step 1 is used to extract the contour, and the minimum circumscribed rectangle of the contour is drawn to obtain a more accurate bounding box of the target, and the long side of the bounding box is calculated as the actual length of the target. Each data point contains a target category label and the actual length of the target. These length data are grouped by category to form a length-category dataset.
[0042] The Gaussian kernel density is estimated for the length data of each category, and then these Gaussian kernels are superimposed to obtain the overall probability density function, thus obtaining the length distribution model of each category;
[0043] Based on the length distribution model, the probability distribution of the corresponding category is obtained according to the length predicted by the model. The binary cross entropy loss is used to calculate the difference between the predicted category and the category distribution that the predicted length should correspond to. During the training process, the model will learn the correspondence between length and category.
[0044] Step 4: Use the backbone network to extract the deep features of the input image, and use the contour feature association module and the key component scattering topology feature association module to extract the contour features of the target and the key component scattering features respectively. The backbone network uses the CNN network framework based on the tilted rectangular box for target detection and recognition. The tilted rectangular box detection is to add an angle θ prediction on the basis of the horizontal box to better fit the target shape. In order to avoid the existing boundary problem, the CSL algorithm is used to transform the angle regression problem into a classification problem, and the continuous problem is directly discretized to avoid the boundary situation.
[0045] Step 5, the deep features extracted in step 4, the contour features of the target, and the scattering features of the key components are fused in the channel dimension in the fully connected layer to achieve feature layer fusion;
[0046] Step 6: In the detection phase, the correspondence between the actual length and the category is established by the length category matching module, which guides the network to focus on the size features and perform decision-level fusion, thereby achieving fine-grained ship recognition.
[0047] Furthermore, the curvature analysis method in step 1.1 comprises the following steps:
[0048] First, determine the threshold of curvature through statistical analysis method, and calculate the mean and standard deviation of the curvature values of all contour points;
[0049] The mean plus a certain multiple of the standard deviation is used as the threshold;
[0050] For each target contour point, the curvature is calculated based on its adjacent points before and after it. Points with high curvature represent sharp turns or complex areas of the target contour. At this time, more target contour points should be selected to ensure the accuracy of the description; points with low curvature represent smooth areas. Fewer target contour points should be selected to reduce redundancy.
[0051] Furthermore, the specific implementation process of step 2.2 is as follows:
[0052] Use Alpha complex filtering to map data points to real numbers;
[0053] Use the value of the filter function to construct a complex shape, and the complex shape formula is shown in Equation 9;
[0054]
[0055] Where X is a metric space, α is a non-negative real number, VR(X, α) is the complex of X, any finite subset σ of X represents a simplex, and d(x, y) represents the distance between points x and y;
[0056] Calculate the homology group of the complex. The calculation formula of the homology group is as follows:
[0057]
[0058] in, is the kth boundary mapping, and for all k, satisfies represents the kernel of the nth boundary map, represents the image of the n+1th boundary mapping;
[0059] Generate persistence graphs based on the "birth" and "death" times of homology groups;
[0060] Finally, the persistence graph is converted into a fixed-dimensional feature vector, namely a persistence image. Based on the above method, a persistence image describing the topological characteristics of key components of different types of ship targets can be constructed.
[0061] (III) Effective income
[0062] 1. Based on an in-depth analysis of the main identification features that expert interpretation relies on, this invention proposes a learning framework that fuses multiple physical features and deep features such as size, contour, and key component distribution at multiple levels. Three modules are designed to expand the features of the scattering features to obtain higher-order physical features. The above physical features are expressed by extracting the elliptical Fourier descriptor of the contour, constructing a persistent image of the component distribution, and performing Gaussian kernel density estimation on the size. This allows the contour features and key component distribution features to be fused with the deep features at the feature level, combined with the shape and component information of the target, to enhance the model's ability to capture target details;
[0063] 2. The present invention fuses the size features with the depth features at the decision level, retaining the multi-scale characteristics of the model and avoiding information loss or interference that may be caused by feature-level fusion. Through the fusion of different levels, the model can comprehensively utilize multiple feature information, enhance the ability to distinguish similar targets, and improve the accuracy and reliability of recognition.
[0064] 3. The fine-grained ship recognition method for SAR images based on multi-level fusion of multiple physical features and deep features proposed in the present invention selects multiple core physical features based on the human recognition process of the target, and transforms the multiple physical features through certain mathematical representation methods, converting them into a form suitable for understanding and processing by deep learning networks, thereby realizing the fusion of physical features and deep features, which is more in line with the hierarchy of physical features and the recognition process of deep learning networks.
[0065] 4. The reasonable feature selection and fusion strategy of the present invention is positively correlated with network performance, effectively capturing the main physical identification features of ship targets such as length, outline and distribution of key components, and significantly improving the accuracy and reliability of SAR image ship identification. At the same time, it helps the model to accelerate convergence speed, and even with a small number of samples, it can achieve good results and show excellent performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a schematic diagram of the network design and composition of the present invention. DETAILED DESCRIPTION
[0067] The present invention will be further explained and illustrated below in conjunction with the accompanying drawings.
[0068] A method for fine-grained ship recognition in SAR images based on multi-level fusion of multi-physical features and deep features. Its composition and network design are as follows Figure 1As shown in the figure, it mainly includes the construction of contour feature extraction module (CFAM), key component scattering topological feature extraction module (KCST), and length category matching module (LCMM). Then, the backbone network is used in combination with the multi-level feature fusion network to call the above three modules, and finally the fine-grained recognition of ships in SAR images is realized.
[0069] Step 1: Construct a contour feature extraction module based on elliptical Fourier descriptor
[0070] 1.1 Contour feature extraction
[0071] Contour features, or edge features of targets, play a vital role in image recognition. They refer to places where the attribute distribution (such as the grayscale value or texture of pixels) in an image changes suddenly. These features usually appear as significant changes in pixels around the target, such as step-like or ridge-like changes, providing key information for recognition and detection. In SAR images, the edge of a target is often a series of continuous or intermittent strong scattering points that are approximately linear. These strong scattering points at the edge depict the general outline of the target.
[0072] The present invention uses the Canny operator with better performance to preliminarily extract the target contour, and then obtains a large number of closed contours through morphological opening, expansion, corrosion and other operations. Finally, the area of each closed contour is calculated and the contour with the largest area is found as the target contour. At this time, the obtained target contour is actually composed of a large number of edge points. In order to reduce the amount of calculation, the contour needs to be represented by a certain number of points. Due to the large differences in the size and shape complexity of ship targets, some targets have simple shapes and can be described by only a small number of points, while some targets have complex shapes and need to be described by more points. In order to improve the accuracy and efficiency of the description, the present invention uses a curvature analysis method to dynamically determine the number of points describing the target contour. First, the threshold of the curvature is determined by a statistical analysis method, and the mean and standard deviation of the curvature values of all contour points are calculated, and the mean plus a certain multiple of the standard deviation is used as the threshold. For each contour point, the curvature is calculated according to its adjacent points before and after. Points with high curvature represent sharp turns or complex areas of the contour, and more points are selected to ensure the accuracy of the description; while points with low curvature represent smooth areas, and fewer points are selected to reduce redundancy.
[0073] 1.2 Elliptical Fourier Descriptor of Contours
[0074] The elliptical Fourier descriptor (EFD) approximates the target shape as a superposition of multiple ellipses. Through Fourier transform, the shape is decomposed into ellipse parameters of different frequencies, so that a limited number of ellipse descriptors are used to reconstruct and represent complex contours. This method can effectively capture the characteristics of the shape while retaining key geometric information, making the description smooth and continuous, helping to capture subtle changes in the contour and reduce the impact of noise; it has rotation, scaling and displacement invariance, so that it remains stable under different viewing angles and sizes, thereby improving the robustness of the model; it effectively describes complex contours through limited Fourier coefficients, reduces feature dimensions, and enhances computational efficiency. In addition, its structured features facilitate fusion with deep features.
[0075] First, a series of points (x i ,y i ) for normalization processing, where i = 1, 2, ..., N. First, the centroid (Cx, Cy) of the target contour is calculated, as shown in Formula 4:
[0076]
[0077] The contour point is translated to the centroid (x′) by formula 5. i , y′ i ):
[0078] x′ i =x i -C x ,y′ i =y i -C y (5)
[0079] The size (aspect ratio) of the target contour is calculated using Formula 6, and is scaled using Formula 7 (Formula 7) to complete the normalization process.
[0080]
[0081] Calculate the parametric contour and represent the contour points as complex numbers Z i =x″ i +jy″ i , and perform parameterization. Define parameter t as the arc length of the contour, and use uniformly distributed parameters to represent
[0082] Then calculate the Fourier series Ak and Bk:
[0083]
[0084] Construct an elliptic Fourier descriptor from the Fourier series:
[0085] Dk =A k +jB k (9)
[0086] Usually the first m descriptors (m=5 in the present invention) are selected as the contour feature representation of the target:
[0087] D={D 1 ,D 2 ,...,D m} (10)
[0088] Step 2: Construct a key component scattering topology feature extraction module based on persistence image
[0089] 2.1 Scattering feature extraction of key components
[0090] Since SAR images reflect the size of the backscattering coefficient of the target, the characteristics of the target on the SAR image are mainly reflected in the size of its backscattering coefficient. The attribute scattering center (ASC) is the most prominent feature point of its backscattering coefficient, which represents the information of the part with the strongest scattering characteristics of the target. The key components of ship targets such as bridges, bows, naval guns, vertical launch devices and aprons are made of metal with rough surfaces and generally have high backscattering characteristics. Therefore, the scattering conditions of the key components of the target can be obtained by extracting the attribute scattering center and processing it in a certain way. Image-based attribute scattering center extraction is essentially to fit the image using a scattering model.
[0091] The backscatter of a distributed radar target operating in the high-frequency region can be well approximated as the sum of the responses of a single SC, which can be described as a function of the radar's operating frequency f and the aspect angle ψ, as shown below:
[0092]
[0093] in represents the total scattered field, represents the scattering field of the i-th scattering center, and its corresponding formula is as follows:
[0094]
[0095] where f c represents the center frequency of the radar; θ i The parameter set is [A i ,x i ,y i ,α i ,L i ,φ′ i ,γ i ] together constitute, A i is a scalar that represents the amplitude of the i-th scattered field, αi represents its frequency dependence, [x i ,y i ] represents its geometric coordinates, L i and represents the length and direction angle, γ i Indicates its longitudinal dependence; when the bandwidth and center frequency of the radar system are small, the frequency correlation factor α i can be ignored. At the same time, for the SAR system, γ i It is usually small, so the ASC model can be simplified as follows:
[0096]
[0097] The simplified ASC model can accurately parameterize the target characteristics, and the noise and clutter are also removed. At the same time, ASC can stably extract features from SAR images with different signal-to-noise ratios. By combining the simplified ASC model with the corresponding estimation method, the complex data of the SAR target is represented as several scattering centers with parameters [A i ,x i ,y i ,L i ,φ′ i ], it can be seen that this type of feature extraction method is similar to the scale-invariant feature extraction method. In view of the variability of SAR images in the imaging process and environmental conditions, especially the high sensitivity to the azimuth angle of the target, the scale-invariant feature provides an accurate and stable description method for the key components of the ship target in the SAR image under the background of the simplified ASC model, which is an ideal choice for target detection and recognition.
[0098] According to the simplified ASC model, the present invention uses the ORB (Oriented FAST and Rotated BRIEF) method, which is a rotationally invariant improvement of the original BRIEF algorithm, to extract the scattering center and obtain a large number of feature points and corresponding descriptors. However, the distribution of these feature points in space may be very dense, and there may be a large amount of redundant information. It is necessary to use a clustering algorithm to cluster similar feature points together to form a "cluster". In this way, the center point of each cluster (i.e., the mean of all points in the cluster) can be used to represent all points in the cluster, thereby greatly reducing the number of feature points that need to be processed, while also retaining the main information of the original feature points. The present invention implements the above process through the K-means clustering method. In order to achieve a balance between the amount of calculation and the effect of scattering point extraction, the present invention determines the number of clusters (i.e., the k value) to be 9 based on the number and position of the key components of the ship target after repeated comparison. Each cluster represents the scattering points of a key component of a ship. These 9 cluster centers can be regarded as representatives of the 9 key components of the ship. They reflect the main structure and shape characteristics of the ship, and can relatively completely retain the overall framework and relative distribution of the target while ensuring the calculation efficiency.
[0099] 2.2 Scattering characteristics of key components topology
[0100] The main difference between the scatter plots of key components of different types of ship targets is the difference in distance and distribution between components. The best way to reflect this distribution and distance is the topological structure. Therefore, the distribution characteristics of key components can be transformed into the topological structure of key component scattering characteristics. In order to describe the shape of n-dimensional point cloud data, topological data analysis (TDA) methods are needed. Persistence Diagram (PD) is an important tool in TDA, which provides a powerful method to quantify and understand the topological structure of data. The basic principle is to describe the topological structure of data by quantifying the "birth" and "death" time of the topological features of the data (such as connected components, holes or cavities). The "birth" and "death" times represent the moments when the topological features appear and disappear, respectively.
[0101] However, this method also has two main limitations. First, extracting PD is computationally cumbersome, and the amount of calculation increases significantly with the increase in dimension and the number of samples in the database; the second obstacle is that PD is a collection of multiple points rather than a linear vector, so it is not suitable to be directly used as the input of the deep learning framework. To solve the first problem, while retaining the main scattering information of the key components and taking into account the reasonable amount of calculation, in the specific scenario of the ship target, the number range of key points is set according to experience and actual needs, and then the number of points that reflect the key components of the target is determined based on importance sorting, setting persistence thresholds, and clustering effect comparison. In the embodiment of the present invention, 9 points are selected to reflect the key components of the target. The present invention uses Alpha complex filtering instead of time-consuming VietorisRips or Cech complex filtering. Regarding the second problem, by topologically encoding the data, complex high-dimensional data can be converted into low-dimensional vectors or tensor representations with invariant properties. For example, after the persistence image is quantized, a fixed-length vector is used to describe the topological characteristics of the data. The vector is regarded as a special tensor and can also be called a persistence image, which is suitable for downstream machine learning algorithms.
[0102] First, the data points are mapped to real numbers using Alpha complex filtering. Then, the values of the filter function are used to construct a complex, and the complex formula is shown in Equation 14.
[0103]
[0104] Where X is a metric space, α is a non-negative real number, VR(X,α) is the complex of X, any finite subset σ of X represents a simplex, and d(x,y) represents the distance between points x and y.
[0105] Next, calculate the homology group of the complex. The formula for calculating the homology group is as follows:
[0106]
[0107] in, is the kth boundary mapping, and for all k, satisfies represents the kernel of the nth boundary map, Represents the image of the n+1th boundary map.
[0108] Then, a persistence graph is generated based on the "birth" and "death" times of the homology group. Finally, the persistence graph is converted into a fixed-dimensional feature vector, namely a persistence image. Based on the above method, a persistence image is constructed to describe the topological characteristics of key components of different types of ship targets.
[0109] Step 3: Construct a length category matching module based on Gaussian kernel density estimation
[0110] 3.1 Gaussian Kernel Density Estimation
[0111] The deep learning network can well learn the aspect ratio characteristics of the target and the relative size characteristics in the image, which serve as an important basis for target recognition. However, the aspect ratio of most ship targets is about 7:1, while the real size characteristics are quite different and have a strong correlation with the category. Therefore, for ship targets, the real size characteristics are an important basis for preliminary fine-grained recognition. By establishing the correspondence between the actual length of the target and the category, and incorporating the size prior knowledge into the network training, the model can be assisted to better learn the correspondence between length and category, speed up the convergence of the model, avoid predicting the result of mismatch between length and category, and improve the accuracy of prediction.
[0112] The present invention uses Gaussian Kernel Density Estimation (KDE) to estimate the probability of length-corresponding categories, rather than just giving a single length-category correspondence, mainly to improve the robustness of this correspondence and to better cope with the situation where the length of each type of target changes within a certain range due to various external reasons. The formula for Gaussian Kernel Density Estimation is as follows:
[0113]
[0114] Where n is the number of observations, h is the bandwidth parameter, K is the Gaussian kernel function, Xi is the i-th observation, and x is the point whose probability density you want to estimate.
[0115] The formula of the Gaussian kernel function K is the formula of the standard normal distribution, as shown below:
[0116]
[0117] The bandwidth parameter h plays a very important role in Gaussian kernel density estimation. It controls the width of the Gaussian kernel, thereby affecting the smoothness of the probability density function. The present invention dynamically adjusts the value of the bandwidth parameter h according to the characteristics of the data, and makes the bandwidth parameter h proportional to the standard deviation of the data by calculating the standard deviation of each category of data. In this way, when the standard deviation of a certain category of data is large, a larger bandwidth parameter h can be used to obtain a smoother probability density function; conversely, when the standard deviation is small, a smaller bandwidth parameter h can be used to obtain a more refined probability density function, thereby better capturing the characteristics of each category of target data;
[0118] 3.2 Constructing the matching relationship between length and category
[0119] First, the present invention uses a self-made data set. When constructing the sample set, the actual resolution of each sample is added to construct the correspondence between the actual length of the target label and the category, rather than the relative length. This can prevent confusion when subsequently collecting the length information of the target contained in samples of different resolutions, making the model more instructive and able to more accurately reflect the actual attributes of the target, thereby providing more accurate prediction results.
[0120] Then, the training data set is selected for certain processing to form a length-category data set. In order to avoid the error caused by the tilted box of the manually labeled target, the tilted box is made to fit the target more closely, thereby constructing a more accurate data set. For each image, after extracting the contour using the contour extraction method mentioned above, the minimum circumscribed rectangle of the contour is drawn to obtain a more accurate bounding box of the target, and the long side of the bounding box is calculated as the actual length of the target. Each data point contains a target category label and the actual length of the target. These length data are grouped by category to form a length-category data set.
[0121] Next, we performed a Gaussian kernel density estimation on the length data of each category. This process can be seen as placing a Gaussian kernel (that is, a normal distribution curve) around each length data point, and then superimposing these Gaussian kernels to obtain the overall probability density function. In this way, we get the length distribution model of each category.
[0122] Finally, the established correspondence between length and category is integrated into the model training process. Based on the above length distribution model, the probability distribution of the corresponding category can be obtained according to the length predicted by the model, and the difference between the predicted category and the category distribution that the predicted length should correspond to is calculated using Binary Cross Entropy Loss (BCE), as shown in Formula 18. In this way, the model will learn the correspondence between length and category during the training process.
[0123]
[0124] Where n represents the number of samples, x is the category distribution predicted by the network, y is the category distribution corresponding to the predicted length, and σ(x) is the Sigmoid function that can map x to the (0,1) interval.
[0125] Step 4: Use the backbone network to extract the depth features of the input image, and the contour feature association module and the key component scattering topological feature association module respectively extract the contour features of the target and the key component scattering. The present invention uses a CNN network framework for target detection and recognition based on an inclined rectangular frame as the backbone network.
[0126] Step 5: Fuse the extracted deep features and the above physical features in the channel dimension in the fully connected layer to achieve feature layer fusion.
[0127] Step 6: In the detection phase, the correspondence between the actual length and the category is established by the length category matching module, which further guides the network to focus on the size features and realizes the fusion of the decision layer.
[0128] Example 1
[0129] The specific process of an embodiment provided by the present invention is as follows:
[0130] Due to the large size of the ship, the sample size of the present invention is designed to be 1024*1024. After the image is input into the network, it is downsampled through a convolution layer, the convolution kernel size is 6x6, the step length is 2, and the image size becomes 512x512x64 to reduce the parameters and calculation amount of the network; next, feature extraction and image downsampling are performed through 4 networks composed of C3 modules and convolution layers, the convolution kernel size is 3x3, the step length is 2, and the size of the feature map obtained is 32x32x1024; then, the SPPF module is used to perform spatial pyramid pooling to construct feature maps of different scales and fuse and extract low-level and intermediate features of the image. In the above process, the size of the input image gradually decreases, and the depth of the feature map gradually increases. These features are then processed by a series of convolutions, normalization, RELU activation functions and upsampling, and the outputs of some layers in the connection layer and the previous process are spliced to obtain high-level features. In this process, the scale of the feature map gradually increases, and the depth of the feature map gradually decreases.
[0131] Next, the contour features and key component scattering features extracted by the physical feature extraction module are fused in series in the channel dimension in the fully connected layer. The detection module then generates a detection box from the feature map and makes a category prediction. The module first generates multiple preset anchor boxes at each position of each feature map. Each anchor box has a preset shape and size, which are usually obtained by clustering the real bounding boxes in the training data set. Then, these anchor boxes are fine-tuned to better adapt to the objects in the image. This fine-tuning process is completed through a convolutional layer, the output of which is the four offsets of each anchor box, which correspond to the center position and width and height of the anchor box respectively. Then, a category prediction is made for each anchor box. This process is also completed through a convolutional layer, the output of which is processed by the softmax activation function, which converts the score of each category into a probability so that the sum of the probabilities of all categories is 1, thus obtaining the category probability of each anchor box. These category probabilities represent the probability that each anchor box contains an object of each category.
[0132] Then, through the length category matching module, the correspondence between the length and category of the predicted anchor box is established, and the network parameters are updated through the backward function. Finally, the overlapping detection results are removed through non-maximum suppression (NMS).
[0133] A SAR ship identification method and device based on electromagnetic characteristics and deep learning in the prior art (CN202310042158.4) uses an attribute scattering center model to extract multiple single scattering centers of a SAR image to be identified, and calculates multiple target component images based on the multiple single scattering centers; then the SAR image and each target component image are input into a feature extraction backbone network, and the corresponding target feature vector and multiple component feature vectors are extracted respectively; the target feature vector and the multiple component feature vectors are fused using a component attention unit to obtain a fused feature matrix, wherein, when the target feature vector and the multiple component feature vectors are fused, the features of each channel of the target feature vector are weightedly calculated with the feature vectors of each component to obtain the fused feature vectors corresponding to each channel, and the fused feature matrix is formed by the fused feature vectors of all channels; a convolutional layer is used to perform target recognition on the fused feature matrix, and the target in the SAR image is recognized based on the recognition result.
[0134] Compared with the invention, the method for fine-grained ship recognition in SAR images based on multi-level fusion of multiple physical features and deep features proposed in the present invention selects multiple core physical features based on the human recognition process of the target, and transforms the multiple physical features through certain mathematical expression methods, converting them into a form more suitable for understanding and processing by the deep learning network, and integrating multiple physical features with deep features at the feature level and decision level, respectively, which is more in line with the hierarchy of physical features and the recognition process of the deep learning network. It effectively captures the main physical recognition features of the ship target, such as the length, outline and distribution of key components. Through the fusion of different levels, the model can comprehensively utilize multiple feature information, enhance the ability to distinguish similar targets, and significantly improve the accuracy and reliability of SAR image ship recognition. In addition, reasonable feature selection and fusion strategies show a positive correlation with network performance, effectively accelerating the convergence process of the network. Even under the condition of few samples, the model can show excellent performance. In addition, CN202310042158.4 only selects the key component features of the target and the depth features for fusion, ignoring the complexity and variability of the SAR image scattering characteristics, and the possible deterioration of the stability of the key component features of the target, resulting in slower network convergence and lower accuracy. The additional size features and contour features selected in the present invention have relatively strong stability, and multiple features can be weighted fused to reduce the impact of the instability of key component features, and have better robustness.
[0135] The above contents are further detailed descriptions of the present invention in combination with specific implementation methods, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A method for fine-grained ship recognition in SAR images based on feature fusion, comprising the following steps: Step 1: Construct a contour feature extraction module based on elliptical Fourier descriptor 1.1 Contour feature extraction The Canny operator is used to extract the contour preliminarily, and then the closed contour set is obtained through morphological opening, expansion, and erosion. The area of each closed contour is calculated and the contour with the largest area is taken as the target contour. The curvature analysis method is used to dynamically determine the number of points describing the target contour. 1.2 Elliptical Fourier Descriptor of Contours First, the points (x i ,y i ) is normalized, where i = 1, 2, ..., N; Calculate the centroid (Cx, Cy) of the target contour, as shown in Equation 1: The points of the target contour are translated to the centroid (x′) by formula 2. i , y′ i ): x′ i =x i -C x ,y′ i =y i -C y (2) The size of the target contour is calculated using formula 3, and scaled using formula 4 to complete the normalization process; Calculate the parameterized target contour and represent the points of the target contour as complex numbers Z i =x″ i +jy″ i , and perform parameterization; define parameter t as the arc length of the target contour, and use uniformly distributed parameters to represent Then calculate the Fourier series A k and B k : Construct an elliptic Fourier descriptor from the Fourier series: D k =A k +jB k (6) Select the first m descriptors as the feature representation of the target contour: D={D1,D2,...,D m } (7) Step 2: Construct a key component scattering topology feature extraction module based on persistence image 2.1 Scattering feature extraction of key components Extract the attribute scattering center and obtain the scattering situation of the target key components after processing; The following simplified ASC model is used to accurately parameterize the target scattering characteristics: in represents the scattered field of the i-th scattering center, f is the operating frequency, φ is the function of the aspect angle, θ i For parameter sets Together, A i represents the amplitude of the ith scattered field, [x i ,y i ] represents its geometric coordinates, L i and Indicates length and direction angle; According to the simplified ASC model, the ORB method, which is a rotationally invariant improvement of the original BRIEF algorithm, is used to extract the scattering center, obtain the feature points and the corresponding descriptors, and then the K-means clustering is used to cluster similar feature points into a "cluster", and the mean of all points in each cluster is used to represent all points in the cluster. 2.2 Scattering characteristic topology of key components In order to make the "clusters" obtained after clustering better reflect the scattering characteristics of key components, and at the same time retain the main scattering information of key components while taking into account a reasonable amount of calculation, in the specific scenario of ship targets, the number range of key points is set according to experience and actual needs, and then the number of "clusters" of key components reflecting the target is determined based on importance sorting, setting persistence thresholds, and clustering effect comparison; The distribution characteristics of the above "clusters" are then represented by a persistence graph, thereby obtaining the scattering topological structure of key components. To facilitate the input of the deep learning framework, the persistence image needs to be quantized, and the target data is topologically encoded using Alpha complex filtering to convert complex high-dimensional data into low-dimensional vectors or tensors with invariant properties. Step 3, construct a length category matching module based on Gaussian kernel density estimation; 3.1 Gaussian Kernel Density Estimation The Gaussian kernel density function is used to estimate the probability of the length corresponding to the category. The value of the bandwidth parameter h in the Gaussian kernel density function is dynamically adjusted according to the characteristics of the target data. By calculating the standard deviation of each category of data, the bandwidth parameter h is proportional to the standard deviation of the target data. When the standard deviation of a certain category of target data is large, a larger bandwidth parameter h is used to obtain a smoother probability density function; conversely, when the standard deviation is small, a smaller bandwidth parameter h is used to obtain a more refined probability density function, thereby better capturing the characteristics of each category of target data; 3.2 Constructing the matching relationship between length and category A self-made dataset is used. When constructing a sample set, the actual resolution of each sample is added to construct the correspondence between the actual length and category of the target label, and then a length-category dataset is formed after certain processing. For each image, the contour feature extraction module in step 1 is used to extract the contour, and the minimum circumscribed rectangle of the contour is drawn to obtain a more accurate bounding box of the target, and the long side of the bounding box is calculated as the actual length of the target. Each data point contains a target category label and the actual length of the target. These length data were grouped by category to form a length-category dataset; The Gaussian kernel density is estimated for the length data of each category, and then these Gaussian kernels are superimposed to obtain the overall probability density function, thus obtaining the length distribution model of each category; Based on the length distribution model, the probability distribution of the corresponding category is obtained according to the length predicted by the model. The binary cross entropy loss is used to calculate the difference between the predicted category and the category distribution that the predicted length should correspond to. During the training process, the model will learn the correspondence between length and category. Step 4: Use the backbone network to extract the deep features of the input image, and use the contour feature association module and the key component scattering topology feature association module to extract the contour features of the target and the key component scattering features respectively. The backbone network uses a CNN network framework based on a tilted rectangular box for target detection and recognition. The tilted rectangular box detection is to add an angle θ prediction on the basis of the horizontal box to better fit the target shape. In order to avoid the existing boundary problem, the CSL algorithm is used to transform the angle regression problem into a classification problem, and the continuous problem is directly discretized to avoid the boundary situation. Step 5, the deep features extracted in step 4, the contour features of the target, and the scattering features of the key components are fused in the channel dimension in the fully connected layer to achieve feature layer fusion; Step 6: In the detection phase, the correspondence between the actual length and the category is established by the length category matching module, which guides the network to focus on the size features and perform decision-level fusion, thereby achieving fine-grained ship recognition.
2. According to the method for fine-grained ship recognition in SAR images based on feature fusion according to claim 1, it is characterized in that: The curvature analysis method in step 1.1 comprises the following steps: First, determine the threshold of curvature through statistical analysis method, and calculate the mean and standard deviation of the curvature values of all contour points; The mean plus a certain multiple of the standard deviation is used as the threshold; For each target contour point, the curvature is calculated based on its adjacent points before and after it. Points with high curvature represent sharp turns or complex areas of the target contour. At this time, more target contour points should be selected to ensure the accuracy of the description; points with low curvature represent smooth areas. Fewer target contour points should be selected to reduce redundancy.
3. The method for fine-grained ship recognition in SAR images based on feature fusion according to claim 1 is characterized in that: The specific implementation process of step 2.2 is as follows: Use Alpha complex filtering to map data points to real numbers; Use the value of the filter function to construct a complex shape, and the complex shape formula is shown in Equation 9; Where X is a metric space, α is a non-negative real number, VR(X, α) is the complex of X, any finite subset σ of X represents a simplex, and d(x, y) represents the distance between points x and y; Calculate the homology group of the complex. The calculation formula of the homology group is as follows: in, is the kth boundary mapping, and for all k, satisfies represents the kernel of the nth boundary map, represents the image of the n+1th boundary mapping; Generate persistence graphs based on the "birth" and "death" times of homology groups; Finally, the persistence graph is converted into a fixed-dimensional feature vector, namely a persistence image. Based on the above method, a persistence image describing the topological characteristics of key components of different types of ship targets can be constructed.
Citation Information
Patent Citations
SAR (Synthetic Aperture Radar) target identification method and device based on electromagnetic characteristics and deep learning
CN116051994A