Target identification system based on learning score Gabor transformation and local scattering extraction network

By combining fractional Gabor transform and local scattering extraction network, the problem of low SAR target recognition rate under occlusion conditions is solved, achieving efficient recognition of occluded SAR images and improving recognition accuracy and discriminability.

CN121010754APending Publication Date: 2025-11-25SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202511118112.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing deep learning methods have low accuracy in SAR target recognition under occlusion conditions, struggle to generate a large number of samples with different occlusion directions and proportions, and are unable to effectively extract global and local key information of occluded targets.

Method used

A target recognition system based on learned fractional Gabor transform and local scattering extraction network is adopted. The system uses a learned fractional Gabor transform module to perform multi-scale, multi-directional, and multi-kernel filtering response fusion, combined with a local scattering extraction module to extract key structural features, and optimizes the recognition results through a feature recognition network.

Benefits of technology

It improves the recognition accuracy of occluded SAR images, especially maintaining high recognition performance under high occlusion rates, and enhances the model's ability to distinguish different SAR targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010754A_ABST
    Figure CN121010754A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of radar target recognition, and provides a target recognition method and system based on learning score Gabor transformation and a local scattering extraction network and a training method thereof, and the method comprises the steps: generating a shielding image based on an obtained SAR data set; enabling a learnable score Gabor transformation module to carry out filtering response fusion and feature extraction on the occlusion image to obtain a processed image; performing key structure feature extraction on the processed image by a local scattering extraction module to obtain a feature image; and performing identification optimization on the feature image through a feature identification network to obtain an identification result. According to the method, multi-scale, multi-direction and multi-core global information extraction of the occluded image is realized by utilizing learnable score Gabor transformation, key extraction of residual important features in the occluded image is realized by the local scattering extraction module, and the problem of insufficient global and local information extraction capability of a traditional method is overcome, so that the recognition accuracy of the occluded SAR image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of radar target recognition technology, specifically, it relates to a target recognition system based on learning fraction Gabor transform and local scattering extraction network, and more particularly to an obstructed SAR target recognition method, system and training method based on learning fraction Gabor transform and local scattering extraction network. Background Technology

[0002] Synthetic Aperture Radar (SAR), as a high-resolution imaging sensor, boasts advantages such as all-weather, all-day operation, high resolution, and high precision, and is widely used in fields such as geological disaster monitoring, resource exploration, military, and national defense. With the rapid development of deep learning, Automatic Target Recognition (ATR) in SAR imaging has become a research hotspot in the field.

[0003] Deep learning methods have achieved satisfactory recognition results under standard operating conditions (SOCs), but target occlusion caused by variations in azimuth, incident angle, and noise levels under extended operating conditions (EOCs) severely affects SAR recognition performance, leading to a significant drop in accuracy. To address the occlusion problem in SAR recognition tasks, existing research has proposed many deep learning methods, such as the scattering excitation learning-channel dropout (SEL-CD) method and the centerlocal constraint shadow residual network (ClcsrNet) method.

[0004] However, these methods fail to extract key information effectively when the target's critical structure is damaged, and their generalization performance is poor. Furthermore, the complexity of the model also significantly impacts their effectiveness. Currently, the main technical challenges are: how to generate a large number of occlusion samples with different occlusion directions and proportions using limited angular samples; how to design a global information extraction module to avoid insufficient extraction of occluded target information; and how to design a local key information extraction module to extract key features of the occluded target and enhance recognition capabilities.

[0005] The patent document "A Method for Re-identifying Occluded Pedestrians with Multi-Scale Hypergraph Connections Using KAN Structure" (CN119169525A) discloses a method that integrates 3D human body information into the pedestrian feature extraction part, removes the influence of background noise, uses a KAN network combined with a multi-scale hypergraph to learn and transfer semantic information from different regions, and uses a 3D representation mask to participate in the training of the loss function, assisting in the training of the entire model, eliminating the influence of inactive occlusion features, and enhancing the role of useful features. However, it is designed for optical images and is not fully applicable to SAR images that present the overall structure of scattering points due to speckle noise.

[0006] SAR offers the advantages of all-weather and all-day operation, but it is also affected by electromagnetic and multiplicative speckle noise. Furthermore, it presents the overall structure of scattering points, resulting in lower resolution than optical images. Therefore, identification is more complex than with optical images and cannot be directly processed using optical methods, leading to lower identification accuracy. Regarding obstruction, SAR is mainly affected by electromagnetic and noise, while optical images are primarily affected by weather and environment. Because SAR is widely used in defense, SAR images are often affected by electromagnetic fields, noise, and buildings, as well as multiplicative speckle noise. This weakens the scattering points of some key structures on SAR targets, while simultaneously revealing strong scattering points from false targets—something optical methods cannot detect. Additionally, even slight perturbations can alter the overall structure of the scattering points.

[0007] The patent document "A High-Resolution Radar Target Recognition Algorithm Based on Deep Learning" (CN110969121A) discloses a VGG-Inception network structure, including a radar data preprocessing module, a feature extraction module, and a classifier module. It optimizes the model through supervised learning and learning rate decay, and replaces the Conv3 module of the original VGG network with the Inception module and a 1×1 convolution module. This preserves the differential phase information between data, reduces the number of parameters, and improves robustness. However, it does not address the extraction of information from occluded targets or the recognition of key features.

[0008] The patent document "A Small Sample Target Recognition Method for SAR Images Based on Gated Multi-Scale Matching Network" (CN113283390A) proposes a method based on a gated multi-scale matching network. By introducing a multi-scale feature extraction module and a gating unit, it improves the traditional matching network. It utilizes multi-scale features and gating units to select features according to different recognition tasks, thereby improving the model's generalization ability and recognition accuracy. However, it only improves the recognition of small targets and still lacks the ability to extract key features of occluded targets.

[0009] Therefore, there is an urgent need for a solution to extract the key structures of occluded targets, thereby overcoming the limitations of existing technologies and having significant engineering implications for improving SAR occluded target recognition systems. Summary of the Invention

[0010] To address the shortcomings of existing technologies, the purpose of this invention is to provide a target recognition system based on learning fraction Gabor transform and local scattering extraction network, thereby solving the problem of low recognition rate of occluded SAR targets in existing methods.

[0011] A target recognition system based on a learned fractional Gabor transform and a local scattering extraction network, provided by the present invention, includes: a learnable fractional Gabor transform module, a local scattering extraction module, and a feature recognition network;

[0012] The learnable fractional Gabor transform module performs filtering response fusion and feature extraction on occluded images to obtain the processed image;

[0013] The local scattering extraction module extracts key structural features from the processed image to obtain a feature image;

[0014] The feature image is optimized by using a feature recognition network to obtain the recognition result.

[0015] Preferably, the learnable fractional Gabor transform module is used in multiple scales, directions, and kernels to extract texture and detail information from the occluded image at different scales, directions, and convolution kernels, respectively. The kernel function is:

[0016]

[0017] in:

[0018] x′=xcosθ+ysinθ

[0019] y′=-xsinθ+ycosθ

[0020] Where s∈{0,…,V-1} represents the scaling factor;

[0021] V represents the total number of scales;

[0022] α represents the trainable order;

[0023] θ∈{θ1,…,θ L} represents the direction factor;

[0024] L represents the total number of directions;

[0025] x and y represent the original X and Y axis coordinates of the image, respectively;

[0026] x' and y' represent the new X and Y axis coordinates after rotation, respectively;

[0027] The receptive fields provided by different convolution kernels are used to perform channel-specific processing and then fusion.

[0028] Preferably, the local scattering extraction module includes:

[0029] Module M3.1 uses a dual-branch strategy to extract contours from the original images in the dataset. The first branch is adaptive thresholding, and the second branch is Gaussian filtering and Canny edge detection. The processing results of the two branches are organically fused to obtain the contours.

[0030] Module M3.2 performs region thresholding on the contour to generate a binary mask image:

[0031]

[0032] Among them, Area(W i ) represents the area of ​​the i-th contour region;

[0033] Q represents the set constant value;

[0034] W represents the i-th contour boundary in the occluded image. i The pixel values ​​of the enclosed region;

[0035] Module M3.3 outputs the fusion result of the binary mask image, the occlusion image, and the processed image, and performs element-wise multiplication to extract the overall contour image;

[0036] Module M3.4 detects key peak features of the overall contour image through a convolutional sliding window, determines whether the pixel value of the overall contour image within the window is greater than the adaptive threshold B of the sliding window. If it is greater, the pixel value remains unchanged; if it is less than or equal to the threshold B, the pixel value is set to 0, and the strong scattering points of the target are extracted.

[0037]

[0038] B=μ+kσ

[0039] Where k×k is a 3×3 convolution kernel, and k represents the kernel value;

[0040] I origin I p These represent the pixel values ​​of the overall outline image and the processed image, respectively.

[0041] μ and σ represent the mean and variance of pixels within a 3×3 window, respectively.

[0042] Preferably, the feature recognition network includes a network loss function, which consists of a triplet loss function, a center loss function, and a cross-entropy loss function.

[0043] The cross-entropy loss function is:

[0044]

[0045] Where N represents the total number of batches;

[0046] k∈{1,…M} represents the k-th category;

[0047] M represents the number of categories;

[0048] y i,k p i,k Let represent the true label and the predicted label corresponding to the k-th component in the i-th image, respectively.

[0049] The triplet loss function is:

[0050] L triplet =max{D(T)-D(R)+margin,0}

[0051]

[0052] Where margin represents the minimum interval parameter;

[0053] D(T) represents the distance between samples of the same type;

[0054] D(R) represents the distance between out-of-class samples;

[0055] f(·) represents mapping features from high-dimensional features to feature space;

[0056] A, T, and R represent anchor points, anchor point-like samples, and anchor point-dissimilar samples, respectively.

[0057] The central loss function is:

[0058]

[0059] Where, x i Represents input features;

[0060] This represents the center of the i-th category.

[0061] The network loss function is:

[0062] L total =L cross +λ1L triplet +λ2L center

[0063] Here, λ1 and λ2 both represent constants greater than 0.

[0064] The present invention provides a target recognition method based on learning fraction Gabor transform and local scattering extraction network, wherein the target recognition system based on learning fraction Gabor transform and local scattering extraction network is used to perform target recognition on images.

[0065] According to the present invention, a training method for a target recognition system based on a learning fraction Gabor transform and a local scattering extraction network is provided. The training of the target recognition system based on the learning fraction Gabor transform and local scattering extraction network includes:

[0066] Step S1: Obtain the original image based on the SAR dataset and generate the occlusion image;

[0067] Step S2: The learnable fractional Gabor transform module performs filtered response fusion and feature extraction on the occluded image to obtain the processed image;

[0068] Step S3: Instruct the local scattering extraction module to extract key structural features from the processed image to obtain a feature image;

[0069] Step S4: The feature image is optimized by a feature recognition network to obtain the recognition result.

[0070] Preferably, step S1 includes:

[0071] Step S1.1: Select the top 200 points with the highest pixel values ​​in the original image of the dataset, and set the occlusion directions as left, right, top and bottom for the dataset, and the occlusion ratios as 0%, 10%, 20%, 30%, 40% and 50% respectively.

[0072] Where left indicates processing column by column from the left to the right of the image;

[0073] "right" indicates processing column by column from the right side of the image to the left side.

[0074] top indicates processing line by line from the top of the image to the bottom of the image;

[0075] bottom indicates processing line by line from the bottom of the image to the top of the image;

[0076] Step S1.2: Perform occlusion processing according to the set occlusion direction. When encountering 200 selected points, set the pixel value of the corresponding points to 0 according to the set occlusion ratio to obtain the occluded image.

[0077] Step S1.3: Input the original image as a training set into the target recognition system based on the learning score Gabor transform and local scattering extraction network for training. Input the original image and the occluded image as a test set into the target recognition system based on the learning score Gabor transform and local scattering extraction network to test the recognition performance under different occlusion directions and proportions.

[0078] The SAR dataset is the MSATR dataset.

[0079] Preferably, in step S2, the learnable fractional Gabor transform module is used in multiple scales, directions, and kernels to extract texture and detail information from the occluded image or the original image at different scales, directions, and convolution kernels, respectively. The kernel function is:

[0080]

[0081] in:

[0082] x′=xcosθ+ysinθ

[0083] y′=-xsinθ+ycosθ

[0084] Where s∈{0,…,V-1} represents the scaling factor;

[0085] V represents the total number of scales;

[0086] α represents the trainable order;

[0087] θ∈{θ1,…,θ L} represents the direction factor;

[0088] L represents the total number of directions;

[0089] x and y represent the original X and Y axis coordinates of the image, respectively;

[0090] x' and y' represent the new X and Y axis coordinates after rotation, respectively;

[0091] The receptive fields provided by different convolution kernels are used to perform channel-specific processing and then fusion.

[0092] Preferably, step S3 includes:

[0093] Step S3.1: Use a two-branch strategy to extract contours from the original images in the dataset. The first branch is adaptive thresholding, and the second branch is Gaussian filtering and Canny edge detection. The processing results of the two branches are organically fused to obtain the contours.

[0094] Step S3.2: Perform region thresholding on the contour to generate a binary mask image.

[0095]

[0096] Among them, Area(W i ) represents the area of ​​the i-th contour region;

[0097] Q represents the set constant value;

[0098] W represents the i-th contour boundary in the occluded image. i The pixel values ​​of the enclosed region;

[0099] Step S3.3: Output the fusion result of the binary mask image and the occluded image or the original image and the processed image together, and extract the overall contour image by element-wise multiplication;

[0100] Step S3.4: Detect key peak features of the overall contour image using a convolutional sliding window, and determine whether the pixel value of the overall contour image within the window is greater than the adaptive threshold B of the sliding window. If it is greater, keep the pixel value unchanged; if it is less than or equal to, set the pixel value to 0, and extract the strong scattering points of the target.

[0101]

[0102] B=μ+kσ

[0103] Where k×k is a 3×3 convolution kernel, and k represents the kernel value;

[0104] I origin I p These represent the pixel values ​​of the overall outline image and the processed image, respectively.

[0105] μ and σ represent the mean and variance of pixels within a 3×3 window, respectively.

[0106] Preferably, the feature recognition network includes a network loss function, which consists of a triplet loss function, a center loss function, and a cross-entropy loss function.

[0107] The cross-entropy loss function is:

[0108]

[0109] Where N represents the total number of batches;

[0110] k∈{1,…M} represents the k-th category;

[0111] M represents the number of categories;

[0112] y i,k p i,k Let represent the true label and the predicted label corresponding to the k-th component in the i-th image, respectively.

[0113] The triplet loss function is:

[0114] L triplet =max{D(T)-D(R)+margin,0}

[0115]

[0116] Where margin represents the minimum interval parameter;

[0117] D(T) represents the distance between samples of the same type;

[0118] D(R) represents the distance between out-of-class samples;

[0119] f(·) represents mapping features from high-dimensional features to feature space;

[0120] A, T, and R represent anchor points, anchor point-like samples, and anchor point-dissimilar samples, respectively.

[0121] The central loss function is:

[0122]

[0123] Where, x i Represents input features;

[0124] This represents the center of the i-th category.

[0125] The network loss function is:

[0126] L total =L cross +λ1L triplet +λ2L center

[0127] Here, λ1 and λ2 both represent constants greater than 0.

[0128] Compared with the prior art, the present invention has the following beneficial effects:

[0129] 1. This invention utilizes learnable fractional Gabor transform to achieve multi-scale, multi-directional, and multi-kernel global information extraction from occluded images, overcoming the problem of insufficient global and local information extraction capabilities of traditional methods.

[0130] 2. This invention utilizes a local scattering extraction module to extract key remaining important features in occluded images, enhancing key feature structures such as target texture edges in occluded images, thereby improving the recognition accuracy of occluded SAR images.

[0131] 3. This invention can fully extract global and local information from occluded synthetic aperture radar images, thereby improving the model's ability to distinguish different SAR targets. It constructs occluded images with different directions and proportions. Under the MSTAR benchmark dataset, the recognition accuracy exceeds 90% in all occlusion directions and at high occlusion rates (50%). Attached Figure Description

[0132] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0133] Figure 1 This is a schematic diagram of the target recognition method based on learning score Gabor transform and local scattering extraction network;

[0134] Figure 2 These are optical images and SAR images of ten types of targets from MSTAR, as described in this embodiment of the invention.

[0135] Figure 3 These are SAR images with different occlusion directions and different occlusion ratios in embodiments of the present invention.

[0136] Figure 4 This is a schematic diagram of the learnable fractional Gabor transform module of the present invention;

[0137] Figure 5 This is a schematic diagram of the local scattering information extraction module of the present invention;

[0138] Figure 6 This is a schematic diagram of the recognition performance curves for different occlusion directions and occlusion ratios in an embodiment of the present invention.

[0139] Figure 7 This is a t-SNE visualization diagram of different loss terms in an embodiment of the present invention. Detailed Implementation

[0140] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0141] The present invention provides a target recognition system training method based on learning score Gabor transform and local scattering extraction network, which solves the problem of low recognition rate of occluded SAR targets. Figure 1 For example, specifically including:

[0142] Step S1: Generate occlusion images with different directions and proportions based on the acquired SAR dataset.

[0143] Specifically, it includes the following steps:

[0144] Step S1.1: Set different occlusion directions and occlusion ratios for the MSATR dataset, which is a high-resolution synthetic aperture radar (SAR) dataset. Here, left means processing column by column from the left to the right of the image, right means processing column by column from the right to the left of the image, top means processing row by row from the top to the bottom of the image, and bottom means processing row by row from the bottom to the top of the image.

[0145] The occlusion percentages were 0%, 10%, 20%, 30%, 40%, and 50%, respectively.

[0146] Step S1.2: Select the top 200 points with the highest pixel values ​​in the original image.

[0147] Step S1.3: Perform occlusion processing on the image according to different preset occlusion directions. When encountering 200 selected points, set their pixel values ​​to 0 according to the occlusion ratio, thereby realizing the simulated occlusion of the SAR image.

[0148] Step S2: Based on the learnable fraction Gabor transform module, perform multi-scale, multi-directional, and multi-kernel size filtering response fusion and feature extraction on the occluded image.

[0149] Specifically, the fractional Gabor transform is a generalized form of the traditional Gabor transform, and its mathematical expression is:

[0150]

[0151] Where B(x1,x2,α) represents a transformation kernel. To transform the angle, x1 and x2 are the coordinates in the spatial domain and the fractional Fourier transform domain, respectively.

[0152] The structure of the learnable fractional Gabor transform module is as follows: Figure 4 As shown, this is used to extract global information. Multi-scale, multi-directional, and multi-kernel methods are used to extract texture and detail information from images at different scales, directions, and convolution kernels, respectively, and to fuse them. This allows for the extraction of both texture and detail information from the target, which are then fused to obtain the overall global information. The trainable, multi-directional, and learnable multi-scale multi-kernel fractional Gabor transform kernel function is as follows:

[0153]

[0154] x′=xcosθ+ysinθ

[0155] y′=-xsinθ+ycosθ

[0156] Where s∈{0,…,V-1} is the scaling factor, V is the total number of scales, α is the trainable order, and θ∈{θ1,…,θ...} L} represents the orientation factor, L represents the total number of orientations, x and y represent the original X and Y coordinates of the image, and x' and y' represent the new X and Y coordinates after rotation.

[0157] According to the formula, the learnable fractional Gabor transform module extracts multi-scale and multi-directional features from SAR images using this convolutional kernel. It then performs channel-specific processing using the receptive fields provided by different convolutional kernels, ultimately extracting multi-directional and multi-scale spatial features suitable for SAR target characteristics, which is beneficial for occluded target identification and classification. Furthermore, the multi-directional nature of the Gabor transform allows for enhancing signal features in specific directions while suppressing the response in noise-dominant directions.

[0158] Step S3: Extract key structural features from the image processed by the Gabor module based on the local scattering extraction module.

[0159] Specifically, the results of the local scattering extraction module are as follows: Figure 5 As shown, the residual key scattering structure of the occluded SAR image is extracted. The residual key scattering structure is the overall contour shape and is also key to distinguishing different SAR targets. Step S3 includes the following steps:

[0160] Step S3.1: Contour extraction is performed using a dual-branch strategy to reduce the impact of multiplicative speckle noise unique to SAR. The first branch is adaptive thresholding, and the second branch is Gaussian filtering and Canny edge detection. The two branches are organically fused to avoid contour extraction errors that might occur with a single branch, thus affecting the extraction of the target's strong scattering characteristics.

[0161] Step S3.2: Perform region thresholding on the detected contours to generate a binary mask for the SAR image. The function is:

[0162]

[0163] Among them, Area(W i ) represents the area of ​​the i-th contour region, and Q is a constant value. W represents the i-th contour boundary in the SAR image. i The pixel values ​​of the enclosed region. Contour extraction can completely eliminate SAR multiplicative speckle noise, yielding a binary mask image of the target.

[0164] Step S3.3: Multiply the binary mask image element-wise with the combined output of Gabor transform and convolution transform to extract the overall contour image of the target area in the SAR image and suppress background noise interference.

[0165] Step S3.4: Detect key peak features using the convolutional sliding window method, the calculation formula being:

[0166]

[0167] B=μ+kσ

[0168] Where k×k is a 3×3 convolution kernel, k is a constant, and I origin and I p represents the pixel values ​​of the overall contour image and the local scattering extraction image, i.e., the processed image, respectively, and μ and σ are the mean and variance of the pixels within a 3×3 window, respectively.

[0169] By determining whether the pixels within the window are greater than the adaptive threshold B of the sliding window, the strong scattering points of the target for preliminary feature extraction are obtained, thereby providing key feature information for SAR targets affected by occlusion and enhancing the feature differentiation of different SAR targets.

[0170] Step S4: Optimize the feature recognition network based on the triplet loss function, center loss function and cross-entropy loss function to further improve the recognition effect of occluded images, optimize the feature images for recognition, and obtain the recognition result.

[0171] Specifically, the entire network loss function is composed of the triplet loss function, the center loss function, and the cross-entropy loss function, as follows:

[0172] L total =L cross +λ1L triplet +λ2L center

[0173] In the formula, λ1 and λ2 are constants greater than 0, and the cross-entropy loss function is:

[0174]

[0175] In the formula, N is the number of batches, k∈{1,…M} is the kth category, M is the total number of categories, and y i,k and p i,k Let represent the true label and the predicted label corresponding to the k-th component in the i-th image, respectively.

[0176] The triplet loss function:

[0177] L triplet =max{D(T)-D(R)+margin,0}

[0178]

[0179] In the formula, margin is the minimum margin parameter, D(T) is the distance between samples of the same class, D(R) is the distance between samples of different classes, f(·) represents mapping the features from high-dimensional features to the feature space, and A, T, and R represent the anchor point, the anchor point of the same class, and the anchor point of the different class, respectively. By minimizing the triplet loss function, it can be seen that the distance between samples of the same class will be smaller, and the distance between samples of different classes will be larger, thereby enhancing the separability between classes.

[0180] The central loss function is:

[0181]

[0182] In the formula, N is the number of batches, x i As input features, Let be the center of the i-th category. Introducing a center loss function can further enhance intra-class aggregation, thereby improving recognition accuracy.

[0183] Therefore, it can still extract global information and local key feature information under high occlusion conditions, while enhancing key feature structures such as target texture edges in occluded images, and can achieve excellent recognition accuracy when facing SAR images with different directions and high occlusion ratios.

[0184] In more preferred examples, an Intel i7-13700 CPU platform was used. The GPU was an NVIDIA RTX 4090 with 24GB of VRAM and 64GB of RAM. The PyTorch framework was used for model building, training, and testing, and the operating system was Windows. The hardware and software environment configurations are shown in Tables 1 and 2.

[0185] Table 1. Hardware Environment Configuration:

[0186] part model CPU RTX4090 GPU i7-13700 RAM 64G

[0187] Table 2. Software Environment Configuration:

[0188]

[0189] Based on the benchmark MSTAR dataset, which has a resolution of 0.3 × 0.3 m, and operates in the X-band with HH polarization. Table 3 shows the target categories and corresponding sample sizes at different elevation angles.

[0190] Table 3. Target categories and corresponding sample numbers at different pitch angles:

[0191] Serial Number Category Name 17° training set size 15° test set size 0 2S1 299 274 1 BMP2 232 196 2 BDRM2 298 274 3 BTR60 256 195 4 BTR70 233 196 5 D7 299 274 6 T62 299 273 7 T72 232 196 8 ZIL131 299 274 9 ZSU23 / 4 299 274

[0192] The baseline MSTAR optical image and SAR image in Experiment 1 are as follows: Figure 2 For example, Figure 3 These are SAR images of occlusion from different directions and scales in Experiment 1. For example... Figure 2 and Figure 3 As shown in the comparison, after step S1, some key strong scattering points and key structures of the target disappear.

[0193] After the learnable fractional Gabor transform module, it can be seen that the occluded image can obtain the multi-level information structure of the image target well after processing. After the local scattering information extraction module, it can be seen that key scattering structure information can be extracted from the image processed by the Gabor module.

[0194] like Figure 6 The figure shows the recognition performance curves for different occlusion directions and occlusion ratios. The recognition accuracy is greater than 91% under no occlusion, 10% occlusion, 20% occlusion, 30% occlusion, 40% occlusion, and 50% occlusion, demonstrating the effectiveness and superiority of the method. It can fully extract global and local information from occluded synthetic aperture radar images, while enhancing key feature structures such as target texture edges in occluded images, thereby improving the recognition accuracy of occluded SAR images. Furthermore, it can be seen that the recognition accuracy decreases with increasing occlusion rate. Table 4 lists the recognition performance under four different occlusion directions and high occlusion rates.

[0195] Table 4. Recognition performance under four different occlusion directions and high occlusion rate:

[0196]

[0197] In more preferred examples, a learnable fractional Gabor transform module and a local scattering extraction module are combined to extract dual information from occluded images. The study investigated the two modules in four scenarios: no learnable fractional Gabor transform module and no local scattering extraction module; no learnable fractional Gabor module; no local scattering extraction module; and both modules are present. Experiments were conducted to explore the effectiveness of this module for occluded image recognition. Table 5 shows the ablation experiments for the two modules.

[0198] Table 5. Ablation experimental results of the two modules:

[0199]

[0200] It can be seen that using the learnable fractional Gabor transform module and the local scattering extraction module simultaneously achieves significant improvements of over 17.37% and 32.79% respectively at occlusion rates of 30% and 40%. Using the FGT mode alone achieves significant improvements of over 16.67% and 27.76% respectively at occlusion rates of 30% and 40%. Even without the learnable fractional Gabor transform module, significant improvements of over 5.41% and 6.58% are achieved at occlusion rates of 30% and 40%, respectively. This fully demonstrates the effectiveness and superiority of the method in occluded target recognition.

[0201] In more optimized examples, a multivariate loss function combining the triplet loss function, center loss function, and cross-entropy loss function is used to study the loss term, specifically in four cases: loss with missing triplets and center loss, loss with missing triplets and center loss, and loss with neither loss function missing. Experiments are conducted on these four cases to... Figure 6 For example, here is a t-SNE visualization of different loss terms in Experiment 1.

[0202] like Figure 7 As shown, the intra-class spacing is significantly reduced after adding the center loss function; the inter-class spacing of the ten target classes is significantly increased after adding the triplet loss function; both the intra-class and inter-class spacing of the ten target classes are increased after adding both loss functions. Table 6 lists the ablation experiments for loss functions with different directions and high occlusion.

[0203] Table 6. Ablation experiment results with different loss functions:

[0204]

[0205] It can be seen that adding the two loss functions improves the efficiency by 1.89% and 5.03% respectively, adding the center loss function improves it by 1.28% and 3.38% respectively, and adding the triplet loss function improves it by 1.4% and 4.04% respectively.

[0206] The present invention also provides a target recognition system based on a learned fractional Gabor transform and a local scattering extraction network, trained by the target recognition system training method based on the learned fractional Gabor transform and local scattering extraction network, comprising: a learnable fractional Gabor transform module, a local scattering extraction module, and a feature recognition network.

[0207] The learnable fractional Gabor transform module performs filtering response fusion and feature extraction on occluded images to obtain the processed image;

[0208] The local scattering extraction module extracts key structural features from the processed image to obtain a feature image;

[0209] The feature image is optimized by using a feature recognition network to obtain the recognition result.

[0210] In more preferred embodiments, the learnable fractional Gabor transform module is used in multiple scales, directions, and kernels to extract texture and detail information from the occluded image at different scales, directions, and convolution kernels, respectively. The kernel function is:

[0211]

[0212] in:

[0213] x′=xcosθ+ysinθ

[0214] y′=-xsinθ+ycosθ

[0215] Where s∈{0,…,V-1} represents the scaling factor;

[0216] V represents the total number of scales;

[0217] α represents the trainable order;

[0218] θ∈{θ1,…,θ L} represents the direction factor;

[0219] L represents the total number of directions;

[0220] x and y represent the original X and Y axis coordinates of the image, respectively;

[0221] x' and y' represent the new X and Y axis coordinates after rotation, respectively;

[0222] The receptive fields provided by different convolution kernels are used to perform channel-specific processing and then fusion.

[0223] In more preferred embodiments, the local scattering extraction module includes:

[0224] Module M3.1 uses a dual-branch strategy to extract contours from the original images in the dataset. The first branch is adaptive thresholding, and the second branch is Gaussian filtering and Canny edge detection. The processing results of the two branches are organically fused to obtain the contours.

[0225] Module M3.2 performs region thresholding on the contour to generate a binary mask image:

[0226]

[0227] Among them, Area(W i ) represents the area of ​​the i-th contour region;

[0228] Q represents the set constant value;

[0229] W represents the i-th contour boundary in the occluded image.i The pixel values ​​of the enclosed region;

[0230] Module M3.3 outputs the fusion result of the binary mask image, the occlusion image, and the processed image, and performs element-wise multiplication to extract the overall contour image;

[0231] Module M3.4 detects key peak features of the overall contour image through a convolutional sliding window, determines whether the pixel value of the overall contour image within the window is greater than the adaptive threshold B of the sliding window. If it is greater, the pixel value remains unchanged; if it is less than or equal to the threshold B, the pixel value is set to 0, and the strong scattering points of the target are extracted.

[0232]

[0233] B=μ+kσ

[0234] Where k×k is a 3×3 convolution kernel, and k represents the kernel value;

[0235] I origin I p These represent the pixel values ​​of the overall outline image and the processed image, respectively.

[0236] μ and σ represent the mean and variance of pixels within a 3×3 window, respectively.

[0237] In more preferred embodiments, the feature recognition network includes a network loss function consisting of a triplet loss function, a center loss function, and a cross-entropy loss function.

[0238] The cross-entropy loss function is:

[0239]

[0240] Where N represents the total number of batches;

[0241] k∈{1,…M} represents the k-th category;

[0242] M represents the number of categories;

[0243] y i,k p i,k Let represent the true label and the predicted label corresponding to the k-th component in the i-th image, respectively.

[0244] The triplet loss function is:

[0245] L triplet =max{D(T)-D(R)+margin,0}

[0246]

[0247] Where margin represents the minimum interval parameter;

[0248] D(T) represents the distance between samples of the same type;

[0249] D(R) represents the distance between out-of-class samples;

[0250] f(·) represents mapping features from high-dimensional features to feature space;

[0251] A, T, and R represent anchor points, anchor point-like samples, and anchor point-dissimilar samples, respectively.

[0252] The central loss function is:

[0253]

[0254] Where, x i Represents input features;

[0255] This represents the center of the i-th category.

[0256] The network loss function is:

[0257] L total =L cross +λ1L triplet +λ2L center

[0258] Here, λ1 and λ2 both represent constants greater than 0.

[0259] The present invention provides a target recognition method based on learning fraction Gabor transform and local scattering extraction network, wherein the target recognition system based on learning fraction Gabor transform and local scattering extraction network is used to perform target recognition on images.

[0260] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A target recognition system based on learning fraction Gabor transform and local scattering extraction network, characterized in that, include: It can learn fractional Gabor transform modules, local scattering extraction modules, and feature recognition networks; The learnable fractional Gabor transform module performs filtering response fusion and feature extraction on occluded images to obtain the processed image; The local scattering extraction module extracts key structural features from the processed image to obtain a feature image; The feature image is optimized by using a feature recognition network to obtain the recognition result.

2. The target recognition system based on learning fraction Gabor transform and local scattering extraction network according to claim 1, characterized in that, The learnable fractional Gabor transform module is used in multiple scales, directions, and kernels to extract texture and detail information from the occluded image at different scales, directions, and convolution kernels. The kernel function is: in: x′=xcosθ+ysinθ y′=-xsinθ+ycosθ Where s∈{0,…,V-1} represents the scaling factor; V represents the total number of scales; α represents the trainable order; θ∈{θ1,…,θ L } represents the direction factor; L represents the total number of directions; x and y represent the original X and Y axis coordinates of the image, respectively; x' and y' represent the new X and Y axis coordinates after rotation, respectively; The receptive fields provided by different convolution kernels are used to perform channel-specific processing and then fusion.

3. The target recognition system based on learning fraction Gabor transform and local scattering extraction network according to claim 1, characterized in that, The local scattering extraction module includes: Module M3.1 uses a dual-branch strategy to extract contours from the original images in the dataset. The first branch is adaptive thresholding, and the second branch is Gaussian filtering and Canny edge detection. The processing results of the two branches are organically fused to obtain the contours. Module M3.2 performs region thresholding on the contour to generate a binary mask image: Among them, Area(W i ) represents the area of ​​the i-th contour region; Q represents the set constant value; W represents the i-th contour boundary in the occluded image. i The pixel values ​​of the enclosed region; Module M3.3 outputs the fusion result of the binary mask image, the occlusion image, and the processed image, and performs element-wise multiplication to extract the overall contour image; Module M3.4 detects key peak features of the overall contour image through a convolutional sliding window, determines whether the pixel value of the overall contour image within the window is greater than the adaptive threshold B of the sliding window. If it is greater, the pixel value remains unchanged; if it is less than or equal to the threshold B, the pixel value is set to 0, and the strong scattering points of the target are extracted. B=μ+kσ Where k×k is a 3×3 convolution kernel, and k represents the kernel value; I origin I p These represent the pixel values ​​of the overall outline image and the processed image, respectively. μ and σ represent the mean and variance of pixels within a 3×3 window, respectively.

4. The target recognition method based on learning fraction Gabor transform and local scattering extraction network according to claim 1, characterized in that, The feature recognition network includes a network loss function, which consists of a triplet loss function, a center loss function, and a cross-entropy loss function. The cross-entropy loss function is: Where N represents the total number of batches; k∈{1,…M} represents the k-th category; M represents the number of categories; y i,k p i,k Let represent the true label and the predicted label corresponding to the k-th component in the i-th image, respectively; The triplet loss function is: L triplet =max{D(T)-D(R)+margin,0} Where margin represents the minimum interval parameter; D(T) represents the distance between samples of the same type; D(R) represents the distance between out-of-class samples; f(·) represents mapping features from high-dimensional features to feature space; A, T, and R represent anchor points, anchor point-like samples, and anchor point-dissimilar samples, respectively. The central loss function is: Where, x i Represents input features; Indicates the center of the i-th category; The network loss function is: L total =L cross +λ1L triplet +λ2L center Here, λ1 and λ2 both represent constants greater than 0.

5. A target recognition method based on learning fraction Gabor transform and local scattering extraction network, characterized in that, The target recognition system based on learning fraction Gabor transform and local scattering extraction network as described in any one of claims 1-4 is used to perform target recognition on images.

6. A training method for a target recognition system based on a learning fraction Gabor transform and a local scattering extraction network, wherein the target recognition system based on the learning fraction Gabor transform and the local scattering extraction network is trained, characterized in that, include: Step S1: Obtain the original image based on the SAR dataset and generate the occlusion image; Step S2: The learnable fractional Gabor transform module performs filtered response fusion and feature extraction on the occluded image to obtain the processed image; Step S3: Instruct the local scattering extraction module to extract key structural features from the processed image to obtain a feature image; Step S4: The feature image is optimized by a feature recognition network to obtain the recognition result.

7. The training method for a target recognition system based on learning fraction Gabor transform and local scattering extraction network according to claim 6, characterized in that, Step S1 includes: Step S1.1: Select the top 200 points with the highest pixel values ​​in the original image of the dataset, and set the occlusion directions as left, right, top and bottom for the dataset, and the occlusion ratios as 0%, 10%, 20%, 30%, 40% and 50% respectively. Where left indicates processing column by column from the left to the right of the image; "right" indicates processing column by column from the right side of the image to the left side. top indicates processing line by line from the top of the image to the bottom of the image; bottom indicates processing line by line from the bottom of the image to the top of the image; Step S1.2: Perform occlusion processing according to the set occlusion direction. When encountering 200 selected points, set the pixel value of the corresponding points to 0 according to the set occlusion ratio to obtain the occluded image. Step S1.3: Input the original image as the training set into the target recognition system based on the learning fraction Gabor transform and local scattering extraction network for training. Input the original image and the occluded image as the test set into the target recognition system based on the learning fraction Gabor transform and local scattering extraction network to test the recognition performance under different occlusion directions and proportions. The SAR dataset is the MSATR dataset.

8. The training method for a target recognition system based on learning fraction Gabor transform and local scattering extraction network according to claim 6, characterized in that, In step S2, the learnable fractional Gabor transform module is used in multiple scales, directions, and kernels to extract texture and detail information from the occluded or original image at different scales, directions, and convolution kernels. The kernel function is: in: x′=xcosθ+ysinθ y′=-xsinθ+ycosθ Where s∈{0,…,V-1} represents the scaling factor; V represents the total number of scales; α represents the trainable order; θ∈{θ1,…,θ L } represents the direction factor; L represents the total number of directions; x and y represent the original X and Y axis coordinates of the image, respectively; x' and y' represent the new X and Y axis coordinates after rotation, respectively; The receptive fields provided by different convolution kernels are used to perform channel-specific processing and then fusion.

9. The training method for a target recognition system based on learning fraction Gabor transform and local scattering extraction network according to claim 6, characterized in that, Step S3 includes: Step S3.1: Use a two-branch strategy to extract contours from the original images in the dataset. The first branch is adaptive thresholding, and the second branch is Gaussian filtering and Canny edge detection. The processing results of the two branches are organically fused to obtain the contours. Step S3.2: Perform region thresholding on the contour to generate a binary mask image. Among them, Area(W i ) represents the area of ​​the i-th contour region; Q represents the set constant value; W represents the i-th contour boundary in the occluded image. i The pixel values ​​of the enclosed region; Step S3.3: Output the fusion result of the binary mask image and the occluded image or the original image and the processed image together, and extract the overall contour image by element-wise multiplication; Step S3.4: Detect key peak features of the overall contour image using a convolutional sliding window, and determine whether the pixel value of the overall contour image within the window is greater than the adaptive threshold B of the sliding window. If it is greater, keep the pixel value unchanged; if it is less than or equal to, set the pixel value to 0, and extract the strong scattering points of the target. B=μ+kσ Where k×k is a 3×3 convolution kernel, and k represents the kernel value; I origin I p These represent the pixel values ​​of the overall outline image and the processed image, respectively. μ and σ represent the mean and variance of pixels within a 3×3 window, respectively.

10. The training method for a target recognition system based on learning fraction Gabor transform and local scattering extraction network according to claim 6, characterized in that, The feature recognition network includes a network loss function, which consists of a triplet loss function, a center loss function, and a cross-entropy loss function. The cross-entropy loss function is: Where N represents the total number of batches; k∈{1,…M} represents the k-th category; M represents the number of categories; y i,k p i,k Let represent the true label and the predicted label corresponding to the k-th component in the i-th image, respectively; The triplet loss function is: L triplet =max{D(T)-D(R)+margin,0} Where margin represents the minimum interval parameter; D(T) represents the distance between samples of the same type; D(R) represents the distance between out-of-class samples; f(·) represents mapping features from high-dimensional features to feature space; A, T, and R represent anchor points, samples of the same class as anchor points, and samples of different classes as anchor points, respectively; the central loss function is: Where, x i Represents input features; Indicates the center of the i-th category; The network loss function is: L total =L cross +λ1L triplet +λ2L center Here, λ1 and λ2 both represent constants greater than 0.

Citation Information

Patent Citations

  • High-resolution radar target recognition algorithm based on deep learning

    CN110969121A

  • SAR image small-sample target identification method based on gating multi-scale matching network

    CN113283390A

  • Covered pedestrian re-identification method of multi-scale hypergraph connection with KAN structure

    CN119169525A

Cited By

  • Image restoration method and electronic equipment

    CN121213426A