Vehicle target identification method based on characteristic decomposition and cumulative learning

Through the FDRM ST-PA_RCNN algorithm based on deep neural network, combined with feature decomposition and cumulative learning strategies, the problem of low recognition rate in traditional vehicle detection and recognition methods in complex backgrounds is solved, and higher detection accuracy and robustness are achieved.

CN120279436AActive Publication Date: 2025-07-08FUDAN UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510452749.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-08
Filing Date
2025-04-11
Publication Date
2025-07-08
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

Traditional vehicle detection and identification methods perform poorly in feature extraction, recognition rate and detection robustness, especially when vehicle targets have small physical size, high mobility and complex backgrounds.

Method used

The FDRM ST-PA_RCNN algorithm based on deep neural network is adopted, and through feature decomposition and accumulation learning strategies, combined with SAR image data sets and feature redistribution models, a vehicle target recognition method is constructed, including feature fusion method of feature decomposition and weight redistribution, reducing feature redundancy and improving detection performance.

Benefits of technology

It significantly improves the accuracy of vehicle target detection and suppresses false alarm performance, improves the effectiveness and robustness of the model, especially in complex backgrounds and high noise environments, which can more accurately identify vehicle targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279436A_ABST
    Figure CN120279436A_ABST
Patent Text Reader

Abstract

The invention relates to the field of vehicle target recognition, and discloses a vehicle target recognition method based on characteristic decomposition and cumulative learning, which comprises the following steps: constructing a data set by using a land SAR image vehicle target data set; vehicle target scattering feature extraction and enhancement are carried out; an ST-PARCNN pixel level fusion model is constructed; training a feature level fusion model based on an FDRM module; and testing the performance of the model by using the test data and outputting a result. According to the invention, the vehicle target detection capability of the SAR image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle target recognition. Specifically, it relates to a vehicle target recognition method based on feature decomposition and cumulative learning. Background Art

[0002] Due to the characteristics of small physical size, high mobility, and complex background of vehicle targets, traditional vehicle detection and recognition methods perform poorly in aspects such as feature extraction, recognition rate, and detection robustness.

[0003] In the past decade, deep learning technology has become an important milestone in the development of artificial intelligence. Its innovative breakthroughs have significantly promoted the innovation process in multiple technical fields. This technology system has been successfully applied in multiple fields such as speech recognition, text understanding, visual information processing, dynamic video analysis, and multimedia content analysis, demonstrating excellent application value. Different from the traditional method of pattern recognition that relies on artificial feature engineering, this technology autonomously extracts feature information from massive data, realizing a fundamental change in the way of knowledge representation. This data-driven learning mechanism enables the model to have stronger representation ability and higher computing efficiency, not only improving the model performance, but also providing an innovative solution for processing high-dimensional and unstructured data.

[0004] This method proposes the FDRM ST-PA_RCNN algorithm for vehicle target recognition, using the ST-PA_RCNN based on the ReDet and Oriented RCNN frameworks of deep neural networks as the backbone network, and adding an FDRM module (feature decomposition and feature reassignment model) migration fusion block to reduce feature redundancy through weight reassignment, forming a feature fusion method based on feature decomposition and weight reassignment using a cumulative learning strategy. Verifying this algorithm on a typical target detection dataset, the comprehensive performance has a huge improvement compared with the baseline model ST-PA_RCNN, showing high effectiveness and robustness, and having better detection and false alarm suppression performance. Summary of the Invention

[0005] Therefore, this application provides a vehicle target recognition method based on feature decomposition and cumulative learning to solve the above problems.

[0006] An embodiment of this application provides a vehicle target recognition method based on feature decomposition and cumulative learning, including:

[0007] Using a land SAR (synthetic aperture radar) image vehicle target dataset for dataset construction;

[0008] Performing vehicle target scattering feature extraction and enhancement;

[0009] Constructing an ST-PA_RCNN (region convolutional neural network) pixel-level fusion model;

[0010] Feature-level fusion model training based on the FDRM (Feature Decomposition and Weight Reallocation Model) module;

[0011] Use test data to test the performance of the model and output the results.

[0012] Preferably, the use of the land SAR image vehicle target data set for data set construction includes:

[0013] According to the characteristics of satellite left and right side-looking and ascending and descending orbits, combined with the layout direction of the target in the distribution base, select the key-scene areas to collect sample data;

[0014] After standardization processing, use the target image distribution test to check whether the sample data meets the sample distribution index requirements;

[0015] If it meets the requirements, complete the data set construction;

[0016] If it does not meet the requirements, analyze the missing sample observation parameters, use orbital simulation to perform distribution position coverage analysis on the missing sample observation parameters, and form an optimal reconnaissance imaging plan that can be completed in a short time;

[0017] Use a satellite platform with high SAR timeliness for rapid reconnaissance imaging and perform preprocessing to form training and test samples.

[0018] Preferably, the extraction and enhancement of vehicle target scattering characteristics include:

[0019] Extract the texture features of the SAR image through four algorithms: SAR-SIFT (Scale-Invariant Feature Transform), SAR-HOG (Histogram of Oriented Gradients), NSLP (Non-Subsampled Laplacian Pyramid), and LC (Linear Color Salience);

[0020] Use the SAR-Harris detector (Synthetic Aperture Radar Hollis Corner Detector) and the OPTICS clustering method (Density Clustering Algorithm Based on Reachability and Core Distance) to enhance the extracted features.

[0021] Preferably, the construction of the ST-PA_RCNN pixel-level fusion model includes:

[0022] Construct the ST-PA_RCNN backbone network, combine the Swin Transformer (Sliding Window Hierarchical Attention Network) and PA_FPN (Path Aggregation Multi-Level Feature Pyramid Structure), fuse the original image and the feature map after scattering feature enhancement through the channel fusion method, and train the target detection model.

[0023] Preferably, the feature-level fusion model training based on FDRM includes:

[0024] Based on the ST-PA_RCNN, the FDRM module is introduced by the feature fusion method Dual-Branch FPN (Dual-Branch Feature Pyramid Network). Feature redundancy is reduced and the difference in the feature space is enhanced through feature decomposition and weight redistribution methods.

[0025] Preferably, testing the performance of the model using test data and outputting the results includes:

[0026] Inputting the images in the test dataset into the trained FDRM ST-PA_RCNN model (including pixel-level and feature-level fusion models) for object detection inference, and calculating the performance metrics of the model;

[0027] Evaluating the performance of the model by calculating performance metrics such as precision, recall, and mean average precision (mAP);

[0028] Outputting the detection results (localization boxes and target categories), and visualizing the detection results.

[0029] Preferably, it further includes:

[0030] Performing channel fusion on the original SAR image and the feature map after enhanced scattering features to form multi-channel input data.

[0031] Preferably, it further includes:

[0032] Reducing feature redundancy and enhancing the difference in the feature space through feature decomposition and weight redistribution methods.

[0033] Preferably, the present application also provides an electronic device, including a processor and a memory, wherein: a computer program is stored in the memory; when the computer program in the memory is executed by the processor, the electronic device can implement any of the described methods.

[0034] Preferably, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium includes a computer program, when the computer program runs on an electronic device, the electronic device is enabled to execute the method described in any implementation manner of the first aspect.

[0035] Preferably, an embodiment of the present application provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method described in any implementation manner of the first aspect.

[0036] The vehicle target recognition method, electronic device, computer-readable storage medium, and computer program product based on feature decomposition and cumulative learning provided by the embodiments of the present application first use the land SAR image vehicle target data set to construct the data set; then, extract and enhance the vehicle target scattering features; next, construct the ST-PA_RCNN pixel-level fusion model; furthermore, train the feature-level fusion model based on FDRM; finally, use the test data to test the performance of the model and output the results. The present application can improve the vehicle target detection ability of SAR images. Description of the Drawings

[0037] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0038] Figure 1 is a flowchart of a vehicle target recognition method based on feature decomposition and cumulative learning shown in Embodiment 1;

[0039] Figure 2 is a flowchart of another vehicle target recognition method based on feature decomposition and cumulative learning shown in Embodiment 1;

[0040] Figure 3 is the general schematic diagram of the present application;

[0041] Figure 4 is the schematic diagram of the MIX_MSTAR simulation data set;

[0042] Figure 5 is the general flowchart of the present invention, from left to right are the data acquisition and data set construction part, the model training part, and the model testing part;

[0043] Figure 6 is the schematic diagram of the FDRM module;

[0044] Figure 7 is the schematic diagram of the visualization of the vehicle target detection result in the present invention. Detailed Embodiments

[0045] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the scope of protection of the present application.

[0046] It should be noted that the terms "first", "second", etc. in the description, claims, and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0047] Embodiment 1

[0048] An embodiment of this application provides a vehicle target recognition method based on feature decomposition and cumulative learning. Refer to Figure 1 the flowchart of a vehicle target recognition method based on feature decomposition and cumulative learning shown in the figure. The specific implementation scheme is as follows:

[0049] Step S101: Use the vehicle target dataset of land SAR images to construct a dataset;

[0050] Step S102: Extract and enhance the scattering features of vehicle targets;

[0051] Step S103: Construct an ST-PA_RCNN pixel-level fusion model;

[0052] Step S104: Train a feature-level fusion model based on the FDRM module;

[0053] Step S105: Use test data to test the performance of the model and output the results.

[0054] Another embodiment of this application also provides a vehicle target recognition method based on feature decomposition and cumulative learning. Refer to Figure 2 the flowchart of a vehicle target recognition method based on feature decomposition and cumulative learning shown in the figure. The specific implementation scheme is as follows:

[0055] Step S201: According to the characteristics of satellite left and right side views and ascending and descending orbits, and combined with the layout direction of the target in the distribution base, select the key scene areas of concern for sample data collection.

[0056] Step S202: After standardization processing, use the target map distribution to check whether the sample data meets the requirements of the sample distribution index.

[0057] Step S203: If it meets the requirements, complete the construction of the dataset.

[0058] Step S204, if not, analyze the missing sample observation parameters, and use orbital simulation to conduct distribution position coverage analysis on the observation parameters of the missing samples to form an optimal detection and imaging plan that can be completed in a short time.

[0059] Step S205, use satellite platforms with high timeliness such as SAR for rapid detection and imaging, and perform preprocessing to form training and test samples.

[0060] Step S206, extract the texture features of the SAR image through four algorithms: SAR-SIFT, SAR-HOG, NSLP, and LC.

[0061] Step S207, use the SAR-Harris detector and the OPTICS clustering method to enhance the extracted features.

[0062] Step S208, construct the ST-PA_RCNN backbone network, combine Swin Transformer and PA_FPN, and fuse the original image with the feature map after scattering feature enhancement through the channel fusion method, and train the object detection model.

[0063] Step S209, based on the ST-PA_RCNN, introduce the FDRM module based on the feature fusion method Dual-Branch FPN, and reduce feature redundancy and enhance the difference in the feature space through feature decomposition and weight redistribution methods.

[0064] Step S210, input the images in the test dataset into the trained FDRM ST-PA_RCNN model (including pixel-level and feature-level fusion models) for object detection inference, and calculate the performance metrics of the model.

[0065] Step S211, evaluate the performance of the model by calculating performance metrics such as precision, recall, and mean average precision (mAP).

[0066] Step S212, output the detection results (localization boxes and target categories), and visualize the detection results.

[0067] Step S213, fuse the original SAR image with the feature map after scattering feature enhancement through channels to form multi-channel input data.

[0068] Step S214, reduce feature redundancy and enhance the difference in the feature space through feature decomposition and weight redistribution methods.

[0069] The vehicle target recognition method based on feature decomposition and cumulative learning provided by the embodiments of this application first constructs a data set using the vehicle target data set of land SAR images; then, extracts and enhances the scattering features of vehicle targets; next, constructs an ST-PA_RCNN pixel-level fusion model; furthermore, trains a feature-level fusion model based on the FDRM module; finally, tests the performance of the model using test data and outputs the results. This application can improve the vehicle target detection ability of SAR images.

[0070] Construct a data set using the collected high-resolution SAR images. The specific implementation method is as follows:

[0071] Step 1: Combine the layout direction of the target at the distribution base, and select the key-scenario areas to collect vehicle target SAR image data through satellites.

[0072] Step 2: Standardize the collected SAR images.

[0073] Step 3: Use the target map distribution to check whether the sample data meets the requirements of the sample distribution index. If it meets the requirements, the data set construction is completed. If it does not meet the requirements, it is necessary to analyze the missing sample observation parameters, and use orbital simulation to analyze the distribution position coverage of the missing sample observation parameters to form an optimal detection plan that can be completed in a short time.

[0074] Step 4: Preprocess the image data to form training and test samples.

[0075] Step 5: Use an algorithm for batch operation to downsample these high-resolution pictures to a unified size (512*512) as the data set.

[0076] Step 6: Divide the training set and the test set in a ratio of 8:2.

[0077] Extract and enhance the scattering features of vehicle targets. The specific implementation method is as follows:

[0078] Step 1: Use the SAR-SIFT, SAR-HOG, NSLP, and LC algorithms to extract the vehicle target features respectively. Among them, the SAR-SIFT algorithm enhances the robustness to multiplicative noise in SAR images by improving gradient calculation and is used to extract key-point features; the SAR-HOG algorithm extracts stable structural features through ratio gradient calculation and is used to extract texture features; the NSLP algorithm retains the spatial resolution of multi-scale texture features through non-downsampled Laplacian pyramids and is used to extract contour features; the LC algorithm extracts visual saliency features through a linear color model.

[0079] Step 2: Use the SAR-Harris detector and the OPTICS clustering method to enhance the extracted features.

[0080] Step 3: Perform channel fusion on the original image and the feature map after scattering feature enhancement to construct multi-channel input data.

[0081] Then, construct a pixel-level fusion model ST-PA_RCNN as the baseline (benchmark model), and the specific implementation method is as follows:

[0082] Step 1: Construct the ST-PA_RCNN backbone network, combine Swin Transformer and PA-FPN, and initialize the pre-trained parameters;

[0083] Step 2: Input the multi-channel data after channel fusion into the network, and train the object detection model through the pixel-level fusion module;

[0084] Step 3: Optimize the model parameters using cross-entropy loss and Smooth L1 loss.

[0085] Then, construct a feature-level fusion model FDRMST-PA_RCNN based on the FDRM module on the basis of the baseline (benchmark model), and use the training data for training. The specific implementation method is as follows:

[0086] Step 1: Construct a Dual-Branch PA-FPN module, which is shown as PA-FPN on the left and right sides respectively. Then, input the dual-branch features {P2, P3, P4, P5} obtained from PA-FPN into the feature fusion block to obtain a new feature map {M2, M3, M4, M5}.

[0087] Step 2: Construct an improved feature fusion block, the FDRM module. The FDRM module can reduce more feature redundancy. In addition, through the cumulative learning strategy, re-weighting of the two-branch features can be achieved, enhancing the interaction between the dual feature spaces and enabling the network to learn and utilize scattering characteristics more effectively.

[0088] Step 3: Input the features extracted by DB-PA-FPN from the original image and the feature image after scattering enhancement into the FDRM module for feature decomposition and re-weighting. Construct orthogonal loss and compactness loss to reduce feature redundancy and enhance feature space difference; adopt a dynamic weighting strategy for weight re-distribution to fuse the dual-branch features to form new features, and adjust the weighting coefficient μ (such as constant, linear, parabolic strategy, etc.) to generate hybrid features.

[0089] The input of the feature decomposition module is the feature map output by the feature extraction module, aiming to decompose the input feature map. This module contains two parallel branches, and decomposes the input features into foreground features and background features relying on the orthogonal constraint and reconstruction constraint between the two branches. The decomposed background features include background features such as complex background clutter, while the foreground features contain the target of interest. The two branches have the same network structure, both containing 3 convolutional layers, followed by ReLU layers after each convolutional layer. The convolutional kernel sizes of the 3 convolutional layers are all 3×3, the convolutional strides are all 1, and the number of convolutional kernels is 512. Since the parameters of the two branches are not shared, they are respectively responsible for foreground feature extraction and background feature extraction. Finally, the obtained foreground features and background features are added and input into the decoder for image reconstruction. The decoder contains 5 deconvolutional layers, where the first 4 deconvolutional layers are followed by ReLU layers. The convolutional kernel sizes of the 5 deconvolutional layers are all 3×3, the convolutional strides are 1, 2, 2, 2, 1 respectively, and the number of convolutional kernels are 512, 256, 128, 64, 3 respectively. The outputs of the feature decomposition module, the foreground features and background features, will be input into the subsequent feature weight reallocation module for weighted distribution.

[0090] The orthogonal loss function is defined as follows:

[0091]

[0092] where N represents the number of training samples in each batch. f n1 represents the column vector converted from the differential feature or background feature obtained by decomposing the i-th sample. f n2 represents the column vector converted from the interference feature or foreground feature obtained by decomposing the i-th sample, and T represents the transpose of the vector.

[0093] The compact loss function is defined as follows:

[0094]

[0095] where ‖·‖2 represents the L2 norm. Ii,j represents the decomposition feature when j = 1, …, M (M = 2 in our method). c j represents the center of the j-th decomposition vector, and it is updated in each batch. In this way, the variation of the same decomposition feature is minimized.

[0096] The final loss function is:

[0097] L total = L cls + L loc + λ opl L opl + λ cmpt L cmpt

[0098] where λ opl and λ cmpt are two balance parameters, set to 1 and 0.1 respectively.

[0099] L cls is the classification loss, which is used to measure the accuracy of the model when predicting the vehicle target category (or tasks such as binary classification of target and non - target).

[0100] L loc is the localization loss, which is used to measure the accuracy of the model when performing regression prediction on the vehicle bounding box (position, size, orientation, etc.).

[0101] The feature weight re - distribution module based on cumulative learning fuses the foreground and background features obtained after the feature decomposition module before sending them to the multi - scale detection module. This module re - distributes the weights of the input foreground and background features, and under different cumulative learning strategies, the focus of network training gradually shifts from background features to foreground features. The parameter μ is introduced as follows:

[0102] F r = μF fore +(1 - μ)F back

[0103] where F r represents the fused feature of the foreground and background features. F fore represents the foreground feature, that is, the feature vector or feature map contained in the target of interest (vehicle) focused by the feature decomposition module. F back represents the background feature, that is, the region features such as the image background or interference noise that do not belong to the vehicle body extracted by the same decomposition module. And we can perform multiple cumulative learning strategies to explore the best method for the best performance of object detection. Given the current training epoch T and the total training epochs T max , μ is different for different strategies. The constant strategy takes 0.5, the linear strategy takes T / T max , and the parabolic strategy takes (T / T max ) 2 . The output of the feature weight re - distribution module is used as the input of the multi - scale detection module to predict the class label and location information.

[0104] Step 4: Connect the obtained mixed features in channels and finally perform downsampling for output.

[0105] Finally, use the test data to test the performance of the model. The specific implementation method is as follows:

[0106] Step 1: Select the images in the test dataset to ensure that these images have not participated in the training process;

[0107] Step 2: Preprocess the test image, including downsampling to a unified size (such as 512x512 pixels) and scattering feature extraction;

[0108] Step 3: Input the preprocessed image into the trained ST-PA_RCNN model (including pixel-level and feature-level fusion models) for object detection inference; the input slice size of the dataset is 512x512, and the Stochastic Gradient Descent (SGD) optimizer is used with an initial learning rate of 5×10 -3 , a momentum of 0.9, and a weight decay of 1.0×10 -4 . Convergence is achieved after 56 rounds of training.

[0109] Step 4: Calculate the performance metrics of the model, including mean Average Precision (mAP), Precision, and Recall. The calculation formulas are as follows: (where TP represents true positives, FP represents false positives, and FN represents false negatives).

[0110] Precision:

[0111] Recall:

[0112] Average Precision (AP):

[0113] Mean Average Precision (mAP):

[0114] Step 5: Analyze the test results, compare the performance differences of different fusion methods (pixel-level, feature-level) on different datasets, and verify the effectiveness and robustness of the model.

[0115] Step 6: Output the detection results, output the bounding boxes and target categories, and visualize the detection results.

[0116] Step 7: The final pixel-level fusion method ST-PA_RCNN can achieve a MAP of 95.72% in the vehicle object detection task, while the model of the feature-level fusion method FDRM ST-PA_RCNN has a 2.03% higher MAP compared to ST-PA_RCNN. The comparison test results are shown in Table 1.

[0117]

[0118]

[0119] Table 1: Comparison test results

[0120] Example 2

[0121] In the FDRM module mode, the Adaptive Feature Gain Allocation Algorithm usually obtains the "Foreground Feature" F g and the "Background Feature" F b in two branches. By defining the "Adaptive Gain Function (AGF)" and introducing the parameter α, the Adaptive Feature Gain Allocation Algorithm automatically amplifies or reduces the weights of the foreground / background during the training process to achieve adaptability to different environments.

[0122] The execution process of the Adaptive Feature Gain Allocation Algorithm includes:

[0123] Receiving F g and F b output from the FDRM module, as well as the initial α (usually set to 0.5 or other empirical values);

[0124] First, perform σ(·) transformation on F g and F b respectively to enhance stability and non-linear representation;

[0125] Obtain W AGF according to the adaptive gain function. If α is slightly greater than 0.5, the foreground feature will be moderately enhanced; if α is significantly less than 0.5, the suppression of the background feature will be emphasized;

[0126] Regard W AGF as the new feature map or channel, and perform concatenation, superposition, or subsequent convolution operations with the original F g or other network features.

[0127] After each epoch or specific training step, by observing metrics such as the mAP of the validation set, the total loss, or the foreground recall rate, use gradient or search strategies to fine-tune α to adapt to the current scenario (noise level, target distribution, etc.).

[0128] As the training process progresses, α will tend to a steady state or change dynamically with external feedback, thus ensuring that the "gain" for the vehicle or the "suppression of the background" is always maintained at a reasonable level in different environments.

[0129] The following is the adaptive gain function (Formula 1):

[0130] W AGF (F g ,F b ; α) = α × σ(F g ) - (1 - α)σ(F b )

[0131] where F g represents the foreground feature, usually from the branch that focuses on the vehicle target after the decomposition of the FDRM module;

[0132] Fb Represents background features, mainly including ground clutter, noise, etc.;

[0133] σ(·) is the effective value obtained after an activation function or normalization operation, such as ReLU, LeakyReLU, or BatchNorm;

[0134] α ∈ (0, 1) is an adjustable gain coefficient that will be adaptively updated according to the loss or validation set metrics during the training phase;

[0135] W AGF Is the adaptive weighting result output by the adaptive feature gain allocation algorithm. It retains more foreground features in the way of weighted difference and partially reduces the energy of background features.

[0136] The technical effect is:

[0137] When the background noise increases, it can automatically lower the weight of F b and moderately amplify F g ;

[0138] In the case of a relatively clean background but a large number of targets, α can also be adjusted to a balanced value to avoid over-weakening the background and missing other potential information.

[0139] For extreme noise scenarios, relying solely on adaptive gain may not be sufficient to accurately lock small targets; when the difference between the foreground and the background is already weak, simply "increasing the foreground and reducing the background" cannot guarantee the retention of local details. Therefore, the cross-branch residual alignment algorithm proposes to construct a "residual alignment" operation between the foreground / background branches and introduce a parameter β to explicitly align the differences between the two during network training to avoid the vehicle features being widely submerged by the background.

[0140] The execution process of the cross-branch residual alignment algorithm includes:

[0141] The input is F g 、F b and α;

[0142] Initialize β, which can be set to a value in the range of 0.3 to 0.8, and can be determined specifically through grid search or experience;

[0143] First calculate αF b -F g . If this value is positive, it means that the background features "exceed" the foreground in a certain sense; if it is negative, it means that the foreground is stronger;

[0144] Multiply the residual from the previous step by β and add it to F g to get R CBRARThis process essentially injects a certain amount of background correction into the foreground branch, enabling small target features to be more robustly retained under huge background noise.

[0145] Use R CBRAR as the new foreground representation or further fuse it with background information to improve the accuracy of target detection.

[0146] Similar to α, β can also be fine-tuned according to the loss during network training to ensure that the residual injection does not over-correct and cause serious damage to the foreground.

[0147] Define the cross-branch residual alignment formula (Formula 2) as follows:

[0148] R CBRA (F g , F b ; α, β) = F g + β × (αF b - F g )

[0149] where α is consistent with the adaptive feature gain allocation algorithm;

[0150] β ∈ (0, 1) is the cross-branch residual coefficient, which determines the strength of residual alignment;

[0151] αF b - F g constructs the cross-branch residual signal: if the background feature is much larger than the foreground, this term presents a negative compensation form; if the foreground itself dominates, the residual tends to be small;

[0152] Finally, R CBRAR is the foreground feature map after alignment, which is equivalent to injecting a part of the information from the background branch onto F g and regulating its influence amplitude by β.

[0153] Technical effects:

[0154] In the case of high noise or partial occlusion, through the residual form of the "background branch - foreground branch", the network can perform additional differential compensation for small targets;

[0155] When the background noise changes violently, the cross-branch alignment mechanism will intelligently correct the network's focus of attention and reduce the interference of sudden background changes on target recognition.

[0156] If the adaptive feature gain allocation algorithm or the cross-branch residual alignment algorithm is used alone, it may be one-sided, or each has its own advantages and disadvantages in different stages. In order to organically combine the two, a cascaded linkage fusion algorithm is proposed, which uses the parameter γ to adaptively combine the feature gain allocation algorithm and the cross-branch residual alignment algorithm, so as to take into account the dual advantages of gain amplification and residual alignment in extreme scenarios.

[0157] The execution process of the cascaded linkage fusion algorithm includes:

[0158] Input W AGF (F g ,F b ; α) and R CBRA (F g ,F b ; α, β);

[0159] γ can be selected according to prior experience (such as 0.5) or through automated search;

[0160] Multiply γ by W AGF (F g ,F b ; α) and (1 - γ) × R CBRA (F g ,F b ; α, β) are added together to generate F Cascade ;

[0161] Output: F Cascade , which can be regarded as the final comprehensive feature map and then sent to FPN or the detection head for vehicle detection;

[0162] According to the performance of the training set or validation set, on the basis of fine-tuning α and β, γ can also be globally tuned regularly to observe its impact on indicators such as mAP, Precision, and Recall, so as to find the optimal configuration.

[0163] The cascaded linkage fusion formula (Formula 3) is as follows:

[0164] F Cascade = γ × W AGF (F g ,F b ; α)+(1 - γ) × R CBRA (F g ,F b ; α, β)

[0165] Among them, W AGF (F g ,F b ; α) is the adaptive gain result generated by the adaptive feature gain allocation algorithm;

[0166] R CBRA (F g ,F b ; α, β) is the cross-branch residual result generated by the cross-branch residual alignment algorithm;

[0167] γ∈(0,1), is the cascade fusion factor, which determines whether the final output is more inclined to the adaptive feature gain allocation algorithm or the cross-branch residual alignment algorithm;

[0168] F Cascade The comprehensive features produced by the cascade linkage fusion algorithm have the advantages of adaptive gain and detailed compensation of residual alignment.

[0169] Technical effect:

[0170] In a variety of complex scenarios, if the vehicle target is extremely small or severely occluded, the output of the adaptive feature gain allocation algorithm can first increase the target weight; if the background noise fluctuates significantly, the cross-branch residual alignment algorithm can also make timely corrections and compensations during this process; the two are fused through the gamma cascade, allowing the network to maintain high sensitivity while taking into account detail alignment;

[0171] Compared with single foreground / background weighting or residual alignment, the cascade linkage fusion algorithm is compatible with a variety of extreme situations, thereby ensuring vehicle detection performance in more complex application scenarios.

[0172] In order to enable those skilled in the art to fully implement this embodiment 2, a clear step-by-step description will be given below from the perspective of data preparation to network training and how to combine it with embodiment 1.

[0173] The second embodiment is consistent with the first embodiment in terms of underlying data and basic feature processing, namely:

[0174] Use satellite or airborne SAR systems to obtain high-resolution images to ensure that the training / testing data covers different angles, resolutions, and terrain features;

[0175] If necessary, missing samples are supplemented through target map distribution testing, orbit simulation, etc.;

[0176] Use SAR-SIFT, SAR-HOG, NSLP, LC, etc. to perform multi-channel fusion on the original image to retain multi-dimensional information such as texture, structure, and saliency;

[0177] In the feature level model stage, the FDRM module is executed to obtain the foreground feature F g With background features F b In Example 1, the FDRM module has detailed how to reduce redundancy and improve feature discrimination based on orthogonal loss and compact loss. So far, the foreground / background features have been well separated.

[0178] In order to enable those skilled in the art to accurately implement and reproduce this embodiment 2, the following is a description of α, β,

[0179] The three new parameters of γ provide feasible heuristic tuning suggestions:

[0180] α is an adaptive gain coefficient. The larger α is, the more it tends to strengthen the foreground feature F. g , and the smaller it is, the more it will relatively focus on the background feature;

[0181] The initial value can be set to 0.5. After training every several rounds, the changes in the Loss curve, mAP, or Precision can be observed. If the miss detection rate increases in a high-noise environment, α needs to be increased;

[0182] If there is little background interference but the false detection rate is found to increase, α can be appropriately reduced to allow the network to retain more background information to avoid the network being overly "sensitive".

[0183] β is a cross-branch residual alignment coefficient, usually between (0, 1); when β is large, the difference between (αF b - F g ) is significantly amplified, which is easy to strengthen the compensation when small targets exist, but being too large may also cause the foreground feature to be impacted;

[0184] The initial value can be in the range of 0.3 - 0.8; if it is detected that the vehicle is largely blocked or the background is extremely complex, β can be appropriately increased to enhance the residual alignment; if the foreground feature is overly perturbed due to too high β, it can be adjusted downwards.

[0185] γ is a cascaded fusion factor, usually between (0, 1); when γ is relatively large, the system relies more on the role of adaptive gain; when γ is relatively small, it relies more on cross-branch residual alignment;

[0186] The initial value can be set to 0.5 to make the outputs of the two algorithms each account for half; in actual verification, if overly relying on AGF, the in-depth mining of residual signals may be ignored; on the contrary, if overly relying on CBRA, unnecessary residual operations may be increased when the background is stable; similar to random search or the dichotomy method can be adopted to perform offline parameter tuning for γ and optimize it on the validation set; it can also be during the online training process, and γ can participate in automatic optimization through backpropagation with a certain learning rate.

[0187] The following are the precautions for this Embodiment 2:

[0188] Add a new script or sub-function at the backend of the FDRM module to implement the Adaptive Feature Gain Distribution Algorithm (AFGD) and the Cross-Branch Residual Alignment Algorithm (CBRA);

[0189] Ensure that the three parameters α, β, γ can be jointly managed by the optimizer of the network (if the learnable method is selected) or an external manager (if the search or custom update method is selected).

[0190] When implementing the cascaded linkage fusion algorithm, it is necessary to add references to the outputs of the adaptive feature gain allocation algorithm and the cross-branch residual alignment algorithm in the code, and complete the secondary weighting and merging according to formula (3).

[0191] Formula details:

[0192] If ReLU is selected as σ(·), it should be noted that foreground / background features will be truncated if they are negative, and this problem can be alleviated by using LeakyReLU or other activation functions;

[0193] For scenarios with high requirements for floating-point precision, if α, β, γ participate in backpropagation, it is necessary to ensure that the training framework can correctly handle these parameters during gradient update and avoid numerical overflow or gradient explosion.

[0194] In the initial stage, α, β, γ can be locked to only allow offline tuning, that is, fix the values first and then train, and then fine-tune after observing convergence;

[0195] In the later stage, if automation is desired, α, β, γ can be declared as learnable variables in the network and written into the loss function for backpropagation update, but care should be taken in the initial value selection.

[0196] When there are problems such as unstable accuracy or overfitting, conventional measures such as reducing the learning rate, reducing the number of network layers, or adjusting some hyperparameters can be tried.

[0197] It is recommended to visualize the key intermediate outputs during implementation to view the distribution of foreground / background fusion and residual signals, which is convenient for quickly locating problems;

[0198] If there are serious misdetections or missed detections on some test sets, the value ranges of β and γ can be focused on to check whether they cause excessive amplification of residuals or insufficient fusion.

[0199] In summary, the technical effects of this Embodiment 2 are as follows:

[0200] When facing extreme situations such as strong noise, non-stationary clutter, local shadows / occlusions, etc., this Embodiment 2 is significantly superior to the traditional scheme in terms of vehicle recognition accuracy and missed detection rate through the combination of Algorithms 1 to 3.

[0201] The traditional FDRM module focuses on reducing feature redundancy but does not explicitly handle the feature interaction between foreground / background during transient drastic changes; this Embodiment 2 ensures that foreground features are not significantly diluted and can impose "fine-grained suppression" on background noise by means of the cross-branch residual alignment mechanism.

[0202] The three major parameters α, β, γ can be regarded as paddles, and the network can be adjusted automatically or semi-automatically at different training stages and different scenario requirements to meet diverse application needs;

[0203] Different fields or application scenarios (such as traffic flow monitoring, battlefield reconnaissance, natural disaster assessment, etc.) can customize more suitable numerical ranges according to the actual noise and target size distribution.

[0204] Explicitly characterize the progressive process of foreground / background branch fusion through formulas (1) - (3), providing a basis for subsequent researchers or engineering implementers to further expand and improve the cumulative learning strategy;

[0205] Further enrich the three major fusion points of "feature difference, residual injection, adjustable gain" in the cutting-edge field of deep learning + SAR target detection, which is a good supplement to existing backbone frameworks such as FPN and Transformer.

[0206] In Example 1, the cumulative learning strategy is used to linearly or simply weight-fuse foreground features and background features, and good vehicle target detection results have been achieved. However, as the background noise and occlusion conditions become increasingly complex, fixed or single weighting methods will cause small target features to be engulfed and background redundancy to be difficult to fully suppress, thus affecting the accuracy and recall rate of vehicle detection. To further improve the performance of Example 1, this Example 2 improves it in its "cumulative learning and feature weighting" stage, introduces three sets of distributed algorithms and supporting formulas, and solves problems such as insufficient background noise suppression and missed detection of small targets around the newly added adjustable parameters, realizing more flexible and higher-precision vehicle target recognition.

[0207] Practice shows that Example 1 can achieve high-precision and high-recall vehicle detection results in most conventional or relatively noise-moderate SAR environments. However, when facing extreme scenarios such as ultra-high noise, strong interference, complex occlusion, and drastic changes in background information, there are still the following potential problems:

[0208] Usually, only linear or simple functions are used to fuse the "foreground branch" and "background branch" during the training process. Once the background energy is too strong or the small target features are relatively weak, it may not be possible to fully "amplify" the key vehicle features;

[0209] Although the FDRM module reduces feature redundancy through orthogonal loss and compact loss, the cross-branch complementarity is not explicitly considered in the specific fusion link;

[0210] When the vehicle target is partially occluded or the surrounding noise is extremely high, simple weight increase or decrease is often insufficient, and problems such as insufficient suppression or over-suppression are likely to occur.

[0211] To improve the robustness and accuracy in extreme scenarios (such as partial vehicle occlusion by buildings, large-area rain and snow noise, non-uniform background, etc.), this Example 2 proposes three key improvements:

[0212] Introduce more adjustable parameters to allow for more flexible adjustment of the foreground / background fusion degree according to the performance on the validation set during the training process;

[0213] In the cumulative learning stage, a cross-branch residual idea is proposed to effectively retain small target information and reduce background false detections through residual alignment;

[0214] Through a cascaded linkage fusion method, multiple feature enhancement strategies are combined so that vehicle targets can still be captured and highlighted under strong noise.

[0215] The improvements in this Embodiment 2 are carried out around the above technical problems respectively.

Claims

1. A vehicle target recognition method based on feature decomposition and cumulative learning, characterized in that Including: Using the vehicle target dataset of land SAR images to construct the dataset; Extracting and enhancing the scattering characteristics of vehicle targets; Constructing the ST-PA_RCNN pixel-level fusion model; Training the feature-level fusion model based on the FDRM module; Using the test data to test the performance of the model and output the results.

2. The vehicle target recognition method based on feature decomposition and cumulative learning according to claim 1, wherein The using of the vehicle target dataset of land SAR images to construct the dataset includes: According to the characteristics of satellite left and right side-looking and ascending and descending orbits, combined with the layout direction of the target in the distribution base, selecting the key-scene areas to collect sample data; After standardization processing, using the target image distribution to check whether the sample data meets the requirements of the sample distribution index; If it meets the requirements, the dataset construction is completed; If it does not meet the requirements, analyze the missing sample observation parameters, use orbit simulation to analyze the distribution position coverage of the missing sample observation parameters, and form an optimal detection plan that can be completed in a short time; Use the SAR satellite platform with strong timeliness such as high altitude to perform rapid detection and preprocessing to form training and test samples.

3. The vehicle target recognition method based on feature decomposition and cumulative learning according to claim 1, wherein The extracting and enhancing of the scattering characteristics of vehicle targets includes: Extracting the texture features of SAR images through four algorithms of SAR-SIFT, SAR-HOG, NSLP and LC; Using the SAR-Harris detector and OPTICS clustering method to enhance the extracted features.

4. The vehicle target recognition method based on feature decomposition and cumulative learning according to claim 1, characterized in that The constructing of the ST-PA_RCNN pixel-level fusion model includes: Constructing the ST-PA_RCNN backbone network, combining SwinTransformer and PA_FPN, and fusing the original image and the feature map after scattering feature enhancement through the channel fusion method, and training the object detection model.

5. The vehicle target recognition method based on feature decomposition and cumulative learning according to claim 1, characterized in that The training of the feature-level fusion model based on the FDRM module includes: Based on ST-PA_RCNN, introducing the FDRM module based on the feature fusion method Dual-Branch FPN, and reducing feature redundancy and enhancing the difference in the feature space through feature decomposition and weight reallocation methods.

6. The vehicle target recognition method based on feature decomposition and cumulative learning according to claim 1, characterized in that, The using of the test data to test the performance of the model and output the results includes: Inputting the images in the test dataset into the trained FDRM ST-PA_RCNN model, the FDRM ST-PA_RCNN model includes pixel-level and feature-level fusion models, performing object detection inference, and calculating the performance indicators of the model; Evaluating the performance of the model by calculating the performance indicators of precision, recall and mean average precision; Outputting the detection results, the detection results include the bounding box and the target category, and visualizing the detection results.

7. The vehicle target recognition method based on feature decomposition and cumulative learning according to claim 1, wherein, Also including: Performing channel fusion on the original SAR image and the feature map after scattering feature enhancement to form multi-channel input data.

8. The vehicle target recognition method based on feature decomposition and cumulative learning according to claim 1, characterized in that, Also including: Reducing feature redundancy through feature decomposition and weight reallocation methods, Enhancing the difference in the feature space.

Citation Information

Patent Citations

  • Multi-source characteristic integrated SAR image automatic object identification method

    CN107239740A

  • Polarimetric SAR (synthetic aperture radar) image target detection method based on multipolarization features and FCN (full convolutional)-CRF (conditional random field) fusion network

    CN107392122A

  • Target detection method based on multi-source heterogeneous data cognitive fusion

    CN112465880A

  • Vehicle-mounted video target detection method based on deep learning

    WO2020181685A1