SAR (Synthetic Aperture Radar) target identification method, device and equipment based on hybrid expert system
Through the hybrid expert system method, combined with physical mechanism and data-driven hybrid autonomous physical expert network, the problem of insufficient generalization ability of SAR target recognition under data-driven is solved, and efficient target recognition in the absence of data is achieved.
Patent Information
- Application Number
- CN202510631011.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-08
AI Technical Summary
The existing SAR target recognition methods lack generalization capabilities under data-driven, especially when there are large differences between small samples and within-classes and uneven sample distribution, it is difficult to achieve robust cross-scene recognition.
Using a hybrid expert system method, a hybrid autonomous physical expert network that integrates physical mechanisms and data-driven hybrid autonomous physical expert network is used to map the image of the attribute scattering center component to the same dimension space of the amplitude feature map, generate sparse weights and activate the expert unit, combine the multi-head mutual attention layer and multi-scale convolution unit for feature extraction, and finally use the detection head to achieve target recognition.
It improves the generalization recognition ability in the absence of data, enhances the accuracy and robustness of SAR target recognition, especially in the condition of a small amount of data, which can achieve efficient target recognition.
Smart Images

Figure CN120451795A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of SAR automatic target recognition, and in particular to a SAR target recognition method, apparatus and device based on a hybrid expert system. Background Art
[0002] Synthetic Aperture Radar (SAR), an active microwave imaging sensor, has become a core technology for Earth observation due to its all-weather, all-day, and high-penetration capabilities. Since its introduction in the 1950s, SAR technology has achieved breakthroughs in resolution improvement and multimodal imaging, and its applications have expanded to include terrain mapping, disaster monitoring, and other scenarios. Automatic Target Recognition (ATR), a core component of SAR image interpretation, aims to implement a three-stage processing flow of target detection, identification, and classification through algorithms. Traditional methods rely on manual feature design and template matching, but lack robustness under extended operating conditions such as target pose and observation angle variations.
[0003] With the rise of deep learning technology, methods based on convolutional neural networks (CNNs) have demonstrated significant advantages in optical remote sensing. However, SAR images exhibit unique microwave scattering characteristics: the intensity of target echoes is closely related to the geometric structure, material properties, and imaging parameters (such as pitch angle and polarization), causing the target image to vary dramatically with observation conditions. This characteristic presents challenges for SAR datasets, such as small sample sizes, large intra-class variability, and uneven sample distribution. Deep learning methods that rely solely on data-driven methods are prone to overfitting and struggle to achieve robust cross-scenario generalization. Summary of the Invention
[0004] Based on this, it is necessary to provide a SAR target recognition method, device and equipment based on a hybrid expert system that has strong generalization recognition capabilities in the absence of data to address the above technical problems.
[0005] A SAR target recognition method based on a hybrid expert system, the method comprising:
[0006] Acquire a SAR image for target recognition, extract an amplitude feature map of the SAR image through a pre-order feature extraction layer, and use a clustering method to distinguish the SAR image according to physical attributes to obtain multiple attribute scattering center component images;
[0007] The amplitude feature map and multiple attribute scattering center component images are input into a hybrid autonomous physical expert network to obtain preliminary fusion features. In the hybrid autonomous physical expert network, a projection layer is used to map each attribute scattering center component image into a space of the same dimension as the amplitude feature map to obtain an ASC component feature. The amplitude feature map is grouped according to components and input into a multi-head mutual attention layer with the corresponding ASC component features to obtain a component-amplitude mutual attention feature. A weight mapping unit is used to generate sparse weights according to the component-amplitude mutual attention feature, and the sparse weights are used to activate or shield an expert unit. After adding the component-amplitude mutual attention feature to the amplitude feature map, the activated expert unit is used to output the preliminary fusion feature based on the features obtained after the addition.
[0008] The preliminary fusion features are combined with the amplitude features after passing through two multi-scale convolution units. Figure 1 And input it into the hybrid autonomous expert network to obtain fusion features;
[0009] A detection head is used to obtain a target recognition result based on the fusion features.
[0010] In one embodiment, in the hybrid autonomous physical expert network, the number of projection layers, multi-head mutual attention layers, and expert units is set based on the number of target components differentiated by a clustering method according to physical properties, wherein the number of expert units is an integer multiple of the number of target components;
[0011] For each target component, a corresponding projection layer, a multi-head mutual attention layer, and two or more expert units are set, where the number of attention heads in the multi-head mutual attention layer is the same as the number of corresponding expert units.
[0012] In one embodiment, the projection layer includes an offset convolution layer, a deformable convolution layer, and a normalization and activation layer;
[0013] Using the attribute scattering center component image as input data of the projection layer;
[0014] After adjusting the convolution sampling position of the input data using the offset convolution layer, the input data is subjected to deformable convolution with the input data via the deformable convolution layer to adaptively extract important spatial information features;
[0015] The normalization and activation layer is used to process the important spatial information features to obtain the output data of the projection layer, namely the ASC component features.
[0016] In one embodiment, the weight mapping unit includes a concatenation layer, a global average pooling layer, a linear layer, and a Softmax function layer;
[0017] The component-amplitude mutual attention features output by each multi-head mutual attention layer are spliced in the channel dimension by using the splicing layer, and the original weights are obtained by passing the global average pooling layer, the linear layer and the softmax function layer;
[0018] After adding learnable noise to the original weights, a Softmax function layer is used to process the noised weights to obtain updated weights.
[0019] The first plurality of weights with the highest probability are selected from the updated weights as the sparse weights.
[0020] In one embodiment, when the activated expert units are used to obtain the preliminary features, the corresponding expert units are activated according to the sparse weights, and the output data of the activated expert units are weighted according to the sparse weights.
[0021] In one embodiment, the expert unit is constructed based on a stacked structure of multiple layers of variable-size convolution kernels and a channel separation strategy.
[0022] In one embodiment, the structure of the hybrid autonomous expert network is the same as the structure of the hybrid autonomous physical expert network;
[0023] In the hybrid autonomous expert network, the preliminary fusion features processed by the two multi-scale convolution units are used as input data, and the query, key, and value inputs of the mutual attention module are replaced by the grouped preliminary fusion features, thereby completing the self-attention operation of the preliminary fusion features.
[0024] In one embodiment, a target neural network is obtained, which is a convolutional type neural network. The hybrid autonomous physical expert network and the hybrid autonomous expert network are embedded in the target neural network. The convolution structure in the target neural network is used as the preceding feature extraction layer and the multi-scale convolution unit to extract the fusion features of the SAR image.
[0025] The present application also provides a SAR target recognition device based on a hybrid expert system, the device comprising:
[0026] An input data acquisition module is used to acquire a SAR image to be used for target recognition, extract an amplitude feature map of the SAR image through a pre-order feature extraction layer, and use a clustering method to distinguish the SAR image according to physical properties to obtain multiple attribute scattering center component images;
[0027] a hybrid autonomous physical expert network data processing module, configured to input the amplitude feature map and multiple attribute scattering center component images into the hybrid autonomous physical expert network to obtain preliminary fusion features, wherein in the hybrid autonomous physical expert network, a projection layer is used to map each attribute scattering center component image into a space of the same dimension as the amplitude feature map to obtain an ASC component feature, the amplitude feature map is grouped according to components and input into a multi-head mutual attention layer along with the corresponding ASC component features to obtain a component-amplitude mutual attention feature, a weight mapping unit is used to generate sparse weights based on the component-amplitude mutual attention feature, and the sparse weights are used to activate or shield an expert unit, the component-amplitude mutual attention feature is added to the amplitude feature map, and the activated expert unit is used to output the preliminary fusion feature based on the features obtained after the addition;
[0028] Hybrid autonomous expert network data processing module, used to combine the initial fusion features with the amplitude features after passing through two multi-scale convolution units Figure 1 And input it into the hybrid autonomous expert network to obtain fusion features;
[0029] The target recognition result obtaining module is used to obtain the target recognition result according to the fusion feature using the detection head.
[0030] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0031] Acquire a SAR image for target recognition, extract an amplitude feature map of the SAR image through a pre-order feature extraction layer, and use a clustering method to distinguish the SAR image according to physical attributes to obtain multiple attribute scattering center component images;
[0032] The amplitude feature map and multiple attribute scattering center component images are input into a hybrid autonomous physical expert network to obtain preliminary fusion features. In the hybrid autonomous physical expert network, a projection layer is used to map each attribute scattering center component image into a space of the same dimension as the amplitude feature map to obtain an ASC component feature. The amplitude feature map is grouped according to components and input into a multi-head mutual attention layer with the corresponding ASC component features to obtain a component-amplitude mutual attention feature. A weight mapping unit is used to generate sparse weights according to the component-amplitude mutual attention feature, and the sparse weights are used to activate or shield an expert unit. After adding the component-amplitude mutual attention feature to the amplitude feature map, the activated expert unit is used to output the preliminary fusion feature based on the features obtained after the addition.
[0033] The preliminary fusion features are combined with the amplitude features after passing through two multi-scale convolution units. Figure 1And input it into the hybrid autonomous expert network to obtain fusion features;
[0034] A detection head is used to obtain a target recognition result based on the fusion features.
[0035] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0036] Acquire a SAR image for target recognition, extract an amplitude feature map of the SAR image through a pre-order feature extraction layer, and use a clustering method to distinguish the SAR image according to physical attributes to obtain multiple attribute scattering center component images;
[0037] The amplitude feature map and multiple attribute scattering center component images are input into a hybrid autonomous physical expert network to obtain preliminary fusion features. In the hybrid autonomous physical expert network, a projection layer is used to map each attribute scattering center component image into a space of the same dimension as the amplitude feature map to obtain an ASC component feature. The amplitude feature map is grouped according to components and input into a multi-head mutual attention layer with the corresponding ASC component features to obtain a component-amplitude mutual attention feature. A weight mapping unit is used to generate sparse weights according to the component-amplitude mutual attention feature, and the sparse weights are used to activate or shield an expert unit. After adding the component-amplitude mutual attention feature to the amplitude feature map, the activated expert unit is used to output the preliminary fusion feature based on the features obtained after the addition.
[0038] The preliminary fusion features are combined with the amplitude features after passing through two multi-scale convolution units. Figure 1 And input it into the hybrid autonomous expert network to obtain fusion features;
[0039] A detection head is used to obtain a target recognition result based on the fusion features.
[0040] The above-mentioned SAR target recognition method, device and equipment based on hybrid expert system extracts features from SAR images by utilizing a hybrid autonomous physical expert network that integrates physical mechanisms and data-driven methods. The hybrid autonomous physical expert network uses a projection layer to map each attribute scattering center component image to a space of the same dimension as the amplitude feature map to obtain ASC component features. The amplitude feature map is grouped according to components and input into a multi-head mutual attention layer with the corresponding ASC component features to obtain component-amplitude mutual attention features. A weight mapping unit is used to generate sparse weights according to the component-amplitude mutual attention features, and the sparse weights are used to activate or shield the expert unit. After the component-amplitude mutual attention features are added to the amplitude feature map, the activated expert unit is used to output preliminary fusion features based on the features obtained after the addition. The preliminary fusion features are further extracted using the hybrid autonomous expert network, and then target recognition is achieved to improve the generalization recognition capability in the absence of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 1 is a flow chart of a SAR target recognition method based on a hybrid expert system in one embodiment;
[0042] Figure 2 A schematic diagram of the structure of a hybrid autonomous physical expert network in one embodiment;
[0043] Figure 3 is a schematic structural diagram of a projection layer in one embodiment;
[0044] Figure 4 Schematic diagram of the key differences between a normal convolution operation and a deformable convolution operation in one embodiment;
[0045] Figure 5 A schematic diagram of the structure of an expert backbone network in one embodiment;
[0046] Figure 6 Schematic diagram of the structure of a first target recognition model in one embodiment;
[0047] Figure 7 is a schematic structural diagram of a second target recognition model in one embodiment;
[0048] Figure 8 A structural block diagram of a SAR target recognition device based on a hybrid expert system in one embodiment;
[0049] Figure 9 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0051] The existing SAR-ATR technology is facing a core contradiction between data-driven and physical guidance, that is, the data-driven method relies on large-scale labeled samples but lacks physical interpretability, while the physical model has mechanism transparency but is limited in generalization ability for complex scenes. In this application, a SAR target recognition method based on a hybrid expert system is provided, such as Figure 1 As shown, the method specifically includes the following steps:
[0052] Step S100 , obtaining a SAR image to be used for target recognition, extracting an amplitude feature map of the SAR image through a pre-order feature extraction layer, and using a clustering method to differentiate the SAR image according to physical attributes to obtain multiple attribute scattering center component images.
[0053] Step S110, input the amplitude feature map and multiple attribute scattering center component images into the hybrid autonomous physical expert network to obtain preliminary fusion features, in the hybrid autonomous physical expert network, use the projection layer to map each attribute scattering center component image to the space of the same dimension as the amplitude feature map to obtain ASC component features, group the amplitude feature map according to components and input it into the multi-head mutual attention layer with the corresponding ASC component features to obtain component-amplitude mutual attention features, use the weight mapping unit to generate sparse weights according to the component-amplitude mutual attention features, and use the sparse weights to activate or shield the expert unit, add the component-amplitude mutual attention features and the amplitude feature map, and use the activated expert unit to output the preliminary fusion features according to the features obtained after the addition.
[0054] Step S120: After the initial fusion features pass through two multi-scale convolution units, they are combined with the amplitude features. Figure 1 And input it into the hybrid autonomous expert network to obtain the fusion features.
[0055] Step S130: Using the detection head to obtain the target recognition result based on the fusion features.
[0056] In this application, a hybrid autonomous physical expert network that integrates data-driven and physical property-driven methods is proposed for feature extraction. By giving the expert system physical perception capabilities, it can independently decide whether to process input features, thereby improving the neural network's ability to dynamically allocate effective computing resources, thereby enhancing the accuracy and generalization ability of SAR target recognition with a small amount of data.
[0057] In step S100, the targets in the SAR image can be high-value entities such as vehicles, aircraft, ships, or buildings to be identified. These entities play a specific role in social production and life, and locating and identifying their individual types facilitates the efficient completion of tasks such as resource statistics and regional monitoring.
[0058] In the following, we first introduce the hybrid autonomous physical expert network structure in step S110, as shown in Figure 2 shown.
[0059] In this embodiment, the hybrid autonomous physical expert network receives two types of inputs: the amplitude feature map passed in by the previous feature extraction layer and the attribute scattering center (ASC) component image distinguished by physical properties. Figure 2 As can be seen from the figure, the attribute scattering center component images include multiple images, which are the attribute scattering center images of each component in the target. The hybrid autonomous physical expert network uses the corresponding projection layer to process each attribute scattering center component image separately.
[0060] In this embodiment, the structure of the projection layer is as follows Figure 3 As shown in the figure, it includes an offset convolution layer, a deformable convolution layer, and a normalization and activation layer. The attribute scattering center component image is used as the input data of the projection layer. The offset convolution layer is used to adjust the convolution sampling position of the input data, and then deformable convolution is performed with the input data through the deformable convolution layer to adaptively extract important spatial information features. The normalization and activation layer is then used to process the important spatial information features to obtain the output data of the projection layer, namely the ASC component features.
[0061] Specifically, although the two-dimensional convolution operation is effective in introducing visual priors to neural networks and reducing the number of parameters, its fixed geometric structure of the sampling grid is difficult to adapt to the specific target shape and posture changes, and there is still room for improvement in dynamic data sampling. Deformable convolution is an improved method that enhances the model's geometric deformation modeling capabilities by dynamically adjusting the sampling position. Its core idea is to give each sampling point of the standard convolution a learnable offset p n , so that the convolution kernel can adaptively adjust the receptive field according to the input features. Figure 4 Demonstrates the key differences between ordinary convolution operations and deformable convolution operations.
[0062] To alleviate the dimensionality and feature space mismatch problem caused by direct input of ASC into the network, in this embodiment, deformable convolution is used to group and project the components of ASC. Given an input image x, the deformable convolution result of any point x(p0) on it is:
[0063]
[0064] In formula (1), p n Represents the offset of each point in the convolution kernel relative to the center. Taking 3*3 convolution as an example, p n Possible values are:
[0065] p n ∈{(-1,-1),(-1,0),(-1,1),...,(0,0),...,(1,1)}(2)
[0066] As shown in formula (2), p n There are 9 possible coordinate values. Therefore, by using offset convolution to learn the p of each position n , which enables the deformable convolution to sample important spatial information in the sparse point-like ASC components, enhances the nonlinear representation ability of the projection layer, and increases the effective feature space dimension.
[0067] In this embodiment, in the hybrid autonomous physical expert network, the number of projection layers, multi-head mutual attention layers, and expert units is set based on the number of target components that are distinguished by clustering methods based on physical properties, where the number of expert units is an integer multiple of the number of target components. Figure 2 In the hybrid autonomous physical expert network structure shown in Figure 2, four input attribute scattering center component images, i.e., four target components, are used as an example for illustration. It should be noted that this method can adjust the number of target components based on the specific application scenario and requirements; this number is not fixed.
[0068] Specifically, a corresponding projection layer, a multi-head mutual attention layer, and two or more expert units are set for each target component, wherein the number of attention heads in the multi-head mutual attention layer is the same as the number of corresponding expert units.
[0069] Furthermore, the following description is made by taking the example of 4 attribute scattering center component images. C (4) Initialize N E Experts (N C A positive integer multiple of Figure 2 Taking 2 times as an example), each type of component has N=N E / N C An expert unit is responsible for handling it.
[0070] In the hybrid autonomous physical expert network, the ASC component features and amplitude features obtained after processing in the projection layer are input into the multi-head mutual attention layer, so that the component attributes in the amplitude feature map are paid attention to and emphasized, and the component-amplitude mutual attention features are obtained.
[0071] Specifically, the amplitude feature map is divided into multiple groups according to the number of components using the channel dimension, and then input into the corresponding multi-head mutual attention layer with the corresponding ASC component features. For a multi-head mutual attention layer, it is first initialized, and then the query, key, and value vectors are calculated based on the input data. For the j-th attention head of the i-th component, we have:
[0072]
[0073] In formula (3), i∈[0,N C -1],j∈[0,N-1]. The flatten function flattens the position-encoded image into a two-dimensional vector in channel-space. They are the Query, Key, and Value projection weights of the jth head of the i-th component, ASC-feat i represents the ASC component feature of the i-th component, X i It is important to note that the number of attention heads in each multi-head mutual attention layer is actually the same as the number of expert units processing the same component.
[0074] Furthermore, the projected ASC component features and the corresponding grouped amplitude features are subjected to two-dimensional position encoding (PE). Let the feature map with C channels, height h, and width w be P0(c, x, y). After 2DPE, the process is expressed as:
[0075]
[0076] In formula (4), and is an integer, and ω k is the frequency coefficient, defined as:
[0077] ω k =10000 -2k / C (5)
[0078] After the position encoding, projection dimensionality reduction and spatial flattening operations in formula (3), the feature vector ASC-feat of the input i-th group of N-head mutual attention is i and X i Reduce the dimension to the dimension of a single attention head and calculate the single-head attention score:
[0079]
[0080] In formula (6), d k Is the number of channels of the vector. Then for the i-th component, the channel dimension concatenation function Concat(·) is used to obtain its multi-head attention feature, which is expressed as:
[0081] x i =Feat i Comp-Mag =Concat(Attention i1 ,Attention i2 ,...,Attention iN ) (7)
[0082] In formula (7), x i That is, the component-amplitude mutual attention feature of a single attribute scattering center component (the output data of the multi-head mutual attention layer).
[0083] Furthermore, considering that each N-head mutual attention layer belongs to N component experts, in order to subsequently assign features to each expert unit, The attention features obtained by the j-th expert for the i-th component, function Split(x,N) j It means to divide x into N parts in the channel dimension and output the jth part.
[0084] After obtaining the features that focus on the relationship between different components and the amplitude image, it is natural to use them as the basis for each expert unit to judge the level of its backbone network's ability to process this sample. In this embodiment, a weight mapping unit is designed based on a sparsely gated mixture of experts and noisy Top-K gating. This unit generates sparse weights based on the component-amplitude mutual attention features output by each multi-head mutual attention layer to activate or block the patent unit.
[0085] In this embodiment, the weight mapping unit includes a splicing layer, a global average pooling layer, a linear layer and a Softmax function layer. The splicing layer is first used to splice the component-amplitude mutual attention features output by each multi-head mutual attention layer in the channel dimension, and then the original weights are obtained through the global average pooling layer, the linear layer and the Softmax function layer. After adding learnable noise to the original weights, the Softmax function layer is used to process the weights after noise to obtain updated weights. The first multiple weights with the largest probability are selected from the updated weights as sparse weights.
[0086] Specifically, the original weights obtained include multiple weights corresponding to each attention head. Taking an N-head mutual attention layer, where N is 2 as an example, the number of original weights is 8. The process of obtaining the original weights is expressed as:
[0087]
[0088] Furthermore, the process of adding learnable Gaussian noise and applying gating operation to obtain sparse weight G(x) is expressed as:
[0089]
[0090] Furthermore, eventually, the number of selected sparse weights is consistent with the number of components.
[0091] Next, let's explain the expert unit. First, it's important to note that the expert unit is the core of the expert module, while the head of the expert module is actually a multi-head mutual attention layer shared by experts belonging to the same component. In other words, the expert module consists of a multi-head mutual attention layer corresponding to a certain type of component, as well as the expert unit.
[0092] Specifically, in the multi-head mutual attention layer, it performs mutual attention calculation by combining the amplitude and single-category attribute scattering center features. On the one hand, after integrating the multi-component features, it shields or empowers the expert backbone part through the weight mapping module. On the other hand, it improves the corresponding experts' understanding of the ASC components of independent categories, driving experts to achieve differentiation and specialization from the feature allocation level.
[0093] After the expert header, in this embodiment, a specific structure of an expert unit is also proposed, such as Figure 5 As shown in the figure, it can extract robust features in the contour / semantic information space in complex scenes by adopting the stacking and channel separation strategy of multiple layers of variable-size convolution kernels.
[0094] In this embodiment, when obtaining preliminary features using the activated expert units, the corresponding expert units are activated according to the sparse weights, and the output data of the activated expert units are weighted according to the sparse weights.
[0095] Specifically, all amplitude feature maps and the component-amplitude mutual attention features averaged in the channel dimension are added together and input into the activated expert to obtain the input of the expert backbone structure. For this process, the output of the nth expert unit, that is, the jth expert under the i-th component, is The output result y of the hybrid autonomous physical expert network is expressed as:
[0096]
[0097] In formula (10), n is the expert index, n = N*i+j. In the specific implementation, after the top K experts are obtained through the Top-K method, the features of this sample no longer pass through the remaining routing experts, achieving the essential goal of improving computational efficiency and the attention of the attribute scattering center component.
[0098] Next, after being processed by the hybrid autonomous physical expert network in step S110, preliminary fusion features are obtained. Then, in step S120, two multi-scale convolutions are used to extract semantic features from the preliminary fusion features in turn, and then the features are input into the hybrid autonomous expert network and autocorrelated with the amplitude feature map to obtain the final fusion features. Finally, in step S130, the detection head is used to achieve high-level target recognition based on the fusion features.
[0099] In this embodiment, the structure of the hybrid autonomous expert network is identical to that of the hybrid autonomous physical expert network. In the hybrid autonomous expert network, the initial fused features processed by two multi-scale convolutional units are used as input data. The query, key, and value inputs of the mutual attention module are replaced with the grouped initial fused features, completing the self-attention operation on the initial fused features. Specifically, the hybrid autonomous expert network uses a multi-head self-attention module guided by non-attribute scattering centers as the expert head, enabling the expert backbone to focus on the intra- and inter-channel correlations of deep semantic information.
[0100] In this embodiment, the aforementioned hybrid autonomous expert network and hybrid autonomous physical expert network are directly embedded into a target convolutional neural network. The convolutional structure of the target neural network serves as a pre-processing feature extraction layer and a multi-scale convolutional unit to extract fused features from SAR images. This allows for the construction of a target recognition model based on the hybrid autonomous expert network and the hybrid autonomous physical expert network, increasing the flexibility of this method across various application scenarios.
[0101] In this application, two target recognition models based on the above-mentioned hybrid autonomous expert network and hybrid autonomous physical expert network are proposed.
[0102] In one embodiment, the hybrid autonomous expert network and the hybrid autonomous physical expert network are embedded into the MSNet, and the structure after embedding is as follows: Figure 6 As shown. The aforementioned hybrid autonomous physical expert network is embedded in a multi-scale convolutional network with parallel information flow to obtain the first target recognition model. The SAR amplitude image is fused by stacking two layers of 3 / 7 / 11 scale convolution kernels, and the feature map is spliced according to the channel dimension and randomly shuffled, extracting rich and robust target features for subsequent expert network decision-making. The hybrid autonomous physical expert uses the mutual attention mechanism of ASC image and amplitude features for weight mapping in the expert head. After cascading two multi-scale convolution modules, the hybrid autonomous expert system uses a multi-head self-attention module guided by non-attribute scattering centers as the expert head, allowing the expert backbone to focus on the intra-channel / inter-channel correlation of deep semantic information. Finally, the feature map is flattened after adjusting the dimension through convolution and input into the fully connected layer + Softmax to obtain the recognition result.
[0103] In another embodiment, a second object recognition model is constructed based on the VGG-19 network, such as Figure 7 As shown in the figure, within the densely stacked convolutional blocks (Conv2d+ReLU) of VGGNet, feature maps at the first and fourth scales are selected and inserted into the hybrid autonomous physical expert network and hybrid autonomous expert network. The two expert networks complete the task of integrating amplitude-ASC information and interacting with deep semantic information, and then output the object recognition results through a fully connected layer and softmax.
[0104] In this embodiment, during the training of the target recognition model constructed by the hybrid autonomous expert network and the hybrid autonomous physical expert network, the experts who initially obtained high gating weights took on more gradient flow and first improved the feature representation ability after backpropagation. Therefore, the gating network tended to continue to assign higher weights. This imbalance is constantly self-reinforcing. To address the load imbalance problem, a soft constraint method is adopted. Let the set of routing expert indices selected in each training batch be I, and then define the importance and loss:
[0105]
[0106] In formula (11), CV(·) 2 Refers to the variance. Such a loss function encourages the gating expert to output a smaller variance, that is, a combination with similar weights between the loads.
[0107] Furthermore, let gt be the true label of the sample, X be the input sample image corresponding to the sample, and M(X) be the recognition result output by the model. The cross entropy loss for constructing the basic classification can be defined as:
[0108]
[0109] In formula (12), num-Batch is the number of samples in each training batch, i represents the sample index, c represents the category index, and M C represents the total number of categories, gt(i,c) is a sign function that takes 1 when c is the true category of sample i and takes 0 otherwise, M(X) ic is the probability that the model predicts that sample i belongs to category c,
[0110] Furthermore, by combining formula (11) and formula (12), we can obtain the total loss function for optimizing the model:
[0111] L Total =L CrossEntropy [gt,M(X)]+α·L Importance (I) (13)
[0112] In formula (13), α is a hyperparameter for adjusting the loss size, which is used to adjust the attention paid to the recognition loss and the balanced expert load loss.
[0113] In this paper, we also conduct experiments based on the first and second target recognition models mentioned above to demonstrate the effectiveness of our method. We conduct experiments on the Moving and Stationary Target Acquisition and Recognition (MSTAR) dataset (where the targets are vehicles), using a once-for-all strategy to evaluate the recognition performance.
[0114] The MSTAR dataset is collected using two data collection methods: Standard Operating Condition (SOC) and Extended Operating Condition (EOC), and there are certain differences between the images. Existing deep learning methods can already achieve extremely high accuracy (over 99%) using traditional SOC data for training and EOC data for testing. This does not provide further guidance for further improving the model's practicality and generalization performance. The training and test sets were reconstructed using a once-for-all (OFA) evaluation strategy. Once-for-all here means that once the model has been trained and validated using known data, it is evaluated on multiple test sets with different data distributions. This differs from traditional SOC and EOC evaluations, which require training multiple models to adapt to different test sets. OFA can more effectively evaluate the robustness and generalization ability of an algorithm.
[0115] In order to evaluate the performance of the algorithm in data-rich and data-limited situations respectively, 90%, 50%, 30% and 10% of the samples are randomly selected as training sets, and 10%, 50%, 50% and 50% of the samples in the remaining 17° overhead view data are used as validation sets accordingly.
[0116] For the two object recognition models constructed above, we set hyperparameters that match their structures and provide training settings for different amounts of training data, as shown in Table 1.
[0117] Table 1 Training hyperparameter settings of the two target recognition models proposed in this method
[0118]
[0119] In Table 1, branch-channel represents the output channel of each branch in the multi-scale convolutional module. The dimension increase of the output channels of the VGG block is performed by the first convolutional layer in each block, while the channel dimensions of the remaining convolutional layers remain unchanged. It is worth noting that to ensure that the model training reaches a basic convergence state under different training data amounts, slightly different training rounds were selected based on observations of the loss function and validation set accuracy.
[0120] As shown in Table 2, the proposed MS-PhysMoE (the first target recognition model) outperforms its baseline model in most test sets under all four training data sizes. In particular, at low data size (10%), the introduction of the hybrid autonomous physical expert system significantly improves the model's accuracy on the OFA-1 and OFA-2 validation sets (+10.3% and +7.7%), demonstrating the method's excellent small-sample learning and generalization capabilities.
[0121] Table 2 Comparison of recognition accuracy between MS-PhysMoE and baseline methods under OFA evaluation strategy*
[0122]
[0123] As shown in Table 3, the VGG-19 network (the second object recognition model) with PhysMoE also surpasses its baseline model at all training data levels. The substantial improvements in OFA-1 to OFA-3 metrics (+11.1%, +8.2%, and 17.3%) at 30% data usage, as well as the dramatic breakthrough in OFA-3 metrics (+12%) at 90% data usage, demonstrate the crucial role this approach plays in VGGNet.
[0124] Table 3 Comparison of recognition accuracy of VGG-PhysMoE and its baseline model under OFA evaluation strategy*
[0125]
[0126] In the above-mentioned SAR target recognition method based on hybrid expert system, a SAR image target model based on hybrid autonomous physical experts is proposed. The physical perception ability of the model is enhanced by introducing the features of ASC components. The hybrid expert system (MoE) method with intelligent distribution of information flow is used to balance computational efficiency and processing power, thereby improving the accuracy and generalization ability of SAR vehicle target recognition with a small amount of data.
[0127] At the same time, experiments using two different target models demonstrate that the hybrid autonomous physical expert network, a plug-and-play "physics-guided + data-driven" module, can dynamically assign experts and perform feature processing on the input amplitude feature map based on the target attribute scattering center image. This achieves a steady increase in vehicle target recognition accuracy as the training data gradually increases, while maintaining computational efficiency. For large overhead perspective differences and target variants, the improved model based on the proposed module achieves recognition accuracy that surpasses the baseline model under the same settings using 10% training data, demonstrating that the recognition method proposed in this patent has strong generalization capabilities in the absence of data.
[0128] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0129] In one embodiment, Figure 8 As shown, a SAR target recognition device based on a hybrid expert system is provided, comprising: an input data acquisition module 200, a hybrid autonomous physical expert network data processing module 210, a hybrid autonomous expert network data processing module 220, and a target recognition result acquisition module 230, wherein:
[0130] The input data acquisition module 200 is used to acquire a SAR image to be used for target recognition, extract an amplitude feature map of the SAR image through a pre-order feature extraction layer, and use a clustering method to distinguish the SAR image according to physical attributes to obtain multiple attribute scattering center component images;
[0131] A hybrid autonomous physical expert network data processing module 210 is configured to input the amplitude feature map and multiple attribute scattering center component images into a hybrid autonomous physical expert network to obtain preliminary fusion features, wherein in the hybrid autonomous physical expert network, each attribute scattering center component image is mapped into a space of the same dimension as the amplitude feature map using a projection layer to obtain an ASC component feature, the amplitude feature map is grouped according to components and input into a multi-head mutual attention layer along with the corresponding ASC component features to obtain a component-amplitude mutual attention feature, a weight mapping unit is used to generate sparse weights based on the component-amplitude mutual attention feature, and the sparse weights are used to activate or mask an expert unit, the component-amplitude mutual attention feature is added to the amplitude feature map, and the activated expert unit is used to output the preliminary fusion feature based on the features obtained after the addition;
[0132] Hybrid autonomous expert network data processing module 220 is used to combine the preliminary fusion features with the amplitude features after passing through two multi-scale convolution units. Figure 1 And input it into the hybrid autonomous expert network to obtain fusion features;
[0133] The target recognition result obtaining module 230 is used to obtain the target recognition result according to the fusion feature using the detection head.
[0134] The specific definitions of the hybrid expert system-based SAR target recognition device can be found in the definitions of the hybrid expert system-based SAR target recognition method described above and will not be further elaborated here. Each module in the hybrid expert system-based SAR target recognition device described above can be implemented in whole or in part via software, hardware, or a combination thereof. Each of these modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0135] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 9As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a SAR target recognition method based on a hybrid expert system is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0136] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0137] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0138] Acquire a SAR image for target recognition, extract an amplitude feature map of the SAR image through a pre-order feature extraction layer, and use a clustering method to distinguish the SAR image according to physical attributes to obtain multiple attribute scattering center component images;
[0139] The amplitude feature map and multiple attribute scattering center component images are input into a hybrid autonomous physical expert network to obtain preliminary fusion features. In the hybrid autonomous physical expert network, a projection layer is used to map each attribute scattering center component image into a space of the same dimension as the amplitude feature map to obtain an ASC component feature. The amplitude feature map is grouped according to components and input into a multi-head mutual attention layer with the corresponding ASC component features to obtain a component-amplitude mutual attention feature. A weight mapping unit is used to generate sparse weights according to the component-amplitude mutual attention feature, and the sparse weights are used to activate or shield an expert unit. After adding the component-amplitude mutual attention feature to the amplitude feature map, the activated expert unit is used to output the preliminary fusion feature based on the features obtained after the addition.
[0140] The preliminary fusion features are combined with the amplitude features after passing through two multi-scale convolution units. Figure 1 And input it into the hybrid autonomous expert network to obtain fusion features;
[0141] A detection head is used to obtain a target recognition result based on the fusion features.
[0142] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0143] Acquire a SAR image for target recognition, extract an amplitude feature map of the SAR image through a pre-order feature extraction layer, and use a clustering method to distinguish the SAR image according to physical attributes to obtain multiple attribute scattering center component images;
[0144] The amplitude feature map and multiple attribute scattering center component images are input into a hybrid autonomous physical expert network to obtain preliminary fusion features. In the hybrid autonomous physical expert network, a projection layer is used to map each attribute scattering center component image into a space of the same dimension as the amplitude feature map to obtain an ASC component feature. The amplitude feature map is grouped according to components and input into a multi-head mutual attention layer with the corresponding ASC component features to obtain a component-amplitude mutual attention feature. A weight mapping unit is used to generate sparse weights according to the component-amplitude mutual attention feature, and the sparse weights are used to activate or shield an expert unit. After adding the component-amplitude mutual attention feature to the amplitude feature map, the activated expert unit is used to output the preliminary fusion feature based on the features obtained after the addition.
[0145] The preliminary fusion features are combined with the amplitude features after passing through two multi-scale convolution units. Figure 1 And input it into the hybrid autonomous expert network to obtain fusion features;
[0146] A detection head is used to obtain a target recognition result based on the fusion features.
[0147] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0148] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0149] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A SAR target recognition method based on a hybrid expert system, characterized in that: The method comprises: Acquire a SAR image for target recognition, extract an amplitude feature map of the SAR image through a pre-order feature extraction layer, and differentiate the SAR image using a clustering method according to physical attributes to obtain multiple attribute scattering center component images; The amplitude feature map and multiple attribute scattering center component images are input into a hybrid autonomous physical expert network to obtain preliminary fusion features. In the hybrid autonomous physical expert network, a projection layer is used to map each attribute scattering center component image into a space of the same dimension as the amplitude feature map to obtain an ASC component feature. The amplitude feature map is grouped according to components and input into a multi-head mutual attention layer with the corresponding ASC component features to obtain a component-amplitude mutual attention feature. A weight mapping unit is used to generate sparse weights according to the component-amplitude mutual attention feature, and the sparse weights are used to activate or shield an expert unit. After adding the component-amplitude mutual attention feature to the amplitude feature map, the activated expert unit is used to output the preliminary fusion feature based on the features obtained after the addition. After the preliminary fusion features pass through two multi-scale convolution units, they are input into the hybrid autonomous expert network together with the amplitude feature map to obtain fusion features; A detection head is used to obtain a target recognition result based on the fusion features.
2. The SAR target recognition method according to claim 1, characterized in that: In the hybrid autonomous physical expert network, the number of projection layers, multi-head mutual attention layers, and expert units is set based on the number of target components differentiated according to physical data, wherein the number of expert units is an integer multiple of the number of target components; For each target component, a corresponding projection layer, a multi-head mutual attention layer, and two or more expert units are set, wherein the number of attention heads in each multi-head mutual attention layer is the same as the number of corresponding expert units.
3. The SAR target recognition method according to claim 2, characterized in that: The projection layer includes an offset convolution layer, a deformable convolution layer, and a normalization and activation layer; Using the attribute scattering center component image as input data of the projection layer; After adjusting the convolution sampling position of the input data using the offset convolution layer, the input data is subjected to deformable convolution with the input data via the deformable convolution layer to adaptively extract important spatial information features; The normalization and activation layer is used to process the important spatial information features to obtain the output data of the projection layer, namely the ASC component features.
4. The SAR target recognition method according to claim 2, characterized in that: The weight mapping unit includes a splicing layer, a global average pooling layer, a linear layer and a Softmax function layer; The component-amplitude mutual attention features output by each multi-head mutual attention layer are spliced in the channel dimension by using the splicing layer, and the original weights are obtained by passing the global average pooling layer, the linear layer and the softmax function layer; After adding learnable noise to the original weights, a Softmax function layer is used to process the noised weights to obtain updated weights. The first plurality of weights with the highest probability are selected from the updated weights as the sparse weights.
5. The SAR target recognition method according to claim 4, characterized in that: When the activated expert units are used to obtain the preliminary features, the corresponding expert units are activated according to the sparse weights, and the output data of the activated expert units are weighted according to the sparse weights.
6. The SAR target recognition method according to claim 5, characterized in that: The expert unit is constructed based on a stacked structure of multiple layers of variable-size convolution kernels and a channel separation strategy.
7. The SAR target recognition method according to claim 1, characterized in that: The structure of the hybrid autonomous expert network is the same as that of the hybrid autonomous physical expert network; In the hybrid autonomous expert network, the preliminary fusion features processed by the two multi-scale convolution units are used as input data, and the query, key, and value inputs of the mutual attention module are replaced by the grouped preliminary fusion features, thereby completing the self-attention operation of the preliminary fusion features.
8. The SAR target recognition method according to any one of claims 1 to 7, characterized in that: A target neural network is obtained, which is a convolutional neural network. The hybrid autonomous physical expert network and the hybrid autonomous expert network are embedded in the target neural network. The convolution structure in the target neural network is used as the preceding feature extraction layer and the multi-scale convolution unit to extract the fusion features of the SAR image.
9. A SAR target recognition device based on a hybrid expert system, characterized in that: The device comprises: An input data acquisition module is used to acquire a SAR image to be used for target recognition, extract an amplitude feature map of the SAR image through a preceding feature extraction layer, and use a clustering method to distinguish the SAR image according to physical attributes to obtain multiple attribute scattering center component images; a hybrid autonomous physical expert network data processing module is used to input the amplitude feature map and multiple attribute scattering center component images into a hybrid autonomous physical expert network to obtain preliminary fusion features, in the hybrid autonomous physical expert network, use a projection layer to map each attribute scattering center component image to a space of the same dimension as the amplitude feature map to obtain an ASC component feature, group the amplitude feature map according to components and input it into a multi-head mutual attention layer with the corresponding ASC component feature to obtain a component-amplitude mutual attention feature, use a weight mapping unit to generate sparse weights according to the component-amplitude mutual attention feature, and use the sparse weights to activate or shield expert units, add the component-amplitude mutual attention feature to the amplitude feature map, and use the activated expert unit to output the preliminary fusion feature based on the features obtained after the addition; A hybrid autonomous expert network data processing module is used to input the preliminary fusion features into the hybrid autonomous expert network together with the amplitude feature map after passing through two multi-scale convolution units to obtain fusion features; The target recognition result obtaining module is used to obtain the target recognition result according to the fusion feature using the detection head.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.