SAR Target Recognition Method and Device Based on Graph Structure Consistency Alignment

By using a method based on graph structure consistency alignment in SAR target recognition, multi-scale feature maps and graph structures are constructed, and the dependence on manual design parameters and consistency of recognition results in the prior art is solved, achieving higher recognition accuracy and robustness.

CN120014374BActive Publication Date: 2025-06-17NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510487867.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-17
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing SAR aircraft recognition technology relies on manual design parameters and rules, limiting its application in actual scenarios, and deep neural networks are difficult to capture the aircraft's geometric and topological consistency when processing SAR images.

Method used

Using a method based on graph structure consistency alignment, the SAR target recognition network is trained by constructing multi-scale feature maps and graph structures, using vertex correlation, edge similarity difference and graph distance loss functions to ensure the geometric and topological consistency of the recognition results.

Benefits of technology

It improves the accuracy and robustness of SAR target recognition, can more effectively capture the aircraft's geometry and topological structure, and reduces the dependence on manual design parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014374B_ABST
    Figure CN120014374B_ABST
Patent Text Reader

Abstract

The present application relates to a SAR target recognition method and device based on graph structure consistency alignment. By constructing a structure model graph for each SAR sample image in the sample dataset according to the target reference graph, training samples are generated and a training set is constructed. The training set is input into the SAR target recognition network. The sample image and the model graph are segmented into blocks and recombined into multi-scale feature maps according to a preset size. Then, target classification is performed based on the multi-scale sample feature maps, and the classification loss is calculated. Next, a graph structure is constructed to calculate the vertex correlation, edge similarity difference, and graph spectrum distance loss. The network is trained with these losses until convergence, and the trained network is used to recognize the SAR image to be processed. Using this method can effectively improve the target recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computer target recognition, and particularly to a SAR target recognition method and device based on graph structure consistency alignment. Background Art

[0002] As an active sensor, a synthetic aperture radar (SAR) system can measure the backscattering characteristics of a target by transmitting radar waves. Due to its all-weather, all-day, and high-resolution capabilities, SAR has been widely used in civilian applications. The SAR automatic target recognition (ATR) process refers to identifying the category of an object of interest based on the backscattering characteristics of SAR images. Among different targets, aircraft target recognition plays an important role in civilian airport management, and many researchers have tried to construct a stable and robust feature representation for SAR aircraft recognition in the model space.

[0003] In traditional SAR aircraft recognition methods, the key to the algorithm quality lies in feature extraction. Generally, SAR images reflect geometric and dielectric characteristics. The point responses of aircraft sub-components construct lines and regions, which serve as the primitives of SAR images. Traditional methods attempt to map SAR images to the feature or template space using structural and scattering characteristics. In some techniques, the Hough transform combined with Gaussian kernel smoothing is used to extract the aircraft skeleton, and then, aircraft structure prior segmentation is used to estimate sub-part parameters. There are also existing technologies that study the backscattering characteristics of civilian aircraft to extract significant point vectors with rotational invariance in a small range. In some methods, a Gaussian mixture model (GMM) is also introduced to analyze the scattering intensity distribution, and a sampling selection scheme is designed to achieve target classification. Generally speaking, traditional methods mainly focus on establishing a stable structural representation using scattering characteristics (including points, lines, and statistical distributions). However, they rely to a large extent on manually designed parameters and rules, which limit their application in actual scenarios.

[0004] With its end-to-end characteristics, deep neural networks have witnessed explosive development in SAR ATR. For SAR aircraft target detection and recognition, the main challenges lie in the discrete appearance caused by the complex electromagnetic scattering mechanism and the sensitivity to azimuth angle changes. A large number of scholars have combined the characteristics of aircraft in SAR images and proposed excellent SAR aircraft detection intelligent algorithms. For SAR aircraft classification, currently, methods that introduce structural topology from scattering points into the neural network framework have been widely studied and applied. Some methods utilize scattering characteristics and integrate topological representations into the classification branch. However, they usually require an additional scattering feature extraction branch based on traditional algorithms, which increases the complexity of the algorithm. Moreover, methods that utilize additional high-level information only use rough information and do not fully exploit geometric structure inspiration.

[0005] In the existing technology, an effective alternative to establishing a robust topological representation is to utilize additional high-level information about the structure and component attributes. This idea can originate from the research on aircraft target recognition in optical remote sensing images. The empirical knowledge of remote sensing image interpretation experts effectively promotes the recognition of aircraft targets. In some methods, aircraft classification is transformed into a point regression task, and the shape of the aircraft is described by eight aircraft contour points. There is also a segmentation scheme based on the octagon of the aircraft contour, which provides basic shape information related to the aircraft for target recognition. In addition, there are methods that propose an aircraft important component detection module, in which the aircraft wing and the tail engine are extracted as the main component clues to enhance classification. There are also some methods that construct a fine-grained component parsing database by annotating aircraft parts at the contour level and pixel level. Summary of the Invention

[0006] Based on this, in view of the above technical problems, it is necessary to provide a SAR target recognition method and device based on graph structure consistency alignment that can achieve accurate recognition by fully mining geometric structures.

[0007] A SAR target recognition method based on graph structure consistency alignment, the method comprising:

[0008] Obtain a sample data set, where the sample data set includes multiple SAR sample images and corresponding target category labels;

[0009] For each SAR sample image in the sample data set, construct a corresponding structure model diagram according to the reference diagram of the target type, and generate a set of training samples based on the associated structure model diagram, SAR sample image, and reference diagram, and construct a training data set;

[0010] Input a set of training samples in the training data set into the SAR target recognition network. In the SAR target recognition network, respectively divide the SAR sample image and the structure model diagram in a set of training samples into non-overlapping image blocks of the same size, and re-group the image blocks according to a preset multiple of feature unit sizes to obtain multi-scale feature maps including multiple feature units of the same size, namely multi-scale sample feature maps and multi-scale template feature maps;

[0011] Perform target classification according to the multi-scale sample feature map to obtain a predicted target category, and calculate a classification loss according to the predicted target category and the corresponding target category label;

[0012] Construct corresponding graph structures based on the multi-scale sample feature maps and multi-scale template feature maps respectively, and obtain the multi-scale sample graph structure and the multi-scale template graph structure respectively. Calculate the vertex correlation loss, edge similarity difference loss, and graph spectrum distance loss at the corresponding scale according to the sample graph structure and the template graph structure at the same scale.

[0013] Train the SAR target recognition network according to the vertex correlation loss, edge similarity difference loss, graph spectrum distance loss, and classification loss at different scales until the trained SAR target recognition network is obtained.

[0014] Obtain the SAR image to be target-recognized, and use the trained SAR target recognition network to perform target recognition on the SAR image.

[0015] In one embodiment, the generation of a set of training samples according to the associated structure template graph, SAR sample image, and reference image includes: the reference image is a target optical image, and the structure template graph is a binary graph obtained based on the target optical image.

[0016] Align the structure template graph according to the SAR sample image to obtain the rotated and translated structure template graph.

[0017] Then, perform image matching between the rotated and translated structure template graph and the original structure template graph to obtain the geometric transformation matrix of the target, and extract the relevant information of the target based on the reference image.

[0018] Use the SAR sample image, the corresponding target class label, the structure template graph, the geometric transformation matrix, and the relevant information as a set of training samples.

[0019] In one embodiment, when training the SAR target recognition network, the SAR target recognition network includes a first sub-feature extraction network, a second sub-feature extraction network, and a classifier.

[0020] The first sub-feature extraction network is used to perform multi-scale feature extraction on the SAR sample image to obtain the multi-scale sample feature map.

[0021] The second sub-feature extraction network is used to perform multi-scale feature extraction on the structure template graph to obtain the multi-scale template feature map.

[0022] The classifier includes a global average pooling layer, a fully connected layer, and a Softmax function layer connected in sequence, and is used to perform target class prediction according to the multi-scale sample feature map to obtain the predicted target class.

[0023] In one embodiment, the first sub - feature extraction network adopts the Swin Transformer network structure, and the second sub - feature extraction network adopts the Tokenizer network structure.

[0024] In one embodiment, when constructing the corresponding graph structures according to the multi - scale sample feature maps and multi - scale template feature maps respectively, wherein when constructing the corresponding graph structure according to the multi - scale sample feature maps:

[0025] For each sample feature map of different scales, each feature unit in the sample feature map is used as a node in the graph structure;

[0026] The relationship between the feature unit and other feature units is used as the edge between the corresponding nodes, and the weight of the edge is calculated according to the foreground prior, local constraint, and cosine similarity of the region where the feature unit is located.

[0027] In one embodiment, the sample graph structure of a certain scale is: , where:

[0028] ;

[0029] In the above formula, and represent the feature units at positions and position , represents the domain of all spatial coordinates, represents the weight function, represents the foreground prior of the SAR sample image, represents the local constraint, represents the cosine similarity, and the superscript S represents the parameters related to the sample graph structure.

[0030] In one embodiment, calculating the vertex correlation loss corresponding to the same scale according to the sample graph structure and the template graph structure of the same scale includes:

[0031] For the corresponding feature units in the sample graph structure and the template graph structure of the same scale, a normalized correlation layer is used to obtain a correlation graph;

[0032] A consistency mask is obtained by using a geometric transformation matrix, and the vertex correlation loss is obtained according to the correlation graph and the consistency mask.

[0033] In one embodiment, the trained SAR target recognition network includes the first sub - feature extraction network and a classifier.

[0034] The present application also provides a SAR target recognition device based on graph structure consistency alignment. The device includes:

[0035] A dataset acquisition module, configured to acquire a sample dataset, where the sample dataset includes multiple SAR sample images and corresponding target class labels;

[0036] A training dataset construction module, configured to, for each SAR sample image in the sample dataset, construct a corresponding structural model diagram according to a reference diagram of the target type therein, generate a set of training samples based on the associated structural model diagram, SAR sample image, and reference diagram, and construct a training dataset;

[0037] A multi-scale feature extraction module, configured to input a set of training samples in the training dataset into a SAR target recognition network. In the SAR target recognition network, the SAR sample image and the structural model diagram in the set of training samples are respectively segmented into non-overlapping image patches of the same size, and the image patches are regrouped according to a plurality of preset feature unit sizes to obtain multi-scale feature maps including multiple feature units of the same size, namely a multi-scale sample feature map and a multi-scale template feature map;

[0038] A classification loss calculation module, configured to perform target classification according to the multi-scale sample feature map to obtain a predicted target class, and calculate a classification loss according to the predicted target class and the corresponding target class label;

[0039] A graph structure consistency loss calculation module, configured to respectively construct corresponding graph structures according to the multi-scale sample feature map and the multi-scale template feature map, respectively obtain a multi-scale sample graph structure and a multi-scale model graph structure, and calculate a vertex correlation loss, an edge similarity difference loss, and a graph spectrum distance loss at the corresponding scale according to the sample graph structure and the model graph structure of the same scale;

[0040] A SAR target recognition network training module, configured to train the SAR target recognition network according to the vertex correlation loss, edge similarity difference loss, graph spectrum distance loss, and classification loss at different scales until a trained SAR target recognition network is obtained;

[0041] A SAR image target recognition module, configured to acquire a SAR image to be subjected to target recognition, and perform target recognition on the SAR image by using the trained SAR target recognition network.

[0042] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:

[0043] Obtain a sample data set, where the sample data set includes multiple SAR sample images and corresponding target class labels;

[0044] For each SAR sample image in the sample data set, construct a corresponding structural template diagram according to the reference diagram of the target type therein, generate a set of training samples based on the associated structural template diagram, SAR sample image and reference diagram, and construct a training data set;

[0045] Input a set of training samples in the training data set into the SAR target recognition network. In the SAR target recognition network, divide the SAR sample image and the structural template diagram in a set of training samples into multiple non-overlapping image blocks of the same size respectively, and re-group the image blocks according to a preset multiple of feature unit sizes to obtain multi-scale feature maps including multiple feature units of the same size, namely multi-scale sample feature maps and multi-scale template feature maps;

[0046] Perform target classification according to the multi-scale sample feature map to obtain a predicted target class, and calculate a classification loss according to the predicted target class and the corresponding target class label;

[0047] Construct corresponding graph structures according to the multi-scale sample feature map and the multi-scale template feature map respectively to obtain a multi-scale sample graph structure and a multi-scale template graph structure, and calculate the vertex correlation loss, edge similarity difference loss and graph spectrum distance loss at the corresponding scale according to the sample graph structure and the template graph structure of the same scale;

[0048] Train the SAR target recognition network according to the vertex correlation loss, edge similarity difference loss, graph spectrum distance loss and classification loss at different scales until a trained SAR target recognition network is obtained;

[0049] Obtain a SAR image to be target-recognized, and use the trained SAR target recognition network to perform target recognition on the SAR image.

[0050] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0051] Obtain a sample data set, where the sample data set includes multiple SAR sample images and corresponding target class labels;

[0052] For each SAR sample image in the sample data set, construct a corresponding structural template diagram according to the reference diagram of the target type therein, generate a set of training samples based on the associated structural template diagram, SAR sample image and reference diagram, and construct a training data set;

[0053] Input a set of training samples in the training dataset into the SAR target recognition network. In the SAR target recognition network, split the SAR sample images and the structure template images in a set of training samples into multiple non-overlapping image patches of the same size, and re-group the image patches according to a preset multiple of feature unit sizes to obtain multi-scale feature maps including multiple feature units of the same size, namely multi-scale sample feature maps and multi-scale template feature maps respectively;

[0054] Perform target classification based on the multi-scale sample feature maps to obtain predicted target categories, and calculate classification losses according to the predicted target categories and the corresponding target category labels;

[0055] Construct corresponding graph structures based on the multi-scale sample feature maps and the multi-scale template feature maps respectively to obtain a multi-scale sample graph structure and a multi-scale template graph structure respectively, and calculate vertex correlation losses, edge similarity difference losses, and graph spectrum distance losses at corresponding scales according to the sample graph structure and the template graph structure of the same scale;

[0056] Train the SAR target recognition network according to the vertex correlation losses, edge similarity difference losses, graph spectrum distance losses, and classification losses at different scales until the trained SAR target recognition network is obtained;

[0057] Obtain a SAR image to be target-recognized, and perform target recognition on the SAR image by using the trained SAR target recognition network.

[0058] Beneficial effects: The above SAR target recognition method and device based on graph structure consistency alignment construct structure template images according to target reference images for each SAR sample image in the sample dataset, generate training samples and construct a training set, input it into the SAR target recognition network, split the sample images and the template images into blocks, re-group them according to preset sizes to obtain multi-scale feature maps, then perform target classification based on the multi-scale sample feature maps and calculate classification losses, and then construct graph structures to calculate vertex correlation, edge similarity difference, and graph spectrum distance losses. Use these losses to train the network until convergence, and use the trained network to recognize the SAR image to be processed. Using this method can effectively improve the target recognition accuracy. Description of the Drawings

[0059] Figure 1 It is a schematic flowchart of the SAR target recognition method based on graph structure consistency alignment in an embodiment;

[0060] Figure 2 It is a schematic diagram of graph structure consistency constructed by the similarity between SAR images and structure template images at different azimuth angles in an embodiment, where Figure 2 (a)andFigure 2 (c) is the SAR image of the same aircraft at different azimuth angles, Figure 2 (b) is the structural template diagram;

[0061] Figure 3 is a schematic diagram of the aircraft structure reference template in an embodiment;

[0062] Figure 4 is a schematic diagram of the training of the SAR target recognition network in an embodiment;

[0063] Figure 5 is a schematic diagram of encoding the structural template diagram using a tokenizer in an embodiment;

[0064] Figure 6 is a structural block diagram of a SAR target recognition device based on graph structure consistency alignment in an embodiment;

[0065] Figure 7 is the internal structure diagram of a computer device in an embodiment. Specific implementation manners

[0066] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0067] Considering the technical defects existing in the prior art and further taking into account that different from the optical imaging mechanism, the imaging result of a SAR aircraft consists of a large number of discrete strong scattering points. Although there are supervision signals, the learning process still appears chaotic and unconstrained. The key to the recognition algorithm is to realize the continuous geometric invariant structure representation of the aircraft. Therefore, two sub-problems must be clearly solved in the present application: (1) How to establish a representation closely related to the aircraft structure from existing advanced neural networks. (2) How to maintain the geometric and topological consistency between these representations. Furthermore, as Figure 1 shown, a SAR target recognition method based on graph structure consistency alignment is provided in the application, including the following steps:

[0068] Step S100, obtain a sample data set, which includes multiple SAR sample images and corresponding target category labels.

[0069] Step S110, for each SAR sample image in the sample data set, construct a corresponding structural template diagram according to the reference diagram of the target type, generate a set of training samples according to the associated structural template diagram, SAR sample image and reference diagram, and construct a training data set.

[0070] Step S120: Input a set of training samples in the training dataset into the SAR target recognition network. In the SAR target recognition network, the SAR sample images and the structure template images in the set of training samples are respectively segmented into multiple non-overlapping image patches of the same size, and then regrouped according to a preset multiple of feature unit sizes to obtain multi-scale feature maps including multiple feature units of the same size, namely multi-scale sample feature maps and multi-scale template feature maps.

[0071] Step S130: Perform target classification based on the multi-scale sample feature maps to obtain predicted target categories, and calculate classification losses based on the predicted target categories and the corresponding target category labels.

[0072] Step S140: Construct corresponding graph structures based on the multi-scale sample feature maps and the multi-scale template feature maps respectively to obtain a multi-scale sample graph structure and a multi-scale template graph structure, and calculate the vertex correlation loss, edge similarity difference loss, and graph spectrum distance loss at the corresponding scales based on the sample graph structure and the template graph structure of the same scale.

[0073] Step S150: Train the SAR target recognition network according to the vertex correlation loss, edge similarity difference loss, graph spectrum distance loss, and classification loss at different scales until the trained SAR target recognition network is obtained.

[0074] Step S160: Obtain the SAR image to be target-recognized, and perform target recognition on the SAR image using the trained SAR target recognition network.

[0075] In this application, a method based on graph structure alignment is proposed to improve the accuracy of target recognition. In this method, the SAR image is segmented into multiple tokens by a neural network, and then regrouped according to different sizes to obtain multi-scale feature maps, and then the multi-scale feature maps are transformed into graph structures, and a loss function is used to ensure the structural consistency between the SAR image and the template image.

[0076] Specifically, for the first sub-question proposed above, in this method, a graph structure is introduced to explicitly establish the topological structure of the target. Graph structures have been widely used in real-world applications to process data linked by certain specific attributes, such as biology, social media, and finance. Graph structures can capture the inherent structural information of images and play an important role in visual tasks. Different from the traditional method of constructing graph structures using features, self-similarity is applied to the tokens extracted from the neural network in this method. Nodes and edges are defined by the self-similarity between tokens, and the proposed algorithm can combine a parameter model with end-to-end backpropagation to dynamically update the graph structure.

[0077] Specifically, for the second sub-question raised above, a graph structure alignment method is proposed in this method to capture geometric invariant representations. As Figure 2 shown, airplanes belonging to the same category but with different poses have a consistent topological structure. Therefore, a template graph structure is set for each category of targets in this method, and this template graph structure is constructed based on the tokens extracted from the target structure template. By iteratively restricting the distance between the graph structure generated by the SAR target image tokens and the graph structure generated by the target template image tokens, the structural consistency of the features can be ensured without being affected by pose changes. To this end, graph structure alignment losses are further designed at the vertex, edge, and graph spectrum levels respectively. The vertex and edge distances can measure the structural differences and the graph spectrum distance can measure the signal differences. Utilizing the topological and graph spectrum characteristics, the structural differences between two tokens can be better captured.

[0078] In step S100, according to the specific application background, SAR sample images of relevant targets are obtained. In this application, airplanes are used as an example for illustration. In fact, the method in this application can also be applied to the recognition of other targets. The sample dataset includes SAR sample images of various different types of airplanes and the corresponding airplane category targets.

[0079] Considering that an airplane is a rigid object with a fixed structure, in this method, a template-based airplane geometric annotation method is designed. In step S110, first, according to the target category in the SAR sample image, a corresponding target optical image is generated from the high-resolution RSI obtained by Google Earth, which is the reference image. Then, the target in the reference image is binary segmented to obtain the target binary image, that is, the structure template image.

[0080] Next, the structure template image is aligned according to the corresponding SAR sample image to obtain the structure template image after rotation and translation. Then, the structure template image after rotation and translation is image-matched with the original structure template image to obtain the geometric transformation matrix of the target, and relevant information of the target, including the target length and width information, is extracted based on the reference image. And the SAR sample image, the corresponding target category label, the structure template image, the geometric transformation matrix, and the relevant information are used as a set of training samples. According to each SAR sample image in the sample dataset, a corresponding set of training samples is generated, thereby constructing a training dataset.

[0081] As Figure 3 shown, a schematic diagram of the reference image, the structure template image, and the relevant information with an airplane as the target, where the first row is the optical image, that is, the reference image, the second row is the structure template image, and the third and fourth rows are the length and width information of the airplane.

[0082] Specifically, when performing image matching between the structural layout diagram and the corresponding SAR sample image to obtain the geometric transformation matrix of the target , it is expressed as:

[0083] ;

[0084] Furthermore, when converting the point to , the mapping function used is expressed as:

[0085] ;

[0086] Different from other existing aircraft annotation methods, including point-based, polygon-based, pose-based, and component-based annotation methods, the introduced template-based annotation method has the following advantages: First, it avoids fine annotation of pixels or boundaries, thus ensuring the convenience of operation; Second, it can derive information equivalent to that provided by other annotation methods. For example, the pose can be obtained from in , and the foreground and background interpretation of pixels can be obtained from the template at its transformed position according to . Therefore, for each SAR aircraft image (SAR sample image) , this solution provides additional information including the corresponding structural template diagram and the geometric transformation .

[0087] In step S120, a set of training samples in the training dataset is input into the SAR target recognition network and trained in the SAR target recognition network. When training the SAR target recognition network, as Figure 4 shown, the SAR target recognition network includes a first sub-feature extraction network, a second sub-feature extraction network, and a classifier. Among them, the first sub-feature extraction network is used to perform multi-scale feature extraction on the SAR sample image to obtain a multi-scale sample feature map. The second sub-feature extraction network is used to perform multi-scale feature extraction on the structural layout diagram to obtain a multi-scale template feature map. The classifier includes a global average pooling layer, a fully connected layer, and a Softmax function layer connected in sequence, and is used to perform target category prediction according to the multi-scale sample feature map to obtain the predicted target category.

[0088] Specifically, the first sub-feature extraction network adopts the Swin Transformer network structure, and the second sub-feature extraction network adopts the Tokenizer network structure.

[0089] In this embodiment, the Swin Transformer network first serves as the feature extraction network for the SAR sample image. In the Swin Transformer, the size of The image is first segmented into non-overlapping blocks, and each block is regarded as a token. After a series of Swin Transformer blocks and Patching Merge blocks, the concatenated input tokens are mapped to tokens of various sizes. In this method, the re-partitioned tokens are also represented by feature units, that is, multiple feature maps with different scales of feature units are obtained. It should be noted here that in the same feature map, the sizes of the feature units are the same.

[0090] Specifically, in the Swin Transformer network, after the SAR sample image is initially divided into multiple non-overlapping tokens, the multiple tokens are reorganized according to a preset multiple different sizes to obtain new tokens, that is, feature units. In this way, a SAR sample image can generate sample feature maps of multiple scales under feature units of different sizes.

[0091] In one embodiment, the SAR sample image passes through the Swin Transformer network to obtain sample feature maps of three scales, and their scale sizes include , and .

[0092] In step S130, after the Swin Transformer network, a classifier is also connected. Among them, global average pooling (GAP), fully connected (FC) layer and Softmax function are applied to the last token to obtain the classification probability. In terms of classification loss, cross entropy (CE) is used to evaluate the difference between the prediction and the true label, as follows:

[0093] ;

[0094] In the above formula, represents the number of categories.

[0095] In this embodiment, an independent second sub-feature extraction network, namely the Tokenizer network, is also designed to encode the structure layout diagram into a template feature map with the same size as the sample feature map. As Figure 5 shown, two examples are shown, that is, the template map with a size of is converted into tokens with sizes of and , that is, multi-scale layout diagrams.

[0096] In one embodiment, the downsampling rate in the Tokenizer network Set to 8, 16, and 32. In particular, for each downsampling, the tokenizer pixels are sampled and the structural template graph is divided into Then, these feature maps are stacked into a scale of , the dimension is The feature map of . Then a 3×3 convolutional layer is applied to align the token dimension with the output of the Swin Transformer. It can be observed that the designed tokenizer does not lose any spatial structure information, that is, the position relationship of the pixels does not change before and after encoding.

[0097] In step S140, when constructing corresponding graph structures according to the multi-scale sample feature map and the multi-scale template feature map respectively, wherein, when constructing the corresponding graph structure according to the multi-scale sample feature map: for sample feature maps of different scales, each feature unit in the sample feature map is used as a node in the graph structure, and the relationship between the feature unit and other feature units is used as the edge between the corresponding nodes, and the weight of the edge is calculated based on the foreground prior, local constraints and cosine similarity of the area where the feature unit is located.

[0098] In this example, a series of tokens extracted from SAR images are ,The graph structure can effectively capture the key information and local structure of the image.,Therefore, in this method, a corresponding graph structure is constructed for each scale of,the sample feature graph.

[0099] Specifically, the sample graph structure of a certain scale is: ,in:

[0100] ;

[0101] In the above formula, and Indicates location and location The characteristic unit of represents the domain of all spatial coordinates, represents the weight function, and Respectively represent the foreground priors of the SAR sample image and the structure template image, satisfying as well as , represents a local constraint, It represents cosine similarity, and the superscript S represents the parameter related to the sample graph structure.

[0102] Among them, local constraints It is expressed as:

[0103] ;

[0104] Among them, the cosine similarity is expressed as:

[0105] .

[0106] In this embodiment, local constraints are used to highlight the importance of the edges with a short distance between the starting point and the ending point, and this item is crucial for inheriting the intrinsic distance attributes in the Euclidean space of the original feature map.

[0107] Similarly, in the graph structure, the aircraft structure information can be effectively characterized at the token level. To explore the differences in the graphs, a similar graph structure is also constructed from the template feature maps at various scales, expressed as:

[0108] ;

[0109] Furthermore, the consistency of the graph structure is determined by measuring the distances between the multi-scale sample graph structures and the multi-scale model graph structures. It can be known that the graph structure information can be found in its topology and spectrum. In particular, the edge weights determine the amount of information diffused from a vertex to its neighbors, and the largest spectrum can be used to detect the energy changes caused by abnormal vertices. Using these two attributes, the potential differences in the graph structures can be measured. Therefore, a graph structure alignment method including three different aspects is proposed, as shown in the three sub-graphs below in Figure 4 , that is, three loss functions are used to ensure the consistency of the target structures between the SAR sample graph and the structure model graph during iterative training.

[0110] In this embodiment, the vertex correlation loss is introduced first. Calculate the vertex correlation loss at the corresponding scale according to the sample graph structure and the model graph structure at the same scale, including: for the corresponding feature units in the sample graph structure and the model graph structure at the same scale, use the normalized correlation layer to obtain the correlation graph, use the geometric transformation matrix to obtain the consistency mask, and obtain the vertex correlation loss according to the correlation graph and the consistency mask.

[0111] Specifically, for use the normalized correlation layer to obtain the correlation graph , the calculation is as follows:

[0112] ;

[0113] In the above formula, the normalized inner product of the token at position in and the token at position in Finally, it includes the one-way correlation of all feature pairs in the SAR image and the template.

[0114] Similarly, the correlation graph in the opposite direction is:

[0115] ;

[0116] Given the geometric transformation matrix , the feature pairs with geometric consistency can be confirmed, and the incorrect matches that do not conform to the geometric correspondence can be eliminated. This is achieved through two consistency masks, where the two consistency masks can be expressed as:

[0117] ;

[0118] ;

[0119] In the above formula, is the L2 distance function, and the preset threshold is 1. This definition only considers the vertex correspondence when the projection error is not greater than . Then, by multiplying the correlation graph and the consistency mask and summing them, the loss of GSA at the vertex level, that is, the vertex correlation loss, can be obtained, which is expressed as:

[0120] ;

[0121] The training objective is to maximize the correlation between vertex pairs and geometric consensus. Note that due to the normalized correlation operation, it also penalizes the ambiguous correlations with inconsistent correspondences.

[0122] In this embodiment, since the weighted edges contain rich topological structure relationships, the edge similarity difference can be directly used to measure the difference of the graph, and its calculation is:

[0123] ;

[0124] To maintain the structural consistency, it is necessary to reduce the difference of the semantically coherent edges in the graphs constructed from the SAR image and the template. "Semantically coherent edges" refer to the edge pairs where the starting and ending points of the two domains are geometrically consistent with the transformation , such as the left wing to the left wing, rather than the left wing to the right wing. A simple and intuitive method is to introduce an edge correspondence mask, as follows:

[0125] ;

[0126] However, the loss at the edge level, that is, the edge similarity difference, is expressed as:

[0127] ;

[0128] Although this definition is intuitive and easy to understand, the space complexity of summing all edge pairs over two domains . Suppose the graph has 100 vertices, and the dimension of the edge similarity difference matrix and the consistency mask matrix is , which is difficult to compute and store. Further observation shows that is a sparse matrix, that is, among the elements, only elements are 1.

[0129] Therefore, instead of directly comparing the edge differences between two graphs, geometrically transform the graph into an isomorphic graph of , expressed as: ;

[0130] ;

[0131] Finally, the loss of the graph structure at the edge level is expressed as:

[0132] ;

[0133] In the above formula, it is equivalent to directly summing the non-zero terms in and reducing the space complexity to .

[0134] In this embodiment, the graph Laplacian is a linear shift-invariant signal filter. For the graph , the graph filter generally includes the standard Laplacian matrix , the symmetric normalized Laplacian matrix and the random walk normalized Laplacian matrix , where is the degree matrix satisfying , is the edge matrix. The graph spectrum distance can quantitatively measure the structural differences between two isomorphic graphs with the same number of nodes .

[0135] and 's graph spectrum representations adopt eigenvalue decomposition and , and the eigenvalues are expressed as:

[0136] ;

[0137] ;

[0138] However, the loss of GSA at the spectral level can be represented by the distance between two spectral vectors as follows:

[0139] ;

[0140] Based on , the topological structures between the tokens extracted from the SAR image and the template can be aligned to the same spectral graph and finally have the same spectral graph properties.

[0141] During the implementation process, the dual-path feature extractor generates features of different scales, namely , and . Then, multi-scale graphs are constructed from these features, which are represented as and , where . This can capture the topological structures between aircraft components of different sizes. Therefore, the final optimization objective function is:

[0142] ;

[0143] In the above formula, , and represent the weight terms of the loss, , and represent the GSA losses obtained from graphs and .

[0144] In step 150, under the training of the final optimization objective function, the SAR target recognition network is obtained after training convergence. The trained SAR target recognition network includes the first sub-feature extraction network and the classifier, and the second sub-feature extraction network is removed. In step S160, when actually applying the trained SAR target recognition network, only the SAR image to be detected needs to be input into the SAR target recognition network to achieve target recognition and detection.

[0145] As shown in Algorithm 1 in Table 1, the content is the training process of the SAR target recognition network (GSA-Trans). The GSA loss provides spatial regularization for identifying structural class relationships and effectively improves the discriminability of features. In the testing phase, only the feature extraction and classifier of the SAR image are retained to generate prediction labels. Therefore, it is still an end-to-end feature;

[0146] Table 1

[0147]

[0148] In the above SAR target recognition method based on graph structure consistency alignment, for each SAR sample image in the sample dataset, a structural model graph is constructed according to the target reference graph to generate training samples and construct a training set. The training set is input into the SAR target recognition network. The sample image and the model graph are segmented into blocks and recombined into a multi-scale feature map according to a preset size. Then, target classification is performed based on the multi-scale sample feature map, and the classification loss is calculated. Next, a graph structure is constructed to calculate the vertex correlation, edge similarity difference, and graph spectrum distance loss. The network is trained with these losses until convergence, and the trained network is used to recognize the SAR image to be processed. This method proposes a graph structure construction method, which converts the feature map extracted by the network into a graph structure closely related to the target structure, and designs a graph alignment network loss function for the structural consistency feature by measuring graph differences at the vertices, edges, and graph spectra of the graph structure respectively. The superiority of the proposed method is verified by experiments on public and self-built SAR aircraft recognition datasets.

[0149] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0150] In one embodiment, as Figure 6 shown, a SAR target recognition device based on graph structure consistency alignment is provided, including: a dataset acquisition module 200, a training dataset construction module 210, a multi-scale feature extraction module 220, a classification loss calculation module 230, a graph structure consistency loss calculation module 240, a SAR target recognition network training module 250, and a SAR image target recognition module 260, where:

[0151] The dataset acquisition module 200 is configured to acquire a sample dataset, where the sample dataset includes multiple SAR sample images and corresponding target category labels;

[0152] The training dataset construction module 210 is configured to, for each SAR sample image in the sample dataset, construct a corresponding structural model graph according to the reference graph of the target type, and generate a set of training samples according to the associated structural model graph, SAR sample image, and reference graph, and construct a training dataset;

[0153] The multi-scale feature extraction module 220 is configured to input a set of training samples in the training dataset into the SAR target recognition network. In the SAR target recognition network, the SAR sample images and the structure template images in a set of training samples are respectively segmented into multiple non-overlapping image patches of the same size, and re-grouped according to a plurality of preset feature unit sizes to obtain multi-scale feature maps including a plurality of feature units of the same size, namely, a multi-scale sample feature map and a multi-scale template feature map;

[0154] The classification loss calculation module 230 is configured to perform target classification based on the multi-scale sample feature map to obtain a predicted target category, and calculate a classification loss according to the predicted target category and the corresponding target category label;

[0155] The graph structure consistency loss calculation module 240 is configured to respectively construct corresponding graph structures based on the multi-scale sample feature map and the multi-scale template feature map, obtain a multi-scale sample graph structure and a multi-scale template graph structure respectively, and calculate the vertex correlation loss, the edge similarity difference loss, and the graph spectrum distance loss at the corresponding scale according to the sample graph structure and the template graph structure at the same scale;

[0156] The SAR target recognition network training module 250 is configured to train the SAR target recognition network according to the vertex correlation loss, the edge similarity difference loss, the graph spectrum distance loss, and the classification loss at different scales until the trained SAR target recognition network is obtained;

[0157] The SAR image target recognition module 260 is configured to obtain a SAR image to be target-recognized, and perform target recognition on the SAR image by using the trained SAR target recognition network.

[0158] For the specific limitations of the SAR target recognition device based on graph structure consistency alignment, reference may be made to the limitations of the SAR target recognition method based on graph structure consistency alignment in the above text, which will not be elaborated here. Each module in the above SAR target recognition device based on graph structure consistency alignment can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0159] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as shown in Figure 7As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a SAR target recognition method based on graph structure consistency alignment. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0160] Those skilled in the art can understand that Figure 7 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0161] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0162] Obtain a sample data set, where the sample data set includes multiple SAR sample images and corresponding target category labels;

[0163] For each SAR sample image in the sample data set, construct a corresponding structure model diagram according to the reference diagram of the target type therein, generate a set of training samples according to the associated structure model diagram, SAR sample image, and reference diagram, and construct a training data set;

[0164] Input a set of training samples in the training data set into the SAR target recognition network. In the SAR target recognition network, divide the SAR sample image and the structure model diagram in a set of training samples into multiple non-overlapping image blocks of the same size, and re-group the image blocks according to a preset multiple of feature unit sizes to obtain a multi-scale feature map including multiple feature units of the same size, namely a multi-scale sample feature map and a multi-scale template feature map;

[0165] Perform target classification according to the multi-scale sample feature map to obtain a predicted target category, and calculate a classification loss according to the predicted target category and the corresponding target category label;

[0166] Construct corresponding graph structures based on the multi-scale sample feature maps and the multi-scale template feature maps respectively, and obtain the multi-scale sample graph structure and the multi-scale template graph structure respectively. Calculate the vertex correlation loss, edge similarity difference loss, and graph spectrum distance loss at the corresponding scale according to the sample graph structure and the template graph structure at the same scale;

[0167] Train the SAR target recognition network according to the vertex correlation loss, edge similarity difference loss, graph spectrum distance loss, and classification loss at different scales until the trained SAR target recognition network is obtained;

[0168] Obtain the SAR image to be target-recognized, and use the trained SAR target recognition network to perform target recognition on the SAR image.

[0169] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0170] Obtain a sample data set, where the sample data set includes multiple SAR sample images and corresponding target class labels;

[0171] For each SAR sample image in the sample data set, construct a corresponding structure template graph according to the reference graph of the target type, and generate a set of training samples according to the associated structure template graph, SAR sample image, and reference graph, and construct a training data set;

[0172] Input a set of training samples in the training data set into the SAR target recognition network. In the SAR target recognition network, divide the SAR sample image and the structure template graph in a set of training samples into non-overlapping image blocks of the same size respectively, and re-group the image blocks according to a preset multiple of feature unit sizes to obtain multi-scale feature maps including multiple feature units of the same size, namely multi-scale sample feature maps and multi-scale template feature maps;

[0173] Perform target classification according to the multi-scale sample feature map to obtain a predicted target class, and calculate a classification loss according to the predicted target class and the corresponding target class label;

[0174] Construct corresponding graph structures based on the multi-scale sample feature maps and the multi-scale template feature maps respectively, and obtain the multi-scale sample graph structure and the multi-scale template graph structure respectively. Calculate the vertex correlation loss, edge similarity difference loss, and graph spectrum distance loss at the corresponding scale according to the sample graph structure and the template graph structure at the same scale;

[0175] Train the SAR target recognition network according to the vertex correlation loss, edge similarity difference loss, graph spectrum distance loss, and classification loss at different scales until the trained SAR target recognition network is obtained;

[0176] Obtain the SAR image to be recognized, and use the trained SAR target recognition network to recognize the target in the SAR image.

[0177] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0178] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0179] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A SAR target recognition method based on graph structure consistency alignment, characterized in that: The method comprises: Acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; For each SAR sample image in the sample data set, a corresponding structure template map is constructed according to a reference map of the target type therein, and a group of training samples are generated according to the associated structure template map, SAR sample image and reference map to construct a training data set; A group of training samples in the training data set is input into a SAR target recognition network. In the SAR target recognition network, a SAR sample image and a structure template map in a group of training samples are respectively divided into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; Performing target classification according to the multi-scale sample feature map to obtain a predicted target category, and performing calculation according to the predicted target category and the corresponding target category label to obtain a classification loss; The corresponding graph structures are constructed according to the multi-scale sample feature graph and the multi-scale template feature graph, respectively, and the multi-scale sample graph structure and the multi-scale template graph structure are obtained respectively. According to the sample graph structure and the template graph structure of the same scale, the vertex correlation loss, the edge similarity difference loss and the graph distance loss at the corresponding scale are calculated respectively; The SAR target recognition network is trained according to vertex correlation loss, edge similarity difference loss, graph distance loss, and classification loss at different scales until a trained SAR target recognition network is obtained; A SAR image to be subjected to target recognition is acquired, and target recognition is performed on the SAR image using the trained SAR target recognition network.

2. The SAR target recognition method based on graph structure consistency alignment according to claim 1 is characterized in that: The step of generating a group of training samples according to the associated structure template image, SAR sample image and reference image comprises: the reference image is a target optical image, and the structure template image is a binary image obtained based on the target optical image; Aligning the structure template image according to the SAR sample image to obtain a rotated and translated structure template image; Then, the rotated and translated structure template image is matched with the original structure template image to obtain the geometric transformation matrix of the target, and relevant information of the target is extracted based on the reference image; The SAR sample image, the corresponding target category label, the structure template map, the geometric transformation matrix and related information are used as a group of training samples.

3. The SAR target recognition method based on graph structure consistency alignment according to claim 2 is characterized in that: When training the SAR target recognition network, the SAR target recognition network includes a first sub-feature extraction network, a second sub-feature extraction network and a classifier; The first sub-feature extraction network is used to perform multi-scale feature extraction on the SAR sample image to obtain the multi-scale sample feature map; The second sub-feature extraction network is used to perform multi-scale feature extraction on the structure template map to obtain the multi-scale template feature map; The classifier includes a global average pooling layer, a fully connected layer and a Softmax function layer connected in sequence, and is used to predict the target category according to the multi-scale sample feature map to obtain the predicted target category.

4. The SAR target recognition method based on graph structure consistency alignment according to claim 3 is characterized in that: The first sub-feature extraction network adopts a Swin Transformer network structure, and the second sub-feature extraction network adopts a Tokenizer network structure.

5. The SAR target recognition method based on graph structure consistency alignment according to any one of claims 3 to 4, characterized in that: When corresponding graph structures are constructed according to the multi-scale sample feature graph and the multi-scale template feature graph respectively, wherein, when corresponding graph structures are constructed according to the multi-scale sample feature graph: For sample feature graphs of different scales, each feature unit in the sample feature graph is used as a node in the graph structure; The relationship between the feature unit and other feature units is used as the edge between corresponding nodes, and the weight of the edge is calculated according to the foreground prior, local constraints and cosine similarity of the area where the feature unit is located.

6. The SAR target recognition method based on graph structure consistency alignment according to claim 5 is characterized in that: The sample graph structure of a certain scale is: ,in: In the above formula, and Indicates location and location The characteristic unit of represents the domain of all spatial coordinates, represents the weight function, represents the foreground prior of the SAR sample image, represents a local constraint, It represents cosine similarity, and the superscript S represents a parameter related to the sample graph structure.

7. The SAR target recognition method based on graph structure consistency alignment according to claim 6 is characterized in that: The step of calculating the vertex correlation loss at corresponding scales according to the sample graph structure and the template graph structure of the same scale includes: For the sample graph structure and the corresponding feature units in the template graph structure at the same scale, a normalized correlation layer is used to obtain a correlation graph; A consistency mask is obtained using a geometric transformation matrix, and the vertex correlation loss is obtained according to the correlation graph and the consistency mask.

8. The SAR target recognition method based on graph structure consistency alignment according to claim 7, characterized in that: The trained SAR target recognition network includes the first sub-feature extraction network and a classifier.

9. A SAR target recognition device based on graph structure consistency alignment, characterized in that: The device comprises: A data set acquisition module is used to acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; A training data set construction module is used to construct a corresponding structure template map according to a reference map of the target type for each SAR sample image in the sample data set, generate a set of training samples according to the associated structure template map, SAR sample image and reference map, and construct a training data set; A multi-scale feature extraction module, used for inputting a group of training samples in the training data set into a SAR target recognition network, in which a SAR sample image and a structure template map in a group of training samples are respectively segmented into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; A classification loss calculation module is used to classify the target according to the multi-scale sample feature map to obtain a predicted target category, and to calculate the classification loss according to the predicted target category and the corresponding target category label; A graph structure consistency loss calculation module is used to construct corresponding graph structures according to multi-scale sample feature graphs and multi-scale template feature graphs, respectively, to obtain multi-scale sample graph structures and multi-scale template graph structures, and to calculate vertex correlation loss, edge similarity difference loss and graph distance loss at corresponding scales according to the sample graph structure and template graph structure of the same scale; A SAR target recognition network training module is used to train the SAR target recognition network according to vertex correlation loss, edge similarity difference loss, graph distance loss and classification loss at different scales until a trained SAR target recognition network is obtained; The SAR image target recognition module is used to obtain a SAR image to be subjected to target recognition, and to perform target recognition on the SAR image using the trained SAR target recognition network.

Citation Information

Patent Citations

  • Stochastic gradient Bayesian SAR image segmentation method based on sketch structure

    CN106611422A

  • SAR image target identification method and device

    CN110110625A