SAR (Synthetic Aperture Radar) target identification method and device based on graph structure consistency alignment
By constructing multi-scale feature maps and calculating related losses based on the consistent alignment of graph structures, the SAR target recognition network is trained, and the problem that SAR aircraft recognition method relies on manual parameters and rules in the existing technology is solved, achieving high-precision recognition effect.
Patent Information
- Application Number
- CN202510487867.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing SAR aircraft recognition methods rely on manual design parameters and rules, resulting in limited application in actual scenarios and it is difficult to fully explore geometric structures for accurate identification.
Using a method based on graph structure consistency alignment, by obtaining sample data sets, constructing structural template maps and multi-scale feature maps, vertex correlation, edge similarity difference and map distance loss are calculated, and these losses are used to train the SAR target recognition network to achieve accurate recognition of SAR images to be recognized.
It effectively improves the accuracy of SAR target recognition, reduces dependence on manual design parameters, fully explores geometric structure information, and improves the robustness and efficiency of recognition.
Smart Images

Figure CN120014374A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer target recognition, and in particular to a SAR target recognition method and device based on graph structure consistency alignment. Background Art
[0002] As an active sensor, the synthetic aperture radar (SAR) system is able to measure the backscattering characteristics of the target by emitting radar waves. SAR has been widely used in civil applications due to its all-weather, all-day and high-resolution capabilities. The SAR automatic target recognition (ATR) process refers to identifying the class of the object of interest based on the backscattering characteristics of the SAR image. Among different targets, aircraft target recognition plays an important role in civil airport management. Many researchers have tried to build a stable and robust feature representation for SAR aircraft recognition in the model space.
[0003] In traditional SAR aircraft recognition methods, the key to algorithm quality lies in feature extraction. Generally, SAR images reflect geometric and dielectric properties. Point responses of aircraft sub-components construct lines and regions as primitives of SAR images. Traditional methods attempt to map SAR images to feature or template space using structural and scattering properties. Some techniques use Hough transform combined with Gaussian kernel smoothing to extract the aircraft skeleton, and then use the aircraft structure prior segmentation to estimate sub-part parameters. Other existing technologies study the backscattering characteristics of civil aircraft to extract significant point vectors that are rotationally invariant in a small range. Some methods also introduce Gaussian mixture models (GMM) to analyze the scattering intensity distribution and design a sampling selection scheme to achieve target classification. In general, traditional methods mainly focus on using scattering characteristics (including points, lines, and statistical distributions) to establish stable structural representations. However, they rely heavily on manually designed parameters and rules, which limit their application in practical scenarios.
[0004] Deep neural networks have experienced explosive growth in SAR ATR due to their end-to-end characteristics. For SAR aircraft target detection and recognition, the main challenges are the discrete appearance and sensitivity to azimuth changes caused by complex electromagnetic scattering mechanisms. A large number of scholars have proposed intelligent algorithms for SAR aircraft detection with excellent performance, combining the characteristics of aircraft in SAR images. For SAR aircraft classification, the method of introducing structural topology from scattering points into the neural network framework has been widely studied and applied. Some methods use scattering characteristics to integrate topological representation into the classification branch. However, they usually require an additional scattering feature extraction branch based on traditional algorithms, which increases the complexity of the algorithm. If the method uses additional high-level information, it only uses rough information and does not fully explore the geometric structure inspiration.
[0005] In the existing technology, an effective alternative to building a robust topological representation is to utilize additional high-level information about the structure and component properties. This idea can be traced back to the research on aircraft target recognition in optical remote sensing images. The empirical knowledge of remote sensing image interpretation experts effectively promotes the recognition of aircraft targets. In some methods, aircraft classification is converted into a point regression task, and the shape of the aircraft is described by eight aircraft contour points. There is also a segmentation scheme based on the aircraft contour octagon, which provides basic shape information related to the aircraft for target recognition. In addition, there are methods that propose an important aircraft component detection module, in which the aircraft wing and tail engine are extracted as the main component clues to enhance classification. There are also some methods that construct a fine-grained component parsing database by annotating aircraft parts at the contour level and pixel level. Summary of the invention
[0006] Based on this, it is necessary to provide a SAR target recognition method and device based on graph structure consistency alignment, which can achieve accurate recognition by fully mining the geometric structure, in order to address the above technical problems.
[0007] A SAR target recognition method based on graph structure consistency alignment, the method comprising: Acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; For each SAR sample image in the sample data set, a corresponding structure template map is constructed according to a reference map of the target type therein, and a group of training samples are generated according to the associated structure template map, SAR sample image and reference map to construct a training data set; A group of training samples in the training data set is input into a SAR target recognition network. In the SAR target recognition network, a SAR sample image and a structure template map in a group of training samples are respectively divided into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; Performing target classification according to the multi-scale sample feature map to obtain a predicted target category, and performing calculation according to the predicted target category and the corresponding target category label to obtain a classification loss; The corresponding graph structures are constructed according to the multi-scale sample feature graph and the multi-scale template feature graph, respectively, and the multi-scale sample graph structure and the multi-scale template graph structure are obtained respectively. According to the sample graph structure and the template graph structure of the same scale, the vertex correlation loss, the edge similarity difference loss and the graph distance loss at the corresponding scale are calculated respectively; The SAR target recognition network is trained according to vertex correlation loss, edge similarity difference loss, graph distance loss, and classification loss at different scales until a trained SAR target recognition network is obtained; A SAR image to be subjected to target recognition is acquired, and target recognition is performed on the SAR image using the trained SAR target recognition network.
[0008] In one embodiment, generating a set of training samples according to the associated structure template image, SAR sample image and reference image includes: the reference image is a target optical image, the structure template image is a binary image obtained based on the target optical image; Aligning the structure template image according to the SAR sample image to obtain a rotated and translated structure template image; Then, the rotated and translated structure template image is matched with the original structure template image to obtain the geometric transformation matrix of the target, and relevant information of the target is extracted based on the reference image; The SAR sample image, the corresponding target category label, the structure template map, the geometric transformation matrix and related information are used as a group of training samples.
[0009] In one embodiment, when the SAR target recognition network is trained, the SAR target recognition network includes a first sub-feature extraction network, a second sub-feature extraction network, and a classifier; The first sub-feature extraction network is used to perform multi-scale feature extraction on the SAR sample image to obtain the multi-scale sample feature map; The second sub-feature extraction network is used to perform multi-scale feature extraction on the structure template map to obtain the multi-scale template feature map; The classifier includes a global average pooling layer, a fully connected layer and a Softmax function layer connected in sequence, and is used to predict the target category according to the multi-scale sample feature map to obtain the predicted target category.
[0010] In one embodiment, the first sub-feature extraction network adopts a Swin Transformer network structure, and the second sub-feature extraction network adopts a Tokenizer network structure.
[0011] In one embodiment, when corresponding graph structures are constructed according to the multi-scale sample feature graph and the multi-scale template feature graph respectively, when the corresponding graph structure is constructed according to the multi-scale sample feature graph: For sample feature graphs of different scales, each feature unit in the sample feature graph is used as a node in the graph structure; The relationship between the feature unit and other feature units is used as the edge between corresponding nodes, and the weight of the edge is calculated according to the foreground prior, local constraints and cosine similarity of the area where the feature unit is located.
[0012] In one embodiment, the sample graph structure of a certain scale is: ,in: ; In the above formula, and Indicates location and location The characteristic unit of represents the domain of all spatial coordinates, represents the weight function, represents the foreground prior of the SAR sample image, represents a local constraint, It represents cosine similarity, and the superscript S represents the parameter related to the sample graph structure.
[0013] In one embodiment, the step of calculating the vertex correlation loss at corresponding scales according to the sample graph structure and the template graph structure at the same scale respectively includes: For the sample graph structure and the corresponding feature units in the template graph structure at the same scale, a normalized correlation layer is used to obtain a correlation graph; A consistency mask is obtained using a geometric transformation matrix, and the vertex correlation loss is obtained according to the correlation graph and the consistency mask.
[0014] In one embodiment, the trained SAR target recognition network includes the first sub-feature extraction network and a classifier.
[0015] The present application also provides a SAR target recognition device based on graph structure consistency alignment, the device comprising: A data set acquisition module is used to acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; A training data set construction module is used to construct a corresponding structure template map according to a reference map of the target type for each SAR sample image in the sample data set, generate a set of training samples according to the associated structure template map, SAR sample image and reference map, and construct a training data set; A multi-scale feature extraction module, used for inputting a group of training samples in the training data set into a SAR target recognition network, in which a SAR sample image and a structure template map in a group of training samples are respectively segmented into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; A classification loss calculation module is used to classify the target according to the multi-scale sample feature map to obtain a predicted target category, and to calculate the classification loss according to the predicted target category and the corresponding target category label; A graph structure consistency loss calculation module is used to construct corresponding graph structures according to multi-scale sample feature graphs and multi-scale template feature graphs, respectively, to obtain multi-scale sample graph structures and multi-scale template graph structures, and to calculate vertex correlation loss, edge similarity difference loss and graph distance loss at corresponding scales according to the sample graph structure and template graph structure of the same scale; A SAR target recognition network training module is used to train the SAR target recognition network according to vertex correlation loss, edge similarity difference loss, graph distance loss and classification loss at different scales until a trained SAR target recognition network is obtained; The SAR image target recognition module is used to obtain a SAR image to be subjected to target recognition, and to perform target recognition on the SAR image using the trained SAR target recognition network.
[0016] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; For each SAR sample image in the sample data set, a corresponding structure template map is constructed according to a reference map of the target type therein, and a group of training samples are generated according to the associated structure template map, SAR sample image and reference map to construct a training data set; A group of training samples in the training data set is input into a SAR target recognition network. In the SAR target recognition network, a SAR sample image and a structure template map in a group of training samples are respectively divided into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; Performing target classification according to the multi-scale sample feature map to obtain a predicted target category, and performing calculation according to the predicted target category and the corresponding target category label to obtain a classification loss; The corresponding graph structures are constructed according to the multi-scale sample feature graph and the multi-scale template feature graph, respectively, and the multi-scale sample graph structure and the multi-scale template graph structure are obtained respectively. According to the sample graph structure and the template graph structure of the same scale, the vertex correlation loss, the edge similarity difference loss and the graph distance loss at the corresponding scale are calculated respectively; The SAR target recognition network is trained according to vertex correlation loss, edge similarity difference loss, graph distance loss, and classification loss at different scales until a trained SAR target recognition network is obtained; A SAR image to be subjected to target recognition is acquired, and target recognition is performed on the SAR image using the trained SAR target recognition network.
[0017] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps: Acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; For each SAR sample image in the sample data set, a corresponding structure template map is constructed according to a reference map of the target type therein, and a group of training samples are generated according to the associated structure template map, SAR sample image and reference map to construct a training data set; A group of training samples in the training data set is input into a SAR target recognition network. In the SAR target recognition network, a SAR sample image and a structure template map in a group of training samples are respectively divided into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; Performing target classification according to the multi-scale sample feature map to obtain a predicted target category, and performing calculation according to the predicted target category and the corresponding target category label to obtain a classification loss; The corresponding graph structures are constructed according to the multi-scale sample feature graph and the multi-scale template feature graph, respectively, and the multi-scale sample graph structure and the multi-scale template graph structure are obtained respectively. According to the sample graph structure and the template graph structure of the same scale, the vertex correlation loss, the edge similarity difference loss and the graph distance loss at the corresponding scale are calculated respectively; The SAR target recognition network is trained according to vertex correlation loss, edge similarity difference loss, graph distance loss, and classification loss at different scales until a trained SAR target recognition network is obtained; A SAR image to be subjected to target recognition is acquired, and target recognition is performed on the SAR image using the trained SAR target recognition network.
[0018] Beneficial effect: The above-mentioned SAR target recognition method and device based on graph structure consistency alignment, by constructing a structural template map according to the target reference map for each SAR sample image in the sample data set, generating training samples and constructing a training set, inputting it into the SAR target recognition network, dividing the sample image and the template map into blocks, reorganizing the multi-scale feature map according to the preset size, and then classifying the target based on the multi-scale sample feature map, calculating the classification loss, and then constructing a graph structure to calculate the vertex correlation, edge similarity difference, and spectrum distance loss. Use these losses to train the network until convergence, and use the trained network to identify the SAR image to be processed. The use of this method can effectively improve the accuracy of target recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic diagram of a flow chart of a SAR target recognition method based on graph structure consistency alignment in one embodiment; Figure 2 A schematic diagram of graph structure consistency constructed based on the similarity between SAR images and structure template graphs at different azimuth angles in one embodiment, wherein: Figure 2 (a) and Figure 2 (c) is the SAR image of the same aircraft at different azimuth angles. Figure 2 (b) is the structural template diagram; Figure 3 A schematic diagram of an aircraft structure reference template in one embodiment; Figure 4 A schematic diagram of SAR target recognition network training in one embodiment; Figure 5 A schematic diagram of encoding a structure template graph using a tokenizer in one embodiment; Figure 6 is a structural block diagram of a SAR target recognition device based on graph structure consistency alignment in one embodiment; Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0021] In view of the technical defects of the existing technology, and further considering that the imaging results of SAR aircraft are composed of a large number of discrete strong scattering points, which is different from the optical imaging mechanism, the learning process still appears chaotic and unconstrained despite the existence of supervisory signals. The key to the recognition algorithm is to achieve the continuous geometric invariant structure representation of the aircraft. Therefore, this application proposes two sub-problems that must be clearly solved: (1) How to establish a representation that is closely related to the aircraft structure from the existing advanced neural network. (2) How to maintain the geometric and topological consistency between these representations. Furthermore, if Figure 1 As shown, a SAR target recognition method based on graph structure consistency alignment is provided in the application, comprising the following steps: Step S100: obtaining a sample data set, where the sample data set includes a plurality of SAR sample images and corresponding target category labels.
[0022] Step S110, for each SAR sample image in the sample data set, a corresponding structure template map is constructed according to the reference map of the target type therein, and a group of training samples are generated according to the associated structure template map, SAR sample image and reference map to construct a training data set.
[0023] Step S120, input a group of training samples in the training data set into the SAR target recognition network. In the SAR target recognition network, the SAR sample images and the structure template images in the group of training samples are respectively divided into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map.
[0024] Step S130, performing target classification according to the multi-scale sample feature map to obtain a predicted target category, and performing calculation according to the predicted target category and the corresponding target category label to obtain a classification loss.
[0025] Step S140, construct corresponding graph structures according to the multi-scale sample feature graph and the multi-scale template feature graph, respectively, obtain the multi-scale sample graph structure and the multi-scale template graph structure, and calculate the vertex correlation loss, edge similarity difference loss and graph distance loss at the corresponding scale according to the sample graph structure and the template graph structure of the same scale.
[0026] Step S150, training the SAR target recognition network according to vertex correlation loss, edge similarity difference loss, graph distance loss and classification loss at different scales until a trained SAR target recognition network is obtained.
[0027] Step S160, obtaining a SAR image for target recognition, and using the trained SAR target recognition network to perform target recognition on the SAR image.
[0028] In this application, a method based on graph structure alignment is proposed to improve the accuracy of target recognition. In this method, a SAR image is segmented into multiple tokens using a neural network, and then the tokens are recombined according to different sizes to obtain a multi-scale feature map. The multi-scale feature map is then converted into a graph structure, and a loss function is used to ensure the structural consistency between the SAR image and the template image.
[0029] Specifically, for the first sub-problem raised above, in this method, a graph structure is introduced to explicitly establish the topological structure of the target. Graph structures have been widely used in real-world applications to process data linked by certain attributes, such as biology, social media, and finance. Graph structures can capture the inherent structural information of images and play an important role in visual tasks. Different from the traditional way of using features to build graph structures, self-similarity is applied to tokens extracted from neural networks in this method. Nodes and edges are defined by the self-similarity between tokens, and the proposed algorithm can be combined with the parameter model of end-to-end backpropagation to dynamically update the graph structure.
[0030] Specifically, for the second sub-problem raised above, this method proposes a graph structure alignment method to capture geometrically invariant representations, such as Figure 2 As shown in the figure, aircraft belonging to the same category but with different postures have consistent topological structures. Therefore, in this method, a template graph structure is set for each category target, which is constructed based on the tokens extracted from the target structure template. By iteratively limiting the distance between the graph structure generated by the SAR target image token and the graph structure generated by the target template image token, the structural consistency of the features can be guaranteed to be unaffected by posture changes. To this end, graph structure alignment losses are further designed at the vertex, edge, and graph levels. Vertex and edge distances can measure structural differences and graph distances can measure signal differences. By utilizing topological and graph characteristics, the structural differences between two tokens can be better captured.
[0031] In step S100, according to the specific application background, a SAR sample image of a relevant target is obtained. In this application, an aircraft is used as a target for illustration. In fact, the method in this application can also be applied to the recognition of other targets. The sample data set includes SAR sample images of multiple different types of aircraft and corresponding aircraft category targets.
[0032] Considering that the aircraft is a rigid object with a fixed structure, in this method, a template-based aircraft geometry annotation method is designed. In step S110, the corresponding target optical image is first generated in the high-resolution RSI obtained from Google Earth according to the target category in the SAR sample image, that is, the reference image, and then the target in the reference image is binarized and segmented to obtain a target binary image, that is, a structural template image.
[0033] Next, the structure template image is aligned according to the corresponding SAR sample image to obtain the rotated and translated structure template image, and then the rotated and translated structure template image is matched with the original structure template image to obtain the geometric transformation matrix of the target, and the relevant information of the target, including the target length and width information, is extracted based on the reference image. The SAR sample image, the corresponding target category label, the structure template image, the geometric transformation matrix and the relevant information are used as a set of training samples, and a corresponding set of training samples is generated according to each SAR sample image in the sample data set, thereby constructing the obtained training data set.
[0034] like Figure 3 As shown, it is a schematic diagram of a reference image, a structural template diagram and related information with an aircraft as the target, wherein the first row is the optical image, i.e., the reference image, the second row is the structural template diagram, and the third and fourth rows are the length and width information of the aircraft.
[0035] Specifically, the geometric transformation matrix of the target is obtained by matching the structure template image with the corresponding SAR sample image. , expressed as: ; Furthermore, Convert to , the mapping function used is expressed as: ; Different from other existing aircraft annotation methods, including point-based, polygon-based, pose-based and component-based annotation methods, the introduced template-based annotation method has the following advantages: first, it avoids fine annotation of pixels or boundaries, thus ensuring the convenience of operation; second, it can derive information equivalent to that provided by other annotation methods, for example, the pose can be obtained from In The foreground and background interpretation of a pixel can be obtained from the template at its transformed position according to Therefore, for each SAR aircraft image (SAR sample image) , this scheme provides a corresponding structure template diagram and geometric transformations Additional information included.
[0036] In step S120, a set of training samples in the training data set is input into the SAR target recognition network, and the SAR target recognition network is trained. When training the SAR target recognition network, Figure 4 As shown, the SAR target recognition network includes a first sub-feature extraction network, a second sub-feature extraction network and a classifier, wherein the first sub-feature extraction network is used to perform multi-scale feature extraction on the SAR sample image to obtain a multi-scale sample feature map. The second sub-feature extraction network is used to perform multi-scale feature extraction on the structure template map to obtain a multi-scale template feature map. The classifier includes a global average pooling layer, a fully connected layer and a Softmax function layer connected in sequence, which are used to predict the target category according to the multi-scale sample feature map to obtain the predicted target category.
[0037] Specifically, the first sub-feature extraction network adopts the Swin Transformer network structure, and the second sub-feature extraction network adopts the Tokenizer network structure.
[0038] In this embodiment, the Swin Transformer network is first used as a feature extraction network for SAR sample images. In the Swin Transformer, the size is The image is first divided into non-overlapping blocks, each of which is regarded as a token. After a series of Swin Transformer blocks and Patching Merge blocks, the cascaded input tokens are mapped to tokens of various sizes. In this method, the re-divided tokens are also represented by feature units, that is, multiple feature maps with different feature unit scales are obtained. It should be noted that in the same feature map, the size of the feature unit is consistent.
[0039] Specifically, in the Swin Transformer network, after the SAR sample image is initially divided into multiple non-overlapping tokens, multiple tokens are reorganized according to multiple preset sizes to obtain new tokens, namely feature units. In this way, a SAR sample image can generate sample feature maps of multiple scales under feature units of different sizes.
[0040] In one embodiment, the SAR sample image is passed through the Swin Transformer network to obtain sample feature maps at three scales, the scales of which include , and .
[0041] In step S130, a classifier is connected after the Swin Transformer network. Among them, the global average pooling (GAP), the fully connected (FC) layer and the Softmax function are applied to the last token to obtain the classification probability. In terms of classification loss, the cross entropy (CE) is used to evaluate the difference between the prediction and the true label, as shown below: ; In the above formula, Indicates the number of categories.
[0042] In this embodiment, an independent second sub-feature extraction network, namely the Tokenizer network, is also designed to encode the structure template map into a template feature map of the same size as the sample feature map. Figure 5 As shown, two examples are shown, namely, The template image is converted to a size of and The token is the multi-scale template map.
[0043] In one embodiment, the downsampling rate in the Tokenizer network is Set to 8, 16, and 32. In particular, for each downsampling, the tokenizer pixels are sampled and the structural template graph is divided into Then, these feature maps are stacked into a scale of , the dimension is The feature map of . Then a 3×3 convolutional layer is applied to align the token dimension with the output of the Swin Transformer. It can be observed that the designed tokenizer does not lose any spatial structure information, that is, the position relationship of the pixels does not change before and after encoding.
[0044] In step S140, when constructing corresponding graph structures according to the multi-scale sample feature map and the multi-scale template feature map respectively, wherein, when constructing the corresponding graph structure according to the multi-scale sample feature map: for sample feature maps of different scales, each feature unit in the sample feature map is used as a node in the graph structure, and the relationship between the feature unit and other feature units is used as the edge between the corresponding nodes, and the weight of the edge is calculated based on the foreground prior, local constraints and cosine similarity of the area where the feature unit is located.
[0045] In this example, a series of tokens extracted from SAR images are ,The graph structure can effectively capture the key information and local structure of the image.,Therefore, in this method, a corresponding graph structure is constructed for each scale of,sample feature graph.
[0046] Specifically, the sample graph structure of a certain scale is: ,in: ; In the above formula, and Indicates location and location The characteristic unit of represents the domain of all spatial coordinates, represents the weight function, and Respectively represent the foreground priors of the SAR sample image and the structure template image, satisfying as well as , represents a local constraint, It represents cosine similarity, and the superscript S represents the parameter related to the sample graph structure.
[0047] Among them, local constraints It is expressed as: ; Among them, cosine similarity It is expressed as: .
[0048] In this embodiment, the local constraint is used to highlight the importance of edges with close distances between the start and end points, which is crucial for inheriting the intrinsic distance property of the Euclidean space in the original feature graph.
[0049] Similarly, in the graph structure, the aircraft structure information can be effectively represented at the token level. In order to explore the differences in the graphs, a similar graph structure is also constructed from the template feature graphs of each scale, represented as: ; Furthermore, the consistency of the graph structure is determined by measuring the distance between the multi-scale sample graph structure and the multi-scale template graph structure. As we know, the graph structure information can be found in its topology and spectrum. In particular, the edge weight determines the amount of information diffused from a vertex to its neighbors, and the maximum spectrum can be used to detect energy changes caused by abnormal vertices. Using these two properties, the potential graph structure differences can be measured. Therefore, a graph structure alignment method with three different aspects is proposed, such as Figure 4 As shown in the three sub-figures in the middle and lower part, three loss functions are used to ensure the consistency of the target structure between the SAR sample image and the structural template image during iterative training.
[0050] In this embodiment, vertex correlation loss is first introduced. The vertex correlation loss at the corresponding scale is calculated according to the sample graph structure and the template graph structure at the same scale, including: for the corresponding feature units in the sample graph structure and the template graph structure at the same scale, a normalized correlation layer is used to obtain a correlation graph, a consistent mask is obtained using a geometric transformation matrix, and the vertex correlation loss is obtained according to the correlation graph and the consistent mask.
[0051] Specifically, yes Use the normalized correlation layer to get the correlation graph , calculated as follows: ; In the above formula, Middle position Token and Middle position The normalized inner product of the token. The final result contains the one-way correlations of all feature pairs in the SAR image and the template.
[0052] A similar correlation plot in the opposite direction is: ; Given a geometric transformation matrix , it is possible to confirm the presence of geometrically consistent feature pairs and eliminate false matches that do not conform to the geometric correspondence. This is achieved through two consistency masks, where the two consistency masks can be expressed as: ; ; In the above formula, It is the L2 distance function with a preset threshold 1. This definition is valid only when the projection error is no greater than Then, by multiplying the correlation graph and the consistency mask and summing them, we can get the loss of GSA at the vertex level, that is, the vertex correlation loss, which is expressed as: ; The training objective is to maximize the correlation of vertex pairs with the geometric consensus. Note that due to the normalized correlation operation, it also penalizes ambiguous correlations with inconsistent correspondences.
[0053] In this embodiment, since the weighted edges contain rich topological structural relationships, the edge similarity difference can be directly used to measure the difference of the graph, which is calculated as: ; In order to maintain structural consistency, it is necessary to narrow the differences in semantically coherent edges in the graph constructed from the SAR image and the template. A “semantically coherent edge” refers to an edge whose start and end points of the two domains are geometrically consistent with the transformation. Consistent edge pairs, such as left wing to left wing, rather than left wing to right wing. A simple and intuitive approach is to introduce an edge correspondence mask, as follows: ; However, the loss at the edge level, namely the edge similarity difference, is expressed as: ; Although this definition is intuitive, the space complexity of computing the sum of all edge pairs over two domains is Assume that the graph has 100 vertices and the edge similarity difference matrix is and the consistent mask matrix The dimension is , which is difficult to calculate and store. Further observation shows that is a sparse matrix, that is, The elements contain only The element is 1.
[0054] Therefore, instead of comparing the edge differences between the two graphs directly, The geometric transformation is Isomorphic graph of , expressed as: ; Finally, the loss of the graph structure at the edge level is expressed as: ; In the above formula, it is equivalent to directly The sum of the non-zero terms in , and reduces the space complexity to .
[0055] In this embodiment, the Laplacian of FIG is a linear translation-invariant signal filter. , graph filters generally include the standard Laplacian matrix , the symmetric normalized Laplacian matrix And the random walk normalized Laplacian matrix ,in is a degree matrix satisfying , is the edge matrix. The graph distance can quantitatively measure the number of identical nodes between two nodes. The structural differences between isomorphic graphs.
[0056] and The spectral representation of and , the eigenvalue is expressed as: ; ; However, the loss of GSA at the spectral level can be expressed as the distance between two spectral vectors: ; based on , the topological structures between the SAR image and the tokens extracted from the template can be aligned to the same atlas and finally have the same atlas properties.
[0057] In the implementation process, the two-way feature extractor produces features of different scales, namely , and Then a multi-scale graph is constructed from these features, which is represented as and ,in This is able to capture the topological structure between aircraft components of different sizes. Therefore, the final optimization objective function is: ; In the above formula, , and represents the weight term of the loss, , and Indicates that from the figure and The GSA loss obtained in .
[0058] In step 150, under the training of the final optimization objective function, the SAR target recognition network obtained after training convergence. The trained SAR target recognition network includes the first sub-feature extraction network and the classifier, and the second sub-feature extraction network is removed. In step S160, when the trained SAR target recognition network is actually applied, it is only necessary to input the SAR image to be detected into the SAR target recognition network to achieve target recognition detection.
[0059] As shown in Algorithm 1 in Table 1, this is the training process of the SAR target recognition network (GSA-Trans). The GSA loss provides spatial regularization for identifying structural class relationships, effectively improving the discriminability of features. In the test phase, only the feature extraction and classifier of the SAR image are retained to generate the predicted label. Therefore, it is still an end-to-end feature; Table 1
[0060] In the above-mentioned SAR target recognition method based on graph structure consistency alignment, a structural template map is constructed according to the target reference map for each SAR sample image in the sample data set, and training samples are generated and a training set is constructed. The sample image and the template map are divided into blocks, and a multi-scale feature map is reorganized according to a preset size. Then, the target is classified based on the multi-scale sample feature map, and the classification loss is calculated. Then, the graph structure is constructed to calculate the vertex correlation, edge similarity difference, and graph distance loss. The network is trained with these losses until convergence, and the trained network is used to recognize the processed SAR image. This method proposes a graph structure construction method, which converts the feature map extracted by the network into a graph structure closely related to the target structure, and measures the graph differences at the vertices, edges, and graphs of the graph structure, respectively, and designs a graph alignment network loss function for the structural consistency feature. The superiority of the proposed method is verified by experiments on public and self-built SAR aircraft recognition datasets.
[0061] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0062] In one embodiment, Figure 6 As shown, a SAR target recognition device based on graph structure consistency alignment is provided, comprising: a data set acquisition module 200, a training data set construction module 210, a multi-scale feature extraction module 220, a classification loss calculation module 230, a graph structure consistency loss calculation module 240, a SAR target recognition network training module 250 and a SAR image target recognition module 260, wherein: The data set acquisition module 200 is used to acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; The training data set construction module 210 is used to construct a corresponding structure template map according to a reference map of the target type for each SAR sample image in the sample data set, generate a set of training samples according to the associated structure template map, SAR sample image and reference map, and construct a training data set; A multi-scale feature extraction module 220 is used to input a group of training samples in the training data set into a SAR target recognition network, in which a SAR sample image and a structure template map in a group of training samples are respectively divided into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; A classification loss calculation module 230 is used to perform target classification according to the multi-scale sample feature map to obtain a predicted target category, and to perform calculation according to the predicted target category and the corresponding target category label to obtain a classification loss; A graph structure consistency loss calculation module 240 is used to construct corresponding graph structures according to the multi-scale sample feature graph and the multi-scale template feature graph, respectively, to obtain the multi-scale sample graph structure and the multi-scale template graph structure, and to calculate the vertex correlation loss, edge similarity difference loss and graph distance loss at the corresponding scales according to the sample graph structure and the template graph structure of the same scale; A SAR target recognition network training module 250 is used to train the SAR target recognition network according to vertex correlation loss, edge similarity difference loss, graph distance loss and classification loss at different scales until a trained SAR target recognition network is obtained; The SAR image target recognition module 260 is used to obtain a SAR image to be subjected to target recognition, and to perform target recognition on the SAR image using the trained SAR target recognition network.
[0063] For the specific definition of the SAR target recognition device based on graph structure consistency alignment, please refer to the definition of the SAR target recognition method based on graph structure consistency alignment above, which will not be repeated here. Each module in the above-mentioned SAR target recognition device based on graph structure consistency alignment can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0064] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a SAR target recognition method based on graph structure consistency alignment is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0065] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0066] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented: Acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; For each SAR sample image in the sample data set, a corresponding structure template map is constructed according to a reference map of the target type therein, and a group of training samples are generated according to the associated structure template map, SAR sample image and reference map to construct a training data set; A group of training samples in the training data set is input into a SAR target recognition network. In the SAR target recognition network, a SAR sample image and a structure template map in a group of training samples are respectively divided into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; Performing target classification according to the multi-scale sample feature map to obtain a predicted target category, and performing calculation according to the predicted target category and the corresponding target category label to obtain a classification loss; The corresponding graph structures are constructed according to the multi-scale sample feature graph and the multi-scale template feature graph, respectively, and the multi-scale sample graph structure and the multi-scale template graph structure are obtained respectively. According to the sample graph structure and the template graph structure of the same scale, the vertex correlation loss, the edge similarity difference loss and the graph distance loss at the corresponding scale are calculated respectively; The SAR target recognition network is trained according to vertex correlation loss, edge similarity difference loss, graph distance loss, and classification loss at different scales until a trained SAR target recognition network is obtained; A SAR image to be subjected to target recognition is acquired, and target recognition is performed on the SAR image using the trained SAR target recognition network.
[0067] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented: Acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; For each SAR sample image in the sample data set, a corresponding structure template map is constructed according to a reference map of the target type therein, and a group of training samples are generated according to the associated structure template map, SAR sample image and reference map to construct a training data set; A group of training samples in the training data set is input into a SAR target recognition network. In the SAR target recognition network, a SAR sample image and a structure template map in a group of training samples are respectively divided into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; Performing target classification according to the multi-scale sample feature map to obtain a predicted target category, and performing calculation according to the predicted target category and the corresponding target category label to obtain a classification loss; The corresponding graph structures are constructed according to the multi-scale sample feature graph and the multi-scale template feature graph, respectively, and the multi-scale sample graph structure and the multi-scale template graph structure are obtained respectively. According to the sample graph structure and the template graph structure of the same scale, the vertex correlation loss, the edge similarity difference loss and the graph distance loss at the corresponding scale are calculated respectively; The SAR target recognition network is trained according to vertex correlation loss, edge similarity difference loss, graph distance loss, and classification loss at different scales until a trained SAR target recognition network is obtained; A SAR image to be subjected to target recognition is acquired, and target recognition is performed on the SAR image using the trained SAR target recognition network.
[0068] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0069] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0070] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A SAR target recognition method based on graph structure consistency alignment, characterized in that: The method comprises: Acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; For each SAR sample image in the sample data set, a corresponding structure template map is constructed according to a reference map of the target type therein, and a group of training samples are generated according to the associated structure template map, SAR sample image and reference map to construct a training data set; A group of training samples in the training data set is input into a SAR target recognition network. In the SAR target recognition network, a SAR sample image and a structure template map in a group of training samples are respectively divided into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; Performing target classification according to the multi-scale sample feature map to obtain a predicted target category, and performing calculation according to the predicted target category and the corresponding target category label to obtain a classification loss; The corresponding graph structures are constructed according to the multi-scale sample feature graph and the multi-scale template feature graph, respectively, and the multi-scale sample graph structure and the multi-scale template graph structure are obtained respectively. According to the sample graph structure and the template graph structure of the same scale, the vertex correlation loss, the edge similarity difference loss and the graph distance loss at the corresponding scale are calculated respectively; The SAR target recognition network is trained according to vertex correlation loss, edge similarity difference loss, graph distance loss, and classification loss at different scales until a trained SAR target recognition network is obtained; A SAR image to be subjected to target recognition is acquired, and target recognition is performed on the SAR image using the trained SAR target recognition network.
2. The SAR target recognition method based on graph structure consistency alignment according to claim 1 is characterized in that: The step of generating a group of training samples according to the associated structure template image, SAR sample image and reference image comprises: the reference image is a target optical image, and the structure template image is a binary image obtained based on the target optical image; Aligning the structure template image according to the SAR sample image to obtain a rotated and translated structure template image; Then, the rotated and translated structure template image is matched with the original structure template image to obtain the geometric transformation matrix of the target, and relevant information of the target is extracted based on the reference image; The SAR sample image, the corresponding target category label, the structure template map, the geometric transformation matrix and related information are used as a group of training samples.
3. The SAR target recognition method based on graph structure consistency alignment according to claim 2 is characterized in that: When training the SAR target recognition network, the SAR target recognition network includes a first sub-feature extraction network, a second sub-feature extraction network and a classifier; The first sub-feature extraction network is used to perform multi-scale feature extraction on the SAR sample image to obtain the multi-scale sample feature map; The second sub-feature extraction network is used to perform multi-scale feature extraction on the structure template map to obtain the multi-scale template feature map; The classifier includes a global average pooling layer, a fully connected layer and a Softmax function layer connected in sequence, and is used to predict the target category according to the multi-scale sample feature map to obtain the predicted target category.
4. The SAR target recognition method based on graph structure consistency alignment according to claim 3 is characterized in that: The first sub-feature extraction network adopts a Swin Transformer network structure, and the second sub-feature extraction network adopts a Tokenizer network structure.
5. The SAR target recognition method based on graph structure consistency alignment according to any one of claims 2 to 4, characterized in that: When corresponding graph structures are constructed according to the multi-scale sample feature graph and the multi-scale template feature graph respectively, wherein, when corresponding graph structures are constructed according to the multi-scale sample feature graph: For sample feature graphs of different scales, each feature unit in the sample feature graph is used as a node in the graph structure; The relationship between the feature unit and other feature units is used as the edge between corresponding nodes, and the weight of the edge is calculated according to the foreground prior, local constraints and cosine similarity of the area where the feature unit is located.
6. The SAR target recognition method based on graph structure consistency alignment according to claim 5 is characterized in that: The sample graph structure of a certain scale is: ,in: In the above formula, and Indicates location and location The characteristic unit of represents the domain of all spatial coordinates, represents the weight function, represents the foreground prior of the SAR sample image, represents a local constraint, It represents cosine similarity, and the superscript S represents a parameter related to the sample graph structure.
7. The SAR target recognition method based on graph structure consistency alignment according to claim 6 is characterized in that: The step of calculating the vertex correlation loss at corresponding scales according to the sample graph structure and the template graph structure of the same scale includes: For the sample graph structure and the corresponding feature units in the template graph structure at the same scale, a normalized correlation layer is used to obtain a correlation graph; A consistency mask is obtained using a geometric transformation matrix, and the vertex correlation loss is obtained according to the correlation graph and the consistency mask.
8. The SAR target recognition method based on graph structure consistency alignment according to claim 7 is characterized in that: The trained SAR target recognition network includes the first sub-feature extraction network and a classifier.
9. A SAR target recognition device based on graph structure consistency alignment, characterized in that: The device comprises: A data set acquisition module is used to acquire a sample data set, wherein the sample data set includes a plurality of SAR sample images and corresponding target category labels; A training data set construction module is used to construct a corresponding structure template map according to a reference map of the target type for each SAR sample image in the sample data set, generate a set of training samples according to the associated structure template map, SAR sample image and reference map, and construct a training data set; A multi-scale feature extraction module, used for inputting a group of training samples in the training data set into a SAR target recognition network, in which a SAR sample image and a structure template map in a group of training samples are respectively segmented into a plurality of non-overlapping image blocks of the same size, and the image blocks are regrouped according to a plurality of preset feature unit sizes to obtain a multi-scale feature map including a plurality of feature units of the same size, which are a multi-scale sample feature map and a multi-scale template feature map; A classification loss calculation module is used to classify the target according to the multi-scale sample feature map to obtain a predicted target category, and to calculate the classification loss according to the predicted target category and the corresponding target category label; A graph structure consistency loss calculation module is used to construct corresponding graph structures according to multi-scale sample feature graphs and multi-scale template feature graphs, respectively, to obtain multi-scale sample graph structures and multi-scale template graph structures, and to calculate vertex correlation loss, edge similarity difference loss and graph distance loss at corresponding scales according to the sample graph structure and template graph structure of the same scale; A SAR target recognition network training module is used to train the SAR target recognition network according to vertex correlation loss, edge similarity difference loss, graph distance loss and classification loss at different scales until a trained SAR target recognition network is obtained; The SAR image target recognition module is used to obtain a SAR image to be subjected to target recognition, and to perform target recognition on the SAR image using the trained SAR target recognition network.
Citation Information
Patent Citations
Stochastic gradient Bayesian SAR image segmentation method based on sketch structure
CN106611422A
SAR image target identification method and device
CN110110625A
Radar high-resolution range profile recognition method based on multi-scale grouping fusion convolution
CN113988163A
Ground surface abnormity early warning text generation method based on expression knowledge graph
CN117708346A
SAR (Synthetic Aperture Radar) image target identification method and device, computer equipment and storage medium
CN117710728A