Cross-sensor SAR (Synthetic Aperture Radar) image target detection method, device, equipment and medium

By building a cross-domain object detection network, using semantic scattering map structure and global feature alignment loss function, the problem of poor generalization ability of SAR object detection under cross-sensor conditions is solved, and target accurate detection under different imaging conditions is achieved.

CN120071024AActive Publication Date: 2025-05-30NAT UNIV OF DEFENSE TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510530010.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing SAR object detection methods have poor generalization capabilities under cross-sensor conditions, resulting in a significant decline in the performance of models trained under one imaging condition on data under other conditions.

Method used

By obtaining the data sets of the source domain and the target domain, the semantic feature map is extracted and the semantic scatter graph structure is constructed. The parameters in the multi-layer shared feature extractor and detection head are adjusted by using the scatter structure diagram alignment loss function, global feature alignment loss function and detection loss function to build a cross-domain target detection network.

Benefits of technology

It realizes that target accurate detection on the target domain under the training of the target domain without truth-value tags, and improves the generalization ability of the model under cross-sensor conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071024A_ABST
    Figure CN120071024A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-sensor SAR (Synthetic Aperture Radar) image target detection method, device, equipment and medium, and the method comprises the steps: inputting an SAR source domain image with a truth value label and an SAR target domain image without a truth value label into a multi-layer shared feature extractor to obtain corresponding semantic feature maps, and constructing a corresponding map structure based on the feature maps; according to the method, the consistency of a source domain semantic graph structure and a target domain semantic graph structure is constrained during training by constructing a loss function, so that a trained network has good generalization ability on a target domain, and meanwhile, a target detection constraint and a global feature constraint are combined to obtain a target domain image capable of being trained under the training of a target domain image without a truth value label. And target accurate detection of the target domain SAR image is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of SAR target detection, and particularly to a cross-sensor SAR image target detection method, device, equipment, and medium. Background Art

[0002] Synthetic Aperture Radar (SAR) has the ability to work regardless of day or night and is not affected by weather. It captures target information by actively emitting electromagnetic waves and has been widely used in the field of earth observation. In civilian fields such as marine monitoring, disaster assessment, and resource management, SAR plays an indispensable role.

[0003] Traditional SAR target detection methods mainly rely on the physical characteristics of targets, such as geometric shape and polarization characteristics. Among them, the Constant False Alarm Rate (CFAR)-based algorithm is a commonly used target detection method, which identifies targets by analyzing the distribution difference between targets and background clutter. However, the detection accuracy of this method is limited and it is prone to a high false alarm rate. In recent years, deep learning technology has made remarkable progress in the field of remote sensing image analysis. Deep neural networks have gradually been applied to SAR target detection and recognition tasks and have achieved breakthrough results in fields such as aircraft, ships, and vehicles, demonstrating excellent performance.

[0004] However, most current SAR target detection studies assume that the training data and test data come from the same distribution, that is, data obtained from the same sensor. With the continuous increase in SAR imaging platforms, in practical applications, the training data and test data often come from different sensors, resulting in differences in data distribution. This difference mainly stems from the differences in technical indicators and system parameters of different SAR platforms, such as radar frequency, beam width, orbital altitude, and observation angle. In addition, SAR data under different imaging conditions have different probability distributions, and there are also significant differences in image styles and target characteristics. Therefore, when a model trained under a specific imaging condition is applied to data under other conditions, its performance will drop significantly, indicating that the generalization ability of SAR target detection under cross-sensor conditions is poor.

[0005] A common solution is to annotate the data under the new imaging condition and then retrain the model. However, since the interpretation of SAR images requires professional knowledge, the annotation process is both time-consuming and laborious, and it is difficult to achieve efficient and accurate annotation. Therefore, how to enable a detection model trained under one imaging condition to successfully generalize to data under other conditions has become an urgent problem to be solved. Summary of the Invention

[0006] Based on this, it is necessary to provide a cross-sensor SAR image target detection method, device, equipment, and medium that can achieve cross-domain precise detection for the above-mentioned technical problems.

[0007] A cross-sensor SAR image target detection method, the method comprising: Obtain a source domain dataset and a target domain dataset, where the source domain dataset includes multiple SAR source domain images with ground truth labels, and the target domain dataset includes multiple SAR target domain images without ground truth labels; Extract one SAR image from each of the source domain dataset and the target domain dataset as a training pair, input the training pair into a multi-layer shared feature extractor to obtain a source domain semantic feature map and a target domain semantic feature map respectively, and construct corresponding semantic scattering map structures according to the source domain semantic feature map and the target domain semantic feature map, obtaining a source domain semantic scattering map structure and a target domain semantic scattering map structure; Construct a scattering map structure alignment loss function according to the source domain semantic scattering map structure and the target domain semantic scattering map structure; Use a discriminator to construct a multi-level global feature alignment loss function according to the source domain semantic feature map and the target domain semantic feature map; Use a detection head to perform target detection according to the source domain semantic feature map and the target domain semantic feature map respectively, and construct a detection loss function; Use the scattering map structure alignment loss function, the global feature alignment loss function, and the detection loss function to adjust the tunable parameters in the multi-layer shared feature extractor and the detection head until convergence, and construct a cross-domain target detection network according to the trained multi-layer shared feature extractor and the detection head; Obtain a SAR image to be detected in the target domain, and use the cross-domain target detection network to perform target detection on the SAR image to be detected, obtaining the target category and the target location.

[0008] In one embodiment, the constructing corresponding semantic scattering map structures according to the source domain semantic feature map and the target domain semantic feature map includes: For the source domain semantic feature map, use the ground truth boxes in the corresponding ground truth labels to perform spatially uniform sampling of pixel points at different feature layers, obtaining a source domain semantic scattering point set; For the target domain semantic feature map, use the detection head to generate the class scores of each pixel point, and based on the class scores, obtain a target source domain semantic scattering point set; Map the feature points in the source domain semantic scattering point set and the target source domain semantic scattering point set to a low-dimensional vector space respectively, convert them into graph nodes, and construct the corresponding semantic scattering map structures.

[0009] In one embodiment, constructing the scattering structure graph alignment loss function according to the source domain semantic scattering graph structure and the target domain semantic scattering graph structure includes: Calculating the semantic feature point alignment loss function according to the source domain semantic scattering graph structure and the target domain semantic scattering graph structure; For each graph node in the source domain semantic scattering graph structure and the target domain semantic scattering graph structure respectively, using the feature distribution of the graph node and the nearest neighbor algorithm for node enhancement to obtain the source domain node enhanced graph structure and the target domain node enhanced graph structure; After interacting the source domain node enhanced graph structure and the target domain node enhanced graph structure, calculating the semantic scattering point classification loss function; Constructing the semantic similarity graph structure matching loss function according to the maximum node similarity, minimum node similarity and edge similarity of the matching node pairs in the source domain node enhanced graph structure and the target domain node enhanced graph structure; Obtaining the scattering structure graph alignment loss function according to the semantic feature point alignment loss function, the semantic scattering point classification loss function and the semantic similarity graph structure matching loss function.

[0010] In one embodiment, the semantic feature point alignment loss function is expressed as: ; In the above formula, represents the node alignment discriminator network, represents the th graph node in the source domain semantic scattering graph structure, represents the th graph node in the target domain semantic scattering graph structure, represents the domain classification label, 、 respectively represent the number of graph nodes in the source domain semantic scattering graph structure and the target domain semantic scattering graph structure.

[0011] In one embodiment, when enhancing each graph node in the source domain semantic scattering graph structure and the target domain semantic scattering graph structure: Model the distribution of graph nodes in the semantic scattering graph structure as a Gaussian distribution, sample feature points from this Gaussian distribution, and then obtain the enhanced semantic scattering nodes according to the linear projection formula, and construct the first enhanced semantic scattering graph node set; Obtain the second enhanced semantic scattering graph node set by taking the union of the graph nodes in the original semantic scattering graph node set and the first enhanced semantic scattering graph node set; An adjacency matrix is obtained by multiplying the nodes of the original semantic scattering graph set through inner product. A graph convolutional network is used to enhance each graph node in the second enhanced semantic scattering graph node set based on the features of neighboring nodes according to the adjacency matrix, and a third enhanced semantic scattering graph node set is obtained. The graph nodes in the third enhanced semantic scattering graph node set are clustered using the K-nearest neighbor algorithm to obtain the final node-enhanced graph structure.

[0012] In one embodiment, before constructing the first enhanced semantic scattering graph node set, for the target domain semantic scattering graph structure, the features of the graph nodes in the target domain semantic scattering graph structure are enhanced using the distribution variance of the graph structure in the target domain semantic scattering graph structure and the mean of the source domain semantic scattering graph structure.

[0013] The present application also provides a cross-sensor SAR image target detection device, which includes: A dataset acquisition module for acquiring a source domain dataset and a target domain dataset. The source domain dataset includes multiple SAR source domain images with ground truth labels, and the target domain dataset includes multiple SAR target domain images without ground truth labels. A graph structure construction module for respectively extracting a SAR image from the source domain dataset and the target domain dataset as a training pair, inputting the training pair into a multi-layer shared feature extractor to obtain a source domain semantic feature map and a target domain semantic feature map respectively, and constructing corresponding semantic scattering graph structures according to the source domain semantic feature map and the target domain semantic feature map to obtain a source domain semantic scattering graph structure and a target domain semantic scattering graph structure. A graph structure alignment constraint module for constructing a scattering graph structure alignment loss function according to the source domain semantic scattering graph structure and the target domain semantic scattering graph structure. A global feature alignment constraint module for using a discriminator to construct a multi-level global feature alignment loss function according to the source domain semantic feature map and the target domain semantic feature map. A target detection constraint module for using a detection head to perform target detection according to the source domain semantic feature map and the target domain semantic feature map respectively, and constructing a detection loss function. A training module for adjusting the adjustable parameters in the multi-layer shared feature extractor and the detection head using the scattering graph structure alignment loss function, the global feature alignment loss function, and the detection loss function until convergence, and constructing a cross-domain target detection network according to the trained multi-layer shared feature extractor and the detection head. A cross-domain detection module for acquiring a SAR image to be detected in the target domain, and using the cross-domain target detection network to perform target detection on the SAR image to be detected to obtain the target category and the target location.

[0014] A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Obtain a source domain dataset and a target domain dataset, where the source domain dataset includes multiple SAR source domain images with ground truth labels, and the target domain dataset includes multiple SAR target domain images without ground truth labels; Extract one SAR image from the source domain dataset and the target domain dataset respectively as a training pair, input the training pair into a multi-layer shared feature extractor to obtain a source domain semantic feature map and a target domain semantic feature map respectively, and construct corresponding semantic scattering map structures according to the source domain semantic feature map and the target domain semantic feature map respectively, to obtain a source domain semantic scattering map structure and a target domain semantic scattering map structure; Construct a scattering structure map alignment loss function according to the source domain semantic scattering map structure and the target domain semantic scattering map structure; Use a discriminator to construct a multi-level global feature alignment loss function according to the source domain semantic feature map and the target domain semantic feature map; Use a detection head to perform object detection according to the source domain semantic feature map and the target domain semantic feature map respectively, and construct a detection loss function; Use the scattering structure map alignment loss function, the global feature alignment loss function, and the detection loss function to adjust the tunable parameters in the multi-layer shared feature extractor and the detection head until convergence, and construct a cross-domain object detection network according to the trained multi-layer shared feature extractor and the detection head; Obtain a SAR image to be detected in the target domain, and use the cross-domain object detection network to perform object detection on the SAR image to be detected, to obtain the object category and the object location.

[0015] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented: Obtain a source domain dataset and a target domain dataset, where the source domain dataset includes multiple SAR source domain images with ground truth labels, and the target domain dataset includes multiple SAR target domain images without ground truth labels; Extract one SAR image from the source domain dataset and the target domain dataset respectively as a training pair, input the training pair into a multi-layer shared feature extractor to obtain a source domain semantic feature map and a target domain semantic feature map respectively, and construct corresponding semantic scattering map structures according to the source domain semantic feature map and the target domain semantic feature map respectively, to obtain a source domain semantic scattering map structure and a target domain semantic scattering map structure; Construct a scattering structure graph alignment loss function according to the source domain semantic scattering graph structure and the target domain semantic scattering graph structure; Use a discriminator to construct a multi-level global feature alignment loss function according to the source domain semantic feature map and the target domain semantic feature map; Use a detection head to perform object detection according to the source domain semantic feature map and the target domain semantic feature map respectively, and construct a detection loss function; Use the scattering structure graph pair loss function, the global feature alignment loss function, and the detection loss function to adjust the adjustable parameters in the multi-level shared feature extractor and the detection head until convergence, and construct a cross-domain object detection network according to the trained multi-level shared feature extractor and the detection head; Obtain the SAR image to be detected in the target domain, and use the cross-domain object detection network to perform object detection on the SAR image to be detected to obtain the object category and the object location.

[0016] The above cross-sensor SAR image object detection method, device, equipment and medium input the SAR source domain image with a true value label and the SAR target domain image without a true value label into a multi-level shared feature extractor to obtain corresponding semantic feature maps respectively, and construct corresponding graph structures based on the feature maps. By constructing a loss function to constrain the consistency of the source domain semantic graph structure and the target domain semantic graph structure during training, the trained network has good generalization ability in the target domain. At the same time, combined with object detection constraints and global feature constraints, accurate object detection of the target domain SAR image can be achieved under the training of the target domain image without a true value label. Description of the Drawings

[0017] Figure 1 It is a schematic flowchart of a cross-sensor SAR image object detection method in an embodiment; Figure 2 It is a schematic flowchart of semantic scattering graph construction in an embodiment; Figure 3 It is a schematic diagram of semantic scattering graph node enhancement in an embodiment; Figure 4 It is a schematic diagram of semantic scattering graph perceptual association in another embodiment; Figure 5 It is a schematic diagram of semantic scattering structure graph alignment in another embodiment; Figure 6 It is a structural block diagram of a cross-sensor SAR image object detection device in an embodiment; Figure 7 It is an internal structural diagram of a computer device in an embodiment. Detailed Embodiments

[0018] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0019] Regarding the existing cross-sensor SAR image detection methods, data obtained under new imaging conditions (i.e., another sensor) are required for annotation, and then the model needs to be retrained. However, due to the fact that the interpretation of SAR images requires professional knowledge, the annotation process is both time-consuming and laborious, making it difficult to achieve efficient and accurate annotation. In this application, as Figure 1 shown, a cross-sensor SAR image target detection method is provided, which specifically includes the following steps: Step S100, obtain a source domain dataset and a target domain dataset. The source domain dataset includes multiple SAR source domain images with ground truth labels, and the target domain dataset includes multiple SAR target domain images without ground truth labels.

[0020] Step S110, respectively extract one SAR image from the source domain dataset and the target domain dataset as a training pair, input the training pair into a multi-layer shared feature extractor to obtain a source domain semantic feature map and a target domain semantic feature map respectively, and construct corresponding semantic scattering map structures according to the source domain semantic feature map and the target domain semantic feature map respectively, to obtain a source domain semantic scattering map structure and a target domain semantic scattering map structure.

[0021] Step S120, construct a scattering structure map alignment loss function according to the source domain semantic scattering map structure and the target domain semantic scattering map structure.

[0022] Step S130, use a discriminator to construct a multi-level global feature alignment loss function according to the source domain semantic feature map and the target domain semantic feature map.

[0023] Step S140, use a detection head to perform target detection according to the source domain semantic feature map and the target domain semantic feature map respectively, and construct a detection loss function.

[0024] Step S150, use the scattering structure map alignment loss function, the global feature alignment loss function and the detection loss function to adjust the adjustable parameters in the multi-layer shared feature extractor and the detection head until convergence, and construct a cross-domain target detection network according to the trained multi-layer shared feature extractor and the detection head; Step S160, obtain the SAR image to be detected in the target domain, and use the cross-domain target detection network to perform target detection on the SAR image to be detected to obtain the target category and the target location.

[0025] For the SAR target detection task across imaging sensors, the detection model trained on the source domain needs to be generalized to the target domain. Assume that the source domain data with complete ground truth labels is , where the source domain contains samples, represents the source domain samples, represents the class labels. The target domain data without ground truth labels is , which contains samples. The purpose of the proposed method is to obtain the labels and candidate boxes of the target domain data through feature transfer alignment between the source domain and the target domain.

[0026] In step S100, the source domain dataset and the target domain dataset are constructed respectively, and part of the data in the target domain dataset is used for training and part for testing. The source domain dataset and the target domain dataset are obtained by two different sensors detecting the same target scene. Due to the differences in sensors, the detection angles and poses are different, so the content of the source domain SAR image and the target domain SAR image is not the same.

[0027] In step S110, a set of training pairs is input into the multi-layer shared feature extractor to extract multi-level SAR feature maps, and different-dimensional constraints are constructed respectively during the training process according to the SAR feature maps, so that the detection model trained on the source domain can be generalized to the target domain.

[0028] In this embodiment, a training method through graph structure alignment is mainly proposed. First, the corresponding semantic scattering graph structures are constructed according to the source domain semantic feature map and the target domain semantic feature map respectively. The process is as Figure 2 shown, including: for the source domain semantic feature map, the true boxes in the corresponding ground truth labels are used to perform spatially uniform sampling of pixel points at different feature layers to obtain the source domain semantic scattering point set. For the target domain semantic feature map, the class scores of each pixel point are generated by the detection head, and according to the class scores, the target source domain semantic scattering point set is obtained. Then, the feature points in the source domain semantic scattering point set and the target source domain semantic scattering point set are respectively mapped into the low-dimensional vector space, converted into graph nodes, and the corresponding semantic scattering graph structures are constructed.

[0029] In this embodiment, a bidirectional spatial aggregation mechanism is added to the multi-layer shared feature extractor , and the multi-layer shared feature extractor is used to extract the high-dimensional semantic features of the target in the source domain and the target domain to obtain the feature maps of the source domain and the target domain; .

[0030] Specifically, for the source domain semantic feature map, spatial uniform sampling is performed at different feature layers. The pixels within the source domain ground truth box are sampled to obtain target semantic scattering nodes, and the pixels outside the target box are collected as background semantic scattering nodes at a ratio of 1:2. Subsequently, source domain semantic scattering point sets are obtained and denoted as : .

[0031] Specifically, for the target domain image, a detection head is used to generate class scores for each feature scattering pixel point. The points with scores greater than the threshold are used as target feature points, and background semantic scattering nodes are sampled from the pixels with lower scores at a sampling rate of 1:2. Target semantic scattering nodes are dynamically sampled from the pixels with higher scores. Target domain semantic scattering points are obtained, and the target source domain semantic scattering point set is constructed and denoted as: .

[0032] Furthermore, to achieve fine alignment of the target in domain adaptation, high-dimensional semantic scattering feature points are converted into graph nodes, and edges are constructed between each node. The target domain and the source domain are modeled as graph structures, and the source domain semantic scattering graph structure and the target domain semantic scattering graph structure are obtained, which are respectively denoted as: ; ; In the above formula, , , where , respectively represent the node sets in the source domain semantic scattering graph structure and the target domain semantic scattering graph structure. The edge connection between nodes can be expressed as , , where the superscripts i and j respectively represent different nodes.

[0033] In this embodiment, the semantic scattering nodes are mapped to a low-dimensional vector space to retain the graph structure information and semantic node information, improve the effectiveness of subsequent feature alignment, and achieve adaptation and generalization of the target scattering information.

[0034] In step S120, according to the source domain semantic scattering graph structure and the target domain semantic scattering graph structure, constructing the scattering graph structure alignment loss function includes: first, calculating the semantic feature point alignment loss function according to the source domain semantic scattering graph structure and the target domain semantic scattering graph structure. Then, for each graph node in the source domain semantic scattering graph structure and the target domain semantic scattering graph structure respectively, using the feature distribution of the graph node and the nearest neighbor algorithm for node enhancement to obtain the source domain node enhanced graph structure and the target domain node enhanced graph structure. After that, after interacting the source domain node enhanced graph structure and the target domain node enhanced graph structure, calculating the semantic scattering point classification loss function. Finally, according to the maximum node similarity, minimum node similarity and edge similarity of the matching node pairs in the source domain node enhanced graph structure and the target domain node enhanced graph structure, constructing the semantic similarity graph structure matching loss function, and obtaining the scattering graph structure alignment loss function according to the semantic feature point alignment loss function, the semantic scattering point classification loss function and the semantic similarity graph structure matching loss function.

[0035] Specifically, in order to align the nodes of the source domain and the target domain, a semantic feature point alignment loss function is designed, expressed as: ; In the above formula, represents the node alignment discriminator network, represents the th graph node in the source domain semantic scattering graph structure, represents the th graph node in the target domain semantic scattering graph structure, represents the domain classification label, 、 respectively represent the number of graph nodes in the source domain semantic scattering graph structure and the target domain semantic scattering graph structure.

[0036] Furthermore, since the semantic scattering points obtained by randomly sampling the image will miss some key scattering point features, resulting in incomplete semantic information, thus affecting the effect of SAR target cross-sensor domain adaptation detection. In order to further enhance the semantic representation of the source domain and the target domain, node information of different domains is generated according to the sampled graph node distribution. Since there are large differences in the variance information of different domains and it will change with the training process, the variance of the opposite domain is used to ensure the similarity of the node distribution in the semantic space and improve the effect of domain adaptation.

[0037] In this embodiment, when enhancing each graph node in the source domain semantic scattering graph structure and the target domain semantic scattering graph structure: model the distribution of graph nodes in the semantic scattering graph structure as a Gaussian distribution, sample feature points from this Gaussian distribution, then obtain the enhanced semantic scattering nodes according to the linear projection formula, and construct the first enhanced semantic scattering graph node set. By taking the union of the graph nodes in the original semantic scattering graph node set and the first enhanced semantic scattering graph node set, obtain the second enhanced semantic scattering graph node set. Based on the original semantic scattering graph node set, obtain the corresponding adjacency matrix through inner product multiplication. Use a graph convolutional network combined with the adjacency matrix to enhance each graph node in the second enhanced semantic scattering graph node set according to the features of neighboring nodes, obtain the third enhanced semantic scattering graph node set, and perform a clustering operation on the graph nodes in the third enhanced semantic scattering graph node set using the K-nearest neighbor algorithm to obtain the final node-enhanced graph structure.

[0038] Among them, before constructing the first enhanced semantic scattering graph node set, for the target domain semantic scattering graph structure, also utilize the distribution variance of the graph structure in the target domain semantic scattering graph structure and the mean of the source domain semantic scattering graph structure to enhance the features of the graph nodes in the target domain semantic scattering graph structure, so as to reduce the distribution difference between the target domain and the source domain.

[0039] Specifically, model the distribution of graph nodes in the semantic scattering graph structures of the source domain and the target domain as Gaussian distributions, where and represent the feature points sampled from the Gaussian distribution. Through linear projection obtain the enhanced semantic scattering nodes, generate the enhanced scattering point structure diagram, thereby further enhancing the representation ability of the semantic scattering points and reducing the problem of uneven distribution of scattering points caused by differences in sampling methods. The enhanced semantic scattering node sets of the source domain and the target domain, namely the first enhanced semantic scattering graph node set, can be expressed as: ; ; In the above formula, and respectively follow Gaussian distributions and , , respectively represent the first enhanced semantic scattering graph node sets of the source domain and the target domain.

[0040] Then, take the union of the original semantic scattering graph node set and the first enhanced semantic scattering graph node set to obtain the second enhanced semantic scattering graph node set, where and The original semantic scattering graph node sets representing the source domain and the target domain, and the second enhanced semantic scattering graph node sets are respectively represented as: ; ; Among them, the linear projection operation is used to map the sampled feature points to the target space and can be expressed as: ; In the above formula, W and b are weight coefficients.

[0041] Furthermore, the adjacency matrices are respectively constructed for the second enhanced semantic scattering graph node sets of the source domain and the target domain, and the adjacency matrix capable of characterizing the internal node association information is obtained through the inner product multiplication, which is expressed as: ; ; In the above formula, and respectively represent the adjacent matrices in a single graph structure of the source domain and the target domain, which are used to represent the encoded structure information, and is a learnable weight matrix.

[0042] In this embodiment, graph convolution is used for message propagation, and a single-layer GCN layer is used for long-distance semantic transfer to aggregate the semantic knowledge across images, thereby enhancing the representation of nodes. The graph convolution operation can be expressed as: ; In the above formula, is the node feature matrix of the th layer, is the node feature matrix of the th layer, is a learnable weight matrix, is the adjacency matrix, is 's degree matrix, is the activation function.

[0043] Furthermore, in the graph convolution operation, the network aggregates the information of neighboring nodes through the product of the adjacency matrix of the nodes and the node feature matrix, enhances the features of the current node with the neighboring nodes, and obtains the nodes respectively aggregated through graph convolution, that is, the third enhanced semantic scattering graph node sets and , which are respectively expressed as: ; ; In the above formula, and represent and the neighboring nodes of and represent the elements in the adjacency matrix, which are learnable parameters.

[0044] Furthermore, map the semantic scattering point map after completing the semantic graph structure to the established graph structure, and use the KNN clustering method to cluster different scattering points of the target, so that the locally connected semantic scattering points in the graph are similar for those classified into one category, and then obtain different semantic scattering components of the target. Cluster the positive sample nodes of the target to obtain the semantic scattering components of the target. Semantic scattering component clustering can capture the structured information of the target in the SAR image, enhance the network's understanding of the target structure, thereby improving the accuracy of semantic alignment and capturing the hidden structure in the graph structure.

[0045] As Figure 3 shown, it is a schematic diagram of the process of enhancing the nodes of the semantic scattering graph.

[0046] To narrow the feature distance between the source domain and the target and reduce the feature distribution difference, establish a semantic interaction correlation relationship between the source domain and the target domain between the two constructed graphs. By interacting and fusing the information between the two domains, achieve feature alignment, so that each semantic scattering node has the ability of cross-domain perception, and then learn the semantic information of the other domain, as Figure 4 shown.

[0047] Interact the scattering information of the target domain with the semantic scattering information of the source domain itself, and the set of source domain semantic scattering nodes after interaction is: ; Interact the scattering information of the source domain with the semantic scattering information of the target domain itself, and the set of target domain semantic scattering nodes after interaction is: ; In the above two formulas, and represent the enhanced set of semantic scattering nodes, , and are the learned weight matrices.

[0048] Furthermore, use linear transformation and normalization operations to update the node representation, update the obtained nodes, and enable the scattering information of different domains to interact, narrowing the domain distribution difference. Finally, obtain the semantic scattering point classification loss , expressed as: ; In the above formula, and are the ground truth classification labels of the source domain and target domain images respectively, represents the classification branch of the detection head, the classifier.

[0049] Furthermore, by learning the inherent semantic relationship between two graph nodes in the source domain and the target domain, encoding it into an associated representation, and achieving domain adaptation through node-to-node matching, that is, constructing a semantic scattering structure graph alignment loss function, the process is as Figure 4 shown.

[0050] First, construct a similarity matrix to measure the semantic similarity between scattering nodes in the graph. Specifically, project the input node feature matrix, where the node feature matrix is the enhanced graph structure. Map the features to a new space, and then use the cosine similarity to calculate the semantic feature similarity between the feature vectors of two nodes and , and encode it as an element of the node similarity matrix , expressed as: ; Traverse the node pairs between the two domains, combine the similarity values of all node pairs into a matrix form to measure the semantic similarity between nodes in the source domain and target domain graphs, and obtain the final similarity matrix , expressed as: ; Construct a connection indication matrix between the two graph structures in the source domain and the target domain, where the element is If two nodes have the same input category, assign the element of the connection indication matrix to 1, otherwise assign it to 0, expressed as: ; For each source domain node , find the node in the target domain that matches it, which can maximize the similarity between them, enhance the similarity of correctly matched node pairs, and ensure that the similarity of the node pairs is as close to 1 as possible, obtaining the maximum node similarity loss , expressed as: ; In the above formula, represents the Hadamard product.

[0051] Next, traverse each node in the source domain and each node , for node pairs that do not belong to the same category, minimize their similarity to avoid incorrect matching. Suppress the similarity of incorrectly matched node pairs to ensure that the similarity of incorrectly matched node pairs is as close to 0 as possible, obtaining the minimum node similarity loss , expressed as: ; Furthermore, to ensure the consistency of the structural information of the source domain and target domain graphs, enabling the matched node pairs to have similar connection relationships in the graph and maintaining the integrity of the graph structure and the consistency of component composition, an edge structure loss is designed , expressed as: ; In the above formula, and , respectively represent the adjacency matrices of the source domain and target domain graphs, encoding the structural information of the graphs. If the source domain node and the target domain node belong to the same type of sample, it is assigned 1, otherwise 0.

[0052] Then, the semantic similarity graph structure matching loss is expressed as: ; In step S130, source domain and target domain feature maps at different scales and are input into the discriminator , obtaining a multi-level global feature alignment loss, expressed as: ; In the above formula, represents the discriminator, which consists of 3 discriminators and 1 gradient backpropagation layer, represents the number of layers of the multi-level feature extractor, , represent the coordinates of the pixel points.

[0053] In step S140, the detection loss mainly includes the category output loss of the target , the centerness loss and the position regression loss . Among them, uses the focal loss for optimization, uses the binary cross-entropy loss to optimize the centerness of centerness, uses the IOU loss for optimization.

[0054] ; In the above formula, represents the predicted probability of the sample, represents the parameter for adjusting the weights of positive and negative samples, and the superscript r represents the harmonic factor.

[0055] ; In the above formula, is the true label (0 or 1), is the centerness value predicted by the model.

[0056] ; In the above formula, and are the areas of the regions enclosed by the true coordinates and the predicted coordinates, respectively.

[0057] In step S150, in order to optimize the entire cross-sensor SAR target detection model, the losses in the network are integrated, and the resulting final collaborative loss mainly includes the multi-level global alignment loss , the graph scattering node alignment loss , the graph scattering node classification loss , the graph structure matching loss and the detection loss. The total loss function is expressed as: ; In the above formula, and represent the loss weight coefficients.

[0058] After the training converges, the trained multi-level shared feature extractor and the detection head are used to construct a cross-domain target detection network.

[0059] In this paper, the effectiveness of the proposed method is also demonstrated by experimental results. In the experiment, three datasets from different imaging sensors are used to experimentally verify the algorithm. The SAR measured data of MiniSAR and FARAD are used to verify the proposed method in this paper. The imaging sensor band of MiniSAR data is the Ku band, and the shooting location is in the suburbs. The sensor bands of FARAD data are the Ka band and the X band, and the shooting location is in the urban area with dense buildings. The two types of datasets cover three imaging bands, and FARAD and MiniSAR use completely different sensors. Among them, the image size is 512x512. The training set and the test set are divided according to the ratio of 7:3, and the used metrics are Precision (PR), Recall (RE), mAP and F1.

[0060] Two different cross-sensor experimental tasks were designed, including from the FARAD Ka sensor to the MiniSAR sensor, and from the MiniSAR sensor to the FARAD X sensor. Experiments were conducted on two types of cross-sensor tasks and compared with 9 classic deep learning and domain adaptation algorithms. The experimental results are shown in Tables 1 and 2. It can be seen that the algorithm proposed by this method can achieve good experimental results.

[0061] As shown in Table 1, in the Ka—>M cross-sensor task, the proposed method obtained 82.71% and 68.54% in the Precision and Recall metrics respectively. It is about 10% higher than the second-best algorithm with better performance. Compared with the classic Faster-RCNN algorithm, the proposed algorithm is about 40% higher in the mAP metric and about 30% higher in the F1 metric, significantly improving the performance in the cross-sensor object detection task.

[0062] As shown in Table 2, in the M—>X cross-sensor task, the proposed method reached 71.84% and 75.10% in the Precision and Recall metrics respectively. Compared with the EPM method, the mAP value of the proposed algorithm reached 68.27%, about 15% higher than the EPM algorithm. In addition, in terms of the F1 value, the proposed algorithm exceeded other algorithms by about 13%-30%, achieving relatively superior performance.

[0063] Table 1 Results on the Ka—>M cross-sensor task

[0064] Table 2 Results on the M—>X cross-sensor task

[0065] In the above cross-sensor SAR image object detection method, by inputting the SAR source domain image with true value labels and the SAR target domain image without true value labels into the multi-layer shared feature extractor, the corresponding semantic feature maps are obtained respectively, and a corresponding graph structure is constructed based on this feature map. By constructing a loss function to constrain the consistency of the source domain semantic graph structure and the target domain semantic graph structure during training, the trained network has good generalization ability on the target domain. At the same time, combined with the object detection constraint and the global feature constraint, it can achieve accurate object detection of the target domain SAR image under the training of the target domain image without true value labels.

[0066] It should be understood that although Figure 1The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in Figure 1 may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0067] In one embodiment, as Figure 6 shown, a cross-sensor SAR image target detection device is provided, including: a dataset acquisition module 200, a graph structure construction module 210, a graph structure alignment constraint module 220, a global feature alignment constraint module 230, a target detection constraint module 240, a training module 250, and a cross-domain detection module 260, where: The dataset acquisition module 200 is configured to acquire a source domain dataset and a target domain dataset. The source domain dataset includes multiple SAR source domain images with ground truth labels, and the target domain dataset includes multiple SAR target domain images without ground truth labels; The graph structure construction module 210 is configured to extract one SAR image from the source domain dataset and the target domain dataset respectively as a training pair, input the training pair into a multi-layer shared feature extractor to obtain a source domain semantic feature map and a target domain semantic feature map respectively, and construct corresponding semantic scattering graph structures according to the source domain semantic feature map and the target domain semantic feature map respectively, to obtain a source domain semantic scattering graph structure and a target domain semantic scattering graph structure; The graph structure alignment constraint module 220 is configured to construct a scattering graph structure alignment loss function according to the source domain semantic scattering graph structure and the target domain semantic scattering graph structure; The global feature alignment constraint module 230 is configured to use a discriminator to construct a multi-level global feature alignment loss function according to the source domain semantic feature map and the target domain semantic feature map; The target detection constraint module 240 is configured to use a detection head to perform target detection according to the source domain semantic feature map and the target domain semantic feature map respectively, and construct a detection loss function; The training module 250 is configured to use the scattering graph structure pair loss function, the global feature alignment loss function, and the detection loss function to adjust the adjustable parameters in the multi-layer shared feature extractor and the detection head until convergence, and construct a cross-domain target detection network according to the trained multi-layer shared feature extractor and the detection head; The cross-domain detection module 260 is configured to obtain the SAR image to be detected in the target domain, and perform target detection on the SAR image to be detected by using the cross-domain target detection network, so as to obtain the target category and the target location.

[0068] For the specific limitations of the cross-sensor SAR image target detection device, reference can be made to the limitations of the cross-sensor SAR image target detection method in the foregoing text, which will not be elaborated herein. Each module in the above cross-sensor SAR image target detection device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in the form of hardware or independent of it, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0069] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a cross-sensor SAR image target detection method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0070] Those skilled in the art can understand that Figure 7 the structure shown in

[0071] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Obtain a source domain dataset and a target domain dataset, where the source domain dataset includes multiple SAR source domain images with ground truth labels, and the target domain dataset includes multiple SAR target domain images without ground truth labels; Extract a SAR image from the source domain dataset and the target domain dataset respectively as a training pair, input the training pair into the multi-layer shared feature extractor to obtain the source domain semantic feature map and the target domain semantic feature map respectively, and construct corresponding semantic scattering map structures according to the source domain semantic feature map and the target domain semantic feature map, to obtain the source domain semantic scattering map structure and the target domain semantic scattering map structure; Construct a scattering structure map alignment loss function according to the source domain semantic scattering map structure and the target domain semantic scattering map structure; Use a discriminator to construct a multi-level global feature alignment loss function according to the source domain semantic feature map and the target domain semantic feature map; Use a detection head to perform object detection according to the source domain semantic feature map and the target domain semantic feature map respectively, and construct a detection loss function; Use the scattering structure map alignment loss function, the global feature alignment loss function, and the detection loss function to adjust the adjustable parameters in the multi-layer shared feature extractor and the detection head until convergence, and construct a cross-domain object detection network according to the trained multi-layer shared feature extractor and the detection head; Obtain the SAR image to be detected in the target domain, and use the cross-domain object detection network to perform object detection on the SAR image to be detected, to obtain the object category and the object location.

[0072] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Obtain a source domain dataset and a target domain dataset. The source domain dataset includes multiple SAR source domain images with ground truth labels, and the target domain dataset includes multiple SAR target domain images without ground truth labels; Extract a SAR image from the source domain dataset and the target domain dataset respectively as a training pair, input the training pair into the multi-layer shared feature extractor to obtain the source domain semantic feature map and the target domain semantic feature map respectively, and construct corresponding semantic scattering map structures according to the source domain semantic feature map and the target domain semantic feature map, to obtain the source domain semantic scattering map structure and the target domain semantic scattering map structure; Construct a scattering structure map alignment loss function according to the source domain semantic scattering map structure and the target domain semantic scattering map structure; Use a discriminator to construct a multi-level global feature alignment loss function according to the source domain semantic feature map and the target domain semantic feature map; Use a detection head to perform object detection according to the source domain semantic feature map and the target domain semantic feature map respectively, and construct a detection loss function; Adjust the tunable parameters in the multi-layer shared feature extractor and the detection head using the loss function, global feature alignment loss function, and detection loss function based on the scattering structure diagram until convergence, and construct a cross-domain object detection network based on the trained multi-layer shared feature extractor and detection head; Obtain the SAR image to be detected in the target domain, and perform object detection on the SAR image to be detected using the cross-domain object detection network to obtain the object category and object location.

[0073] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0074] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0075] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A cross-sensor SAR image target detection method, characterized in that: The method comprises: Acquire a source domain dataset and a target domain dataset, wherein the source domain dataset includes a plurality of SAR source domain images with true value labels, and the target domain dataset includes a plurality of SAR target domain images without true value labels; Extracting a SAR image from the source domain data set and the target domain data set as a training pair, respectively, inputting the training pair into a multi-layer shared feature extractor to obtain a source domain semantic feature map and a target domain semantic feature map, respectively, and constructing corresponding semantic scattering map structures according to the source domain semantic feature map and the target domain semantic feature map, respectively, to obtain a source domain semantic scattering map structure and a target domain semantic scattering map structure; According to the source domain semantic scattering map structure and the target domain semantic scattering map structure, construct a scattering structure map alignment loss function; Using the discriminator, a multi-level global feature alignment loss function is constructed according to the source domain semantic feature map and the target domain semantic feature map; Using the detection head, performing target detection according to the source domain semantic feature map and the target domain semantic feature map, and constructing a detection loss function; Using the scattering structure graph to adjust the adjustable parameters in the multi-layer shared feature extractor and the detection head according to its loss function, global feature alignment loss function and detection loss function until convergence, and constructing a cross-domain target detection network according to the trained multi-layer shared feature extractor and detection head; A SAR image to be detected in a target domain is obtained, and target detection is performed on the SAR image to be detected using the cross-domain target detection network to obtain a target category and a target position.

2. The cross-sensor SAR image target detection method according to claim 1, characterized in that: The method of constructing a corresponding semantic scattering graph structure according to the source domain semantic feature graph and the target domain semantic feature graph includes: For the source domain semantic feature map, the true value box in the corresponding true value label is used to uniformly sample the space of pixels at different feature layers to obtain a set of source domain semantic scattering points; For the target domain semantic feature map, the detection head is used to generate the category score of each pixel point, and the target source domain semantic scatter point set is obtained based on the category score; The feature points in the source domain semantic scattering point set and the target source domain semantic scattering point set are respectively mapped to a low-dimensional vector space, converted into graph nodes, and a corresponding semantic scattering graph structure is constructed.

3. The cross-sensor SAR image target detection method according to claim 2, characterized in that: The constructing of the scattering structure graph alignment loss function according to the source domain semantic scattering graph structure and the target domain semantic scattering graph structure comprises: Calculating a semantic feature point alignment loss function according to the source domain semantic scattering map structure and the target domain semantic scattering map structure; Respectively enhancing the graph nodes in the source domain semantic scattering graph structure and the target domain semantic scattering graph structure by using the characteristic distribution of the graph nodes and the nearest neighbor algorithm to obtain the source domain node enhanced graph structure and the target domain node enhanced graph structure; After interacting with the source domain node enhanced graph structure and the target domain node enhanced graph structure, calculating a semantic scatter point classification loss function; Constructing a semantic similarity graph structure matching loss function according to the maximum node similarity, the minimum node similarity and the edge similarity of the matching node pairs in the source domain node enhanced graph structure and the target domain node enhanced graph structure; The scattering structure graph alignment loss function is obtained according to the semantic feature point alignment loss function, the semantic scattering point classification loss function and the semantic similarity graph structure matching loss function.

4. The cross-sensor SAR image target detection method according to claim 3, characterized in that: The semantic feature point alignment loss function is expressed as: ; In the above formula, represents the node alignment discriminator network, represents the first Graph nodes, The first Graph nodes, represents the domain classification label, , Respectively represent the number of graph nodes in the source domain semantic scatter graph structure and the target domain semantic scatter graph structure.

5. The cross-sensor SAR image target detection method according to claim 4, characterized in that: When enhancing each graph node in the source domain semantic scatter graph structure and the target domain semantic scatter graph structure: The distribution of graph nodes in the semantic scatter graph structure is modeled as a Gaussian distribution, feature points are sampled from the Gaussian distribution, enhanced semantic scatter nodes are obtained according to a linear projection formula, and a first enhanced semantic scatter graph node set is constructed; The original semantic scatter graph node set and the graph nodes in the first enhanced semantic scatter graph node set are taken as a union to obtain a second enhanced semantic scatter graph node set; Based on the original semantic scatter graph node set, a corresponding adjacency matrix is ​​obtained by inner product multiplication, and each graph node in the second enhanced semantic scatter graph node set is enhanced according to the characteristics of the neighboring nodes by using a graph convolution network combined with the adjacency matrix to obtain a third enhanced semantic scatter graph node set; A K-nearest neighbor algorithm is used to perform a clustering operation on the graph nodes in the third enhanced semantic scatter graph node set to obtain a final node enhanced graph structure.

6. The cross-sensor SAR image target detection method according to claim 5, characterized in that: Before constructing the first enhanced semantic scatter graph node set, for the target domain semantic scatter graph structure, the features of the graph nodes in the target domain semantic scatter graph structure are enhanced by using the distribution variance of the graph structure in the target domain semantic scatter graph structure and the mean of the source domain semantic scatter graph structure.

7. A cross-sensor SAR image target detection device, characterized in that: The device comprises: A data set acquisition module, used to acquire a source domain data set and a target domain data set, wherein the source domain data set includes a plurality of SAR source domain images with true value labels, and the target domain data set includes a plurality of SAR target domain images without true value labels; A graph structure construction module is used to extract a SAR image from the source domain data set and the target domain data set as a training pair, input the training pair into a multi-layer shared feature extractor to obtain a source domain semantic feature map and a target domain semantic feature map, respectively, and construct corresponding semantic scattering graph structures according to the source domain semantic feature map and the target domain semantic feature map, respectively, to obtain a source domain semantic scattering graph structure and a target domain semantic scattering graph structure; A graph structure alignment constraint module, used to construct a scattering structure graph alignment loss function according to the source domain semantic scattering graph structure and the target domain semantic scattering graph structure; A global feature alignment constraint module, used to construct a multi-level global feature alignment loss function based on the source domain semantic feature map and the target domain semantic feature map using a discriminator; A target detection constraint module, used to use a detection head to perform target detection according to the source domain semantic feature map and the target domain semantic feature map, and to construct a detection loss function; A training module, used to adjust the adjustable parameters in the multi-layer shared feature extractor and the detection head by using the scattering structure graph to adjust the loss function, the global feature alignment loss function and the detection loss function until convergence, and to construct a cross-domain target detection network according to the trained multi-layer shared feature extractor and the detection head; The cross-domain detection module is used to obtain the SAR image to be detected in the target domain, and use the cross-domain target detection network to perform target detection on the SAR image to be detected to obtain the target category and target position.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • SAR (Synthetic Aperture Radar) target identification method and device based on depth Grassmann manifold space

    CN117315323A

  • SAR (Synthetic Aperture Radar) image rotating target detection method and system based on scattering key point guided diffusion model

    CN119007030A

  • Cross-sensor-based SAR (Synthetic Aperture Radar) image target detection method, device, equipment and medium

    CN119810668A

  • Synthetic aperture radar (SAR) image target detection method

    US20230169623A1