Construction method of detection model of domain generalization remote sensing image, detection model, detection method of detection model, detection system, storage medium and electronic equipment

By constructing a detection model of pretrained feature extraction unit and multi-scale deformable attention mechanism, the performance degradation of remote sensing image object detection algorithm in complex scenarios is solved, and high-performance and robust remote sensing object detection is achieved.

CN120388277APending Publication Date: 2025-07-29启元实验室
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510311833.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing remote sensing image object detection algorithms have dramatically declined when facing complex scenarios beyond the training distribution, and fine-tuning of models that rely on large amounts of target domain data does not have the best choice for enhanced robustness.

Method used

By acquiring the remote sensing image data set and environmental algorithm, a cross-domain object detection data set is generated, and a pre-trained feature extraction unit, a domain generalized feature extraction unit, an encoder-decoder unit and an object detection head unit are used to construct a detection model. A multi-scale deformable attention mechanism and an object detection head unit are used to determine the detection box result data.

Benefits of technology

Without the need for a large amount of target domain data, high-performance detection of remote sensing targets in complex scenarios is achieved, enhancing the robustness and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388277A_ABST
    Figure CN120388277A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method of a detection model of a domain generalization remote sensing image, the detection model, a detection method of the detection model, a detection system, a storage medium and electronic equipment. The construction method comprises the following steps: determining a remote sensing target detection data set through an acquired remote sensing image data set and an environment algorithm; determining a preset hierarchical feature map according to the remote sensing target detection data set; determining a domain generalization object query quantity and a first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and a preset feature sequence corresponding to the preset hierarchical feature map; determining a fusion object query quantity set and reference coordinates corresponding to the fusion object query quantity set according to the domain generalization object query quantity and the first multi-scale feature map; and determining detection frame result data according to the fusion object query quantity set and the reference coordinates so as to construct a detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image detection. Specifically, it relates to a method for constructing a detection model of domain-generalized remote sensing images, the detection model, its detection method, a detection system, a storage medium, and an electronic device. Background Art

[0002] Inspired by the attention-based encoder-decoder model Transformer, object detection algorithms based on Transformer, represented by DETR (Detection Transformer), have emerged in the field of computer vision. Such algorithms achieve detection tasks through an end-to-end network architecture, without relying on post-processing operations such as Non-Maximum Suppression (NMS), greatly simplifying the object detection processing flow.

[0003] Due to the unique top-down shooting method of remote sensing images, the observed targets are usually small in size and vary in angles. At the same time, affected by clouds and lighting conditions, the background information is complex and changeable, posing higher requirements for the robustness and generalization ability of object detection algorithms.

[0004] The inventors of the present application found that mainstream object detection algorithms that use images taken under ideal and interference-free conditions as training data and highly rely on the assumption of the same distribution between the source domain and the target domain experience a sharp decline in detection performance when facing complex scenarios such as bad weather beyond the training distribution. And fine-tuning the model by collecting a large number of remote sensing images in different scenarios is not the best choice to improve detection performance and enhance the robustness of the model.

[0005] The content of the part of the background art is only the technology known to the applicant and does not necessarily represent the prior art in this field. Summary of the Invention

[0006] According to an aspect of the present application, there is provided a method for constructing a detection model of domain-generalized remote sensing images, where the detection model is used to detect objects in remote sensing images. The construction method includes: determining a remote sensing object detection data set through the obtained remote sensing image data set and an environmental algorithm; determining a preset hierarchical feature map according to the remote sensing object detection data set; determining a domain-generalized object query quantity and a first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and a preset feature sequence corresponding to the preset hierarchical feature map; determining a set of fused object query quantities and reference coordinates corresponding to the set of fused object query quantities according to the domain-generalized object query quantity and the first multi-scale feature map; and determining detection box result data according to the set of fused object query quantities and the reference coordinates to construct the detection model.

[0007] According to one aspect of the present application, a detection model for domain-generalized remote sensing images is provided. The detection model is constructed by the construction method described above. The detection model includes a pre-trained feature extraction unit, a domain-generalized feature extraction unit, an encoder-decoder unit, and an object detection head unit. The pre-trained feature extraction unit is used to extract a remote sensing object detection data set to determine a preset hierarchical feature map, where the remote sensing object detection data set is determined by the acquired remote sensing images and environmental algorithms; the domain-generalized feature extraction unit is used to determine a domain-generalized object query amount and a first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and a preset feature sequence; the encoder-decoder unit has a multi-scale deformable attention mechanism and is used to determine a set of fused object query amounts and reference coordinates corresponding to the fused object query amounts according to the domain-generalized object query amount and the first multi-scale feature map; the object detection head unit is used to determine detection box result data according to the set of fused object query amounts and the reference coordinates.

[0008] According to one aspect of the present application, a detection method for a detection model based on domain-generalized remote sensing images is provided. The detection model is constructed by the construction method described above. The detection method includes: extracting a remote sensing object data set to be detected to determine a preset hierarchical feature map; determining a domain-generalized object query amount and a first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and a preset feature sequence; determining a set of fused object query amounts and reference coordinates corresponding to the fused object query amounts according to the domain-generalized object query amount and the first multi-scale feature map; determining detection box result data according to the set of fused object query amounts and the reference coordinates.

[0009] According to one aspect of the present application, a detection system for a detection model based on domain-generalized remote sensing images is provided. The detection model is constructed by the construction method described above. The detection system includes an image input unit, a pre-trained feature extraction unit, a domain-generalized feature extraction unit, an encoder-decoder unit, and an object detection head unit. The image input unit acquires a remote sensing object data set to be detected; the pre-trained feature extraction unit extracts the remote sensing object data set to be detected to determine a preset hierarchical feature map; the domain-generalized feature extraction unit determines a domain-generalized object query amount and a first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and a preset feature sequence; the encoder-decoder unit has a multi-scale deformable attention mechanism and determines a set of fused object query amounts and reference coordinates corresponding to the fused object query amounts according to the domain-generalized object query amount and the first multi-scale feature map; the object detection head unit determines detection box result data according to the set of fused object query amounts and the reference coordinates.

[0010] According to another aspect of the present application, the present application further provides a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it can implement the method for constructing a detection model of domain-generalized remote sensing images as described above.

[0011] According to another aspect of the present application, the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors can implement the method for constructing a detection model of domain-generalized remote sensing images as described above.

[0012] According to another aspect of the present application, the present application further provides a computer program product, including: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method for constructing a detection model of domain-generalized remote sensing images as described above.

[0013] According to another aspect of the present application, the present application further provides a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it can implement the detection method of the detection model based on domain-generalized remote sensing images as described above.

[0014] According to another aspect of the present application, the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors can implement the detection method of the detection model based on domain-generalized remote sensing images as described above.

[0015] According to another aspect of the present application, the present application further provides a computer program product, including: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the detection method of the detection model based on domain-generalized remote sensing images as described above.

[0016] The construction method provided by the present application can cope with cross-domain distribution differences without a large amount of target domain data and achieve high-performance detection of remote sensing targets in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 Schematic flowchart of the construction method 1000 according to an embodiment of the present application;

[0019] Figure 2 Schematic flowchart of step S130 according to an embodiment of the present application;

[0020] Figure 3 Schematic flowchart of determining the first multi-scale feature map according to an embodiment of the present application;

[0021] Figure 4 Schematic flowchart of step S140 according to an embodiment of the present application;

[0022] Figure 5 Schematic flowchart of step S143 according to an embodiment of the present application;

[0023] Figure 6 Schematic flowchart of step S1431 according to an embodiment of the present application;

[0024] Figure 7 Schematic diagram of the multi-scale deformable attention mechanism according to an embodiment of the present application;

[0025] Figure 8 Schematic flowchart of step S145 according to an embodiment of the present application;

[0026] Figure 9 Schematic flowchart of step S1451 according to an embodiment of the present application;

[0027] Figure 10 Schematic diagram of determining the reference coordinates corresponding to the set of object query amounts according to an embodiment of the present application;

[0028] Figure 11 Schematic flowchart of step S150 according to an embodiment of the present application;

[0029] Figure 12 Schematic diagram of the structure of the detection model according to an embodiment of the present application;

[0030] Figure 13 Schematic flowchart of the detection method 3000 according to an embodiment of the present application;

[0031] Figure 14 Schematic diagram of the structure of the detection system according to an embodiment of the present application.

[0032] Reference numerals

[0033] Detection model 200.

[0034] Pre-trained feature extraction unit 210; Domain generalization feature extraction unit 220; Encoder-decoder unit 230; Target detection head unit 240.

[0035] Detection system 400.

[0036] Image input unit 410. Detailed implementation mode

[0037] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repeated description will be omitted.

[0038] The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of these specific details, or can be implemented in other ways, components, materials, devices, etc. In these cases, well-known structures, methods, devices, implementations, materials, or operations will not be shown or described in detail.

[0039] In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0040] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order.

[0041] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of this application.

[0042] According to one aspect of this application, this application provides a method for constructing a detection model for domain-generalized remote sensing images. The construction method 1000 can be executed by a construction system. The construction system can be a host (or server) with data processing capabilities.

[0043] See Figure 1, the construction method 1000 may include steps S110 - S150.

[0044] In step S110, the construction system determines a remote sensing target detection dataset through the acquired remote sensing image dataset and environmental algorithms.

[0045] According to the exemplary embodiment, the construction system may acquire a publicly released remote sensing image dataset I ∈ R H×W×3 . Wherein, I is a remote sensing image; R H×W×3 is the real number field of R remote sensing images; H×W×3 is the dimension of the remote sensing image; H is the height of the remote sensing image; W is the width of the remote sensing image; 3 indicates that the number of channels of the remote sensing image is 3.

[0046] The remote sensing image may be a remote sensing image with three bands of red R, green G, and blue B. The remote sensing image dataset may be remote sensing image datasets such as the LEVIR dataset (Large-scale Evaluation of Visual Interpretability and Recognition), the DOTA dataset (Dataset for Object detection in Aerial images), the AOD dataset (Aerial Object Detection), etc.

[0047] The environmental algorithm may be an algorithm for simulating environmental processing on the remote sensing image dataset. For example, the environmental algorithm may include an image fogging algorithm based on the atmospheric scattering model and a low-light algorithm based on contrast compression, etc.

[0048] The construction system may acquire a remote sensing image dataset, crop the remote sensing image into image slices of 256×256, and make corresponding target annotation files (for example, VOC format target annotation files) through an annotation tool (such as LabelImg). For example, the construction system may annotate target objects in the remote sensing image, etc. In step S110, the construction system may simulate complex environments such as thick fog and limited lighting through the environmental algorithm, so that the remote sensing image dataset generates a multi-scenario cross-domain remote sensing target detection dataset.

[0049] In step S120, the construction system determines a preset hierarchical feature map according to the remote sensing target detection dataset.

[0050] According to the exemplary embodiment, the construction system may extract the preset hierarchical feature map of the remote sensing target detection dataset through a pre-trained feature extraction unit. The preset hierarchical feature map may be a feature map output at a preset specified level.

[0051] The pre-trained feature extraction unit can use the large model framework DINO (Distillation with No labels) of self-supervised learning in the field of visual understanding as the backbone network. DINO is a lightweight model trained by combining self-supervised learning and knowledge distillation based on the Vision Transformer (ViT) with 1 billion parameters and a large pre-trained set of high-quality natural images, which can sensitively capture global information and learn image generalization features by itself.

[0052] The backbone network of the pre-trained feature extraction unit has 12 Transformer layers, and each Transformer layer has 12 attention heads. For example, the preset levels can be the 2nd, 5th, 8th, and 11th layers. The construction system can freeze the DINO network parameters Φ M .

[0053] The construction system can output the 2nd layer feature map, 5th layer feature map, 8th layer feature map, and 11th layer feature map through the pre-trained feature extraction unit to form the preset level feature map. The outputs of different layers contain different levels of feature information, which are suitable for the object detection task. Among them, the shallow layer such as the 2nd layer is used to capture local features and low-level semantic information (such as edges, textures), the middle layers such as the 5th and 8th layers are used to capture middle-level semantic information (such as object shapes, etc.), and the deep layer such as the 11th layer is used to capture global features and high-level semantic information (such as object categories, the environment where they are located, etc.). Moreover, only using the outputs of some layers can effectively reduce the computational amount and achieve the purpose of lightweight.

[0054] In step S130, the construction system determines the domain generalization object query amount and the first multi-scale feature map of the preset level feature map according to the preset level feature map and the preset feature sequence corresponding to the preset level feature map.

[0055] According to the exemplary embodiment, the construction system can determine the object query amount and the first multi-scale feature map of the preset level feature map based on the domain generalization feature extraction unit (Domain Generalization, DG) according to the preset level feature map and the preset feature sequence corresponding to the preset level feature map.

[0056] The preset feature sequence can be a set of learnable tokens (tokens). The domain generalization object query amount can be the query vector generated by mapping this set of learnable tokens through a multi-layer perceptron. The first multi-scale feature map can be a feature map reflecting the feature information of the remote sensing image at different resolutions.

[0057] The domain generalization feature extraction unit can adopt an attention-based feature enhancement method and consists of a set of learnable token sequences (i.e., preset feature sequences). For example, a set of learnable token sequences can be: T = {T1, T2, …, T N}, N = 300.

[0058] The construction system can perform a multi-layer perceptron linear mapping on the preset feature sequence through the domain generalization feature extraction unit to obtain the domain generalization object query quantity.

[0059] In step S140, the construction system determines a set of fusion object query quantities and corresponding reference coordinates according to the domain generalization object query quantity and the first multi-scale feature map.

[0060] According to the exemplary embodiment, the set of fusion object query quantities can be the set of domain generalization object query quantities after encoding features and decoding features. Each domain generalization object query quantity in the set of fusion object query quantities can represent a potential target. The reference coordinates corresponding to the set of fusion object query quantities can be the position information of the potential target on the feature map.

[0061] The construction system can determine a set of object query quantities and corresponding reference coordinates according to the object query quantity and the first multi-scale feature map through an encoder-decoder unit with a multi-scale deformable attention mechanism.

[0062] In step S150, the construction system determines detection box result data according to the set of fusion object query quantities and the reference coordinates to construct a detection model.

[0063] According to the exemplary embodiment, the detection box result data can be the attribute data of the target in the remote sensing image. The detection box result data can include data such as the detection box position, target category, and probability.

[0064] The construction system can determine the detection box result data according to the set of fusion object query quantities and the reference coordinates through the target detection head unit. The target detection head unit consists of a classification head and a bounding box regression head adopting a fully connected layer structure. The classification head is responsible for predicting the category of each candidate box, and the bounding box regression head is responsible for predicting the position of the target. The construction system can determine the deviation between the detection box result data and the actual detection box data and train the detection model to construct the detection model.

[0065] Through the above embodiments, the present application determines a remote sensing target detection dataset based on the obtained remote sensing image dataset and environmental algorithms. The present application determines a preset hierarchical feature map based on the remote sensing target detection dataset. The present application determines a domain generalization object query quantity and a first multi-scale feature map of the preset hierarchical feature map based on the preset hierarchical feature map and a preset feature sequence corresponding to the preset hierarchical feature map. The present application determines a set of fusion object query quantities and corresponding reference coordinates based on the domain generalization object query quantity and the first multi-scale feature map. The present application determines detection box result data based on the set of fusion object query quantities and the reference coordinates.

[0066] The construction method provided by the present application can cope with cross-domain distribution differences and achieve high-performance detection of remote sensing targets in complex scenarios without the need for a large amount of target domain data.

[0067] Optionally, referring to Figure 2 , step S130 may include steps S131 - S136.

[0068] In step S131, the construction system determines a preset hierarchical similarity sequence based on the preset hierarchical feature map and a preset feature sequence corresponding to the preset hierarchical feature map.

[0069] According to the exemplary embodiment, the construction system may randomly initialize the Token sequence (i.e., the preset feature sequence), and during the fine-tuning process, embed the learned knowledge into the domain generalization feature extraction unit to establish an implicit connection with the instance target.

[0070] The predicted hierarchical similarity sequence may be the result obtained by performing self-attention calculation based on the preset hierarchical feature map and the preset feature sequence. For example, the construction system may determine the preset hierarchical similarity sequence based on the following formula by performing self-attention on the preset hierarchical feature map and the preset feature sequence:

[0071]

[0072] where, F i represents the frozen feature map of the i-th layer, that is, the preset hierarchical feature map of the i-th layer; T i represents the token sequence of the i-th layer; S i is the preset hierarchical similarity sequence of the i-th layer; d is the scaling factor. The value of F i ×T i significantly increases due to dimension accumulation, and the gradient significantly decreases after similarity calculation (softmax), resulting in the problem of gradient disappearance. Setting the scaling factor d can prevent gradient disappearance and stabilize gradient propagation.

[0073] In step S132, the construction system determines a preset hierarchical object query quantity based on the preset feature sequence using a multi-layer perceptron.

[0074] According to the example embodiment, the preset hierarchical object query volume can be a component of the object query volume corresponding to the preset hierarchy.

[0075] For example, the construction system can perform a linear mapping on T through a Multilayer Perceptron (MLP) i to further strengthen the connection relationship with the instance target. For example, the construction system can determine the preset hierarchical object query volume according to the following formula:

[0076]

[0077] where is the weight of the i-th layer; is the bias of the i-th layer; Q i is denoted as the object query volume of the i-th layer, that is, the preset hierarchical object query volume of the i-th layer.

[0078] In step S133, the construction system determines the preset hierarchical correction deviation volume according to the preset hierarchical similarity sequence and the preset hierarchical object query volume.

[0079] According to the example embodiment, the preset hierarchical correction deviation volume can be the deviation volume that needs to be corrected for the i-th layer frozen feature map (the i-th layer preset hierarchical feature map) during the fine-tuning process. For example, the construction system can determine the preset hierarchical correction deviation volume according to the following formula:

[0080] DG(F i ) = S i ×Q i ;

[0081] where DG(F i ) is the preset hierarchical correction deviation volume of the i-th layer.

[0082] In step S134, the construction system determines the preset hierarchical enhanced feature map according to the preset hierarchical correction deviation volume and the preset hierarchical feature map.

[0083] According to the example embodiment, the construction system can merge DG(F i ) with F i and obtain the preset hierarchical enhanced feature map F' i .

[0084] In step S135, the construction system determines the first multi-scale feature map according to the preset hierarchical enhanced feature map.

[0085] According to the example embodiment, the construction system can use F' iIt can be used as the input of the next layer of the domain generalization feature extraction unit, and then the (i + 1)-th layer of the preset hierarchical feature map F enhanced by the domain generalization feature extraction unit is obtained i+1 For example, the construction system can determine the (i + 1)-th layer of the preset hierarchical feature map F according to the following formula i+1 :

[0086] F i+1 = L i+1 (MLP(F i + DG(F i )));

[0087] Wherein, L i+1 is the (i + 1)-th layer; F i+1 is the (i + 1)-th layer of the (i + 1)-th layer of frozen feature map, that is, the (i + 1)-th layer of the preset hierarchical feature map

[0088] For example, referring to Figure 3 , the construction system can perform feature enhancement processing on the second layer feature map (the feature layer corresponding to the second layer feature map, i is 0) through the domain generalization feature extraction unit, and then output F'0. The output F'0 can be the input of the first layer of the preset hierarchical feature map F1 (that is, the feature layer corresponding to the fifth layer feature map, i is 1), and then perform feature enhancement processing to output F'1. The output F'1 can be the input of the second layer of the preset hierarchical feature map F2 (that is, the feature layer corresponding to the eighth layer feature map, i is 2), and then perform feature enhancement processing to output F'2. The output F'2 can be the input of the third layer of the preset hierarchical feature map F3 (that is, the feature layer corresponding to the eleventh layer feature map, i is 3), and then perform feature enhancement processing to output F'3. The output F'3 is the first multi-scale feature map C

[0089] After 4 layers of enhancement by the construction system, the first multi-scale feature map is finally output through the domain generalization feature extraction unit. The first multi-scale feature map can be the first multi-scale feature map generated by the domain generalization feature extraction unit through downsampling of the preset hierarchical feature map. During the downsampling process, the channel dimension of the preset hierarchical feature map remains unchanged, and the spatial dimension will decrease accordingly with the increase of the downsampling rate (8 times, 16 times, 32 times, 64 times), and the obtained multi-scale features are

[0090] Subsequently, the construction system performs 1×1 convolution and normalization operations on the multi-scale features to unify the number of feature map channels to 256, and finally obtains the first multi-scale feature map

[0091] In step S136, the construction system determines the domain generalization object query amount according to the preset hierarchical object query amount

[0092] According to the example embodiment, the construction system can determine the object query volume according to the following formula:

[0093]

[0094] wherein, is the domain generalization object query volume, The number of channels is 256; is the weight; is the bias; N is the number of preset hierarchical object query volumes.

[0095] Through the above embodiments, the construction method provided by the present application designs a domain generalization feature extraction unit in the feature extraction stage. On the premise of freezing the feature extraction parameters and feature layers, the domain generalization feature extraction unit is embedded, and a limited-scale dataset is used to fine-tune only the domain generalization feature extraction unit, further enhancing the multi-scale features generated by pre-training, so that the detection model effectively improves the performance of downstream object detection without increasing the computational overhead.

[0096] Optionally, referring to Figure 4 , step S140 may include steps S141 - S145.

[0097] In step S141, the construction system determines multi-scale position information according to the position information corresponding to the first multi-scale feature map, and the multi-scale position information is a set of position information corresponding to the first-scale feature map.

[0098] According to the example embodiment, the construction system can provide position information through position encoding (PE) so that the detection model can adapt to objects of different scales.

[0099] Position encoding represents the feature information of different spatial positions in the image through sine and cosine functions. For a feature point, the construction system can determine its position encoding according to the following formula:

[0100]

[0101] wherein, (p x , p y ) represents the coordinate position of the feature point in the remote sensing image, d p is the dimension of the position encoding, related to the hidden layer dimension, usually set to 256. i p is the encoding dimension index in the x direction; j p is the encoding dimension index in the y direction.

[0102] The construction system can provide multi-scale position information for the first multi-scale feature map by calculating the position encoding and superimposing randomly initialized and learnable scale-level embedding vectors. Among them, the feature points of each level share an embedding vector. The dimension of the embedding vector is the same as the channel dimension of the feature map (256), so that the position encoding and the embedding vector can be directly added. The embedding vector is randomly initialized in a way similar to the initialization of the weight matrix in a neural network, and its values may be floating-point numbers that follow a normal distribution.

[0103] In step S142, the construction system determines the second multi-scale feature map based on the first multi-scale feature map and the multi-scale position information.

[0104] According to the exemplary embodiment, the second multi-scale feature map is generated by superimposing the position encoding on each layer of features of the first multi-scale feature map, and the second multi-scale feature map has the same dimension as the first multi-scale feature map.

[0105] The construction system assigns a randomly generated and learnable level embedding vector to each level. The construction system superimposes the position encoding and the randomly initialized and learnable scale-level embedding vector on each layer of features, and then generates a multi-scale feature map containing multi-scale position information (that is, generates the second multi-scale feature map C').

[0106] In step S143, the construction system determines the third multi-scale feature map based on the encoder group of the multi-scale deformable attention mechanism according to the second multi-scale feature map.

[0107] According to the exemplary embodiment, the encoder group may include at least two encoders, and each encoder includes a multi-scale deformable attention module (Multi-scale Deformable Attention, MSDA) and a feed-forward connection layer (Feed-Forward Networks, FFN).

[0108] The construction system can perform multi-scale deformable attention processing and feed-forward processing on the second multi-scale feature map through the multi-scale deformable attention module and the feed-forward connection layer in the encoder, so as to determine the third multi-scale feature map.

[0109] Optionally, referring to Figure 5 , step S143 may include step S1431 and step S1432.

[0110] The construction system executes step S1431 at least twice in a loop. In step S1431, the construction system executes the step of determining the feed-forward multi-scale feature map. Referring to Figure 6 , step S1431 may include step S1431a and step S1431b.

[0111] In step S1431a, the construction system determines an enhanced multi-scale feature map based on a multi-scale deformable attention sub-mechanism and according to the second multi-scale feature map.

[0112] According to an exemplary embodiment, the construction system may determine an enhanced multi-scale feature map based on a multi-scale deformable attention module of an encoder and according to the second multi-scale feature map.

[0113] For example, referring to Figure 7 , the query vector z q determines k sampling points through sparse spatial sampling on the feature layer, interacts with the local region features formed by these k points, calculates the attention weights of the multi-scale deformable attention sub-mechanism, and applies them to the corresponding feature values. Since the sampling positions can be dynamically adjusted through learnable offsets, it avoids the long transition from the global range to the local region for gradual learning, enables the detection model to more efficiently focus on the key regions with real detection significance, and reduces the computational amount.

[0114] The construction system may determine the enhanced multi-scale feature map according to the following formula:

[0115]

[0116] where m is the attention head, with a value of 8; x l is the feature map of different layers, L has a value of 4; z q is the query vector feature generated by linearly transforming x l , the position information is represented by the normalized reference point coordinates and is mapped to the feature maps of different scales through ; the number of sampling points K has a value of 4; Δp mlqk represents the position offset of the k-th sampling point in the l-th feature layer relative to the reference point; A mlqk represents the attention weight of the k-th sampling point in the l-th feature layer; is the feature value after interpolation calculation of the sampling points generated by linearly transforming W' m , then the normalized attention weight A mlqk is used to weight the feature values, and finally W m is used for linear transformation to obtain the output results of different attention heads. D ms is the output of the encoder multi-scale deformable attention module (i.e., the enhanced multi-scale feature map), representing the interaction result between the query vector z q and the k key points selected on the feature x l .

[0117] In step S1431b, the construction system determines a feed-forward multi-scale feature map based on a feed-forward sub-mechanism and according to the enhanced multi-scale feature map.

[0118] According to the exemplary embodiment, the attention result (i.e., the enhanced multi-scale feature map) output by the construction system after being processed by the encoder multi-scale deformable attention module will be fed into the encoder feed-forward connection layer for channel dimension expansion and compression to extract more expressive features. The feed-forward connection layer consists of two fully-connected layers, where the output of the first layer has a higher dimension than the input, and after passing through an activation function (such as ReLU), it returns to the original dimension. A residual connection and layer normalization are introduced in each encoder multi-scale deformable attention module and feed-forward connection layer to alleviate the problem of vanishing gradients in deep networks and ensure stable module training.

[0119] According to the exemplary embodiment, the encoder group may include six stacked encoders. The construction system can stack six encoders, and in each encoder layer, the above steps S1431a and S1431b will be repeated. The encoding process of each layer will use the encoded features output by the previous layer (i.e., the encoded features after residual connection and layer normalization processing of the feed-forward pair of scale feature maps) as input. The advantage of stacking is to enhance the feature expression ability layer by layer, enabling the detection model to better capture multi-scale information in remote sensing images while maintaining computational efficiency.

[0120] In step S1432, the construction system determines the third multi-scale feature map based on the feed-forward multi-scale feature map.

[0121] According to the exemplary embodiment, the third multi-scale feature map may be the multi-scale feature map finally output after stacking multiple encoders. The feature dimension of the third multi-scale feature map is the same as that of the first multi-scale feature map.

[0122] For example, the construction system can finally output the third multi-scale feature map through an encoder group stacked by six layers. [[ID=1,5]]

[0123] In step S144, the construction system determines the fused object query amount based on the initial object query amount and the domain generalization object query amount.

[0124] According to the exemplary embodiment, the construction system can determine the fused object query amount through the encoder-decoder unit according to the following formula:

[0125]

[0126] where Q0 is the randomly initialized object query amount, and Q' is the fused object query amount.

[0127] In step S145, the construction system determines the fused object query amount set and the reference coordinates based on the decoder group of the scale variability attention mechanism, according to the third multi-scale feature map and the fused object query amount.

[0128] According to an exemplary embodiment, the decoder group may include at least two decoders. The decoder includes a self-attention module (SA), a multi-scale deformable attention module, and a feed-forward connection layer.

[0129] The construction system may determine a set of fused object query amounts and reference coordinates according to the third multi-scale feature map and the fused object query amount through the multi-scale deformable attention module and the feed-forward connection layer in the decoder.

[0130] See Figure 8 , step S145 may include step S1451 and step S1452.

[0131] The construction system executes step S1451 at least twice in a loop. In step S1451, the construction system executes the step of determining the feed-forward feature information of the fused object query amount. See Figure 9 , step S1451 includes step S1451a - step S1451c.

[0132] In step S1451a, the construction system determines the self-attention object query amount according to the fused object query amount based on the self-attention sub-mechanism.

[0133] For example, see Figure 10 , the construction system may perform self-attention calculation on the fused object query amount through the self-attention module of the decoder. The construction system may perform self-attention calculation on the fused object query amount according to the following formula:

[0134]

[0135] Among them, Q' is the input sequence for self-attention calculation of the fused object query amount, and through the linear mapping (weight of Q q ), (weight of K q ), and (weight of V q ), the three vectors Q q , K q , and V q are obtained respectively. is the scaling factor for self-attention calculation of the fused object query amount. Obtain the similarity between Q q and K q . After weighted aggregation of V q , the self-attention result is obtained. Specifically, Q q represents the feature or target that needs to be focused on, K q represents all positions of the input features, V q carries the feature information of each position, and each Kq There is a corresponding V q . V q does not directly participate in the calculation of similarity Q q K q T in the calculation, but is weighted and aggregated according to the calculated attention weights .

[0136] The role of the self-attention module is to allow each fusion object query volume to interact with other fusion query volumes, thereby enhancing the correlation between fusion object queries and removing duplicate detection box information.

[0137] In step S1451b, the construction system determines the feature information of the fusion object query volume based on the multi-scale deformable attention sub-mechanism, according to the multi-scale feature maps output by the encoder group and the self-attention object query volume.

[0138] According to the exemplary embodiment, the construction system can use the multi-scale deformable attention module of the decoder to take each fusion object query volume as the query vector Q d , the encoded feature (i.e., the third multi-scale feature map) as the value vector V d and the key vector K d to perform multi-scale deformable attention calculation. (For example, the third multi-scale feature map corresponds to Figure 10 the V in d , the third multi-scale feature map corresponds to Figure 10 the K in d , and the fusion object query volume corresponds to Figure 10 the Q in d ). The construction system can focus on a small number of sampling points through the multi-scale deformable attention module of the decoder, focus on important local regions, calculate the attention weights, perform weighted summation on the local regions of the third multi-scale feature map, and extract the feature information related to the target to be detected (i.e., the feature information of the fusion object query volume).

[0139] In step S1451c, the construction system determines the feed-forward feature information of the fusion object query volume based on the feed-forward sub-mechanism, according to the feature information of the fusion object query volume.

[0140] According to the exemplary embodiment, the construction system can further process the features at each position through the feed-forward connection layer of the decoder, further strengthen the feature representation, and extract high-level feature expressions (i.e., obtain the feed-forward feature information of the fusion object query volume). The multi-scale deformable attention module and the feed-forward connection layer of the decoder can also introduce residual connections and layer normalization.

[0141] The decoder group may include an encoder stacked in six layers. The construction system may stack six layers of decoders, and the above steps S1451a - S1451c are repeated for each layer of the decoder. That is, the construction system may stack six layers of decoder groups to process the interaction between the query vector (i.e., the fused object query volume) and the encoder features (i.e., the third multi-scale feature map) layer by layer in each decoder and strengthen the feature representation through the feed-forward sub-mechanism. The feed-forward sub-mechanism may be a Feed-Forward Network (FFN).

[0142] In step S1452, the construction system determines the fused object query volume set and the reference coordinates according to the feed-forward feature information of the fused object query volume.

[0143] According to the exemplary embodiment, the construction system may finally output a fused object query volume set {Q'} ∈ R 6×300×256 and the corresponding reference coordinates Ref ∈ R 6×300×4 . Each fused object query volume in the fused object query volume set {Q'} ∈ R 6 ×300×256 represents a potential target.

[0144] Through the above embodiments, the construction method provided by the present application adopts a multi-scale deformable attention mechanism in the encoder-decoder stage to reduce the resource waste caused by global interaction calculations. At the same time, an initialization scheme for the fused object query volume of the decoder is designed to optimize the initial distribution of the fused object query volume to be close to the distribution in the later stage of training, accelerating the convergence of the detection model and improving the generalization ability of the detection model in different scenarios.

[0145] The inventors also found that the Intersection over Union (IoU) is obtained by calculating the ratio of the intersection and union of two bounding boxes, resulting in a value between 0 and 1, which is used to measure the overlap degree between the predicted box and the true annotation box. As a key metric in object detection tasks, a series of bounding box regression loss functions have been developed based on IoU, such as Generalized IoU (GIoU), etc. However, different from natural images, remote sensing images have characteristics such as large width and small target size. Bounding box regression loss functions represented by IoU are extremely sensitive to target size and position deviation. In the case of the same pixel deviation, the overlapping area between the two boxes significantly decreases for small targets, resulting in a much higher decrease in GIoU than for large targets, thereby affecting the backpropagation of small target feature information in the network and causing a significant decline in the learning ability of some detection models for small targets.

[0146] See Figure 11 , step S150 may include steps S151 - S153.

[0147] In step S151, the construction system determines the loss function of the target detection head unit according to the preset parameters of the target detection head unit. The loss function includes a bounding box regression loss function and a classification loss function.

[0148] According to the exemplary embodiment, the loss function of target detection is divided into a classification loss and a bounding box regression loss. The former is used to measure the difference between the model prediction category and the true label, and the latter is used to quantify the deviation between the predicted box and the true annotation box.

[0149] Since the Gaussian distribution satisfies that the peak is located at the center point μ and the function value gradually decreases away from the center point. To strengthen the influence of the pixels in the neighborhood of the center point on the regression process and weaken the interference of the pixels far from the center point, the predicted box is modeled as a Gaussian distribution X p ~N p (μ p ,ε p ), the target annotation box is modeled as a Gaussian distribution X t ~N t (μ t ,ε t ), and the Kullback-Leibler divergence (KL divergence) is used to measure the similarity between the two Gaussian distributions, obtaining a smooth bounding box regression loss function to replace the commonly used IoU-like loss function. The construction system can determine the loss function improved based on the KL divergence according to the following formula:

[0150]

[0151] Loss KL =1 - 1 / [1 + log(D KL (N p ‖N t ))];

[0152] Where, Loss KL is the loss function improved based on the KL divergence; μ p is the mean of the predicted box, ε p is the variance of the predicted box; μ t is the mean of the true annotation box; ε t is the variance of the true annotation box; D KL is the K-L divergence; D KL (N p ‖N t ) is the K-L divergence calculated after modeling the predicted box and the target annotation box as Gaussian distributions; D KL (N p ‖N t ) represents the K-L divergence calculated after modeling the two boxes as Gaussian distributions.

[0153] X p is a random variable of the prediction box, N p is the Gaussian distribution of the prediction box, X p ~N p means that after the prediction box is modeled as the Gaussian distribution N p X p obeys N p .

[0154] X t is a random variable of the target annotation box, N t is the Gaussian distribution of the target annotation box, X t ~N t means that after the target annotation box is modeled as the Gaussian distribution N t X t obeys N t .

[0155] According to the example embodiment, the construction system can sum the loss function Loss improved based on the KL divergence KL and the L1 loss function Loss L1 to obtain the final bounding box regression loss function Lossb box . Since Loss L1 is insensitive to outliers, combining it with Loss KL can accelerate the convergence speed of the model. The construction system can determine the bounding box regression loss function according to the following formula:

[0156] Loss L1 =|pred - target|;

[0157] Loss bbox =Loss L1 +Loss KL ;

[0158] where target represents the position of the target annotation box; pred represents the position of the prediction box. Loss L1 is the L1 loss function; Loss bbox is the bounding box regression loss function.

[0159] According to the example embodiment, the classification loss uses the focal loss (Focal - Loss) to weight the cross - entropy loss log(p t ), effectively alleviating the problem of positive - negative sample imbalance existing in the cross - entropy loss. The construction system can determine the classification loss using the focal loss function according to the following formula:

[0160] Loss class =-(1 - p t ) γlog(p t );

[0161]

[0162] where p is the probability of being classified as a positive sample; 1 - p is the probability of being classified as a negative sample; p t is the predicted probability of the correct class, and the closer p t is to 1, the higher the prediction confidence of the target detection head unit for the true class; γ is the focal parameter; the adjustment factor (1 - p t ) γ , γ ≥ 0, is used to balance positive and negative samples; Loss class is the classification loss using the focal loss function.

[0163] According to the example embodiment, the construction system can determine the loss function of the target detection head unit according to the following formula:

[0164] Loss = λ1Loss class + λ2Loss L1 + λ3Loss KL ;

[0165] where λ1 is the weight corresponding to the classification loss using the focal loss function; λ2 is the weight corresponding to the L1 loss function; λ3 is the weight corresponding to the loss function improved based on the KL divergence; Loss is the loss function.

[0166] In step S152, the construction system determines the optimization factor of the target detection head unit according to the loss function.

[0167] According to the example embodiment, the optimization factor can be a parameter for improving the performance of the target detection head unit. For example, the optimization factor can include the weight corresponding to the classification loss using the focal loss function, the weight corresponding to the L1 loss function, and the weight corresponding to the loss function improved based on the KL divergence, etc. The construction system can set the IoU threshold to distinguish positive and negative samples, which is used to guide the detection model to perform effective loss calculation. For example, if the IoU is greater than the threshold (set to 0.5), the prediction box and the ground truth box match, and it is regarded as a positive sample, and the classification loss and regression loss are calculated for this box; if the IoU is less than the threshold, the prediction box is regarded as a negative sample, and only the classification loss is calculated to ensure that the negative samples are correctly classified as the background.

[0168] According to the example embodiment, the construction system can set λ1 of the optimization factor to 2, λ2 to 5, and λ3 to 1 when the loss function takes the minimum value.

[0169] In step S153, the construction system determines the detection box result data through the target detection head unit according to the optimization factor of the target detection head unit, the set of fusion object query quantities, and the reference coordinates, so as to construct a detection model.

[0170] According to the exemplary embodiment, after determining the optimization factor, the construction system can determine the detection box result data according to the set of fusion object query quantities and the reference coordinates. For example, the construction system can use the AdamW optimizer to optimize the training of the detection model, set the initial learning rate to 1×10 -4 , and after 40 rounds of training, reduce the current learning rate to 10% of the original, and end the training, thereby constructing a detection model.

[0171] Through the above embodiments, the construction method provided by the present application improves the regression loss function in the target detection stage to adapt to the situation of small target sizes in remote sensing images, effectively avoiding problems such as the existing commonly used intersection over union loss function being unable to perform backpropagation when the gradient is zero, resulting in the detection model being unable to converge.

[0172] According to an aspect of the present application, the present application provides a detection model 200 for domain-generalized remote sensing images. The detection model 200 is constructed by the above-mentioned construction method 1000. Refer to Figure 12 , the detection model 200 includes a pre-trained feature extraction unit 210, a domain-generalized feature extraction unit 220, an encoder-decoder unit 230, and a target detection head unit 240.

[0173] The pre-trained feature extraction unit 210 is used to extract a remote sensing target detection data set to determine a preset hierarchical feature map, wherein the remote sensing target detection data set is determined by the acquired remote sensing images and environmental algorithms.

[0174] The domain-generalized feature extraction unit 220 is used to determine the domain-generalized object query quantity and the first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and the preset feature sequence.

[0175] The encoder-decoder unit 230 has a multi-scale deformable attention mechanism and is used to determine the set of fusion object query quantities and the reference coordinates corresponding to the fusion object query quantities according to the domain-generalized object query quantity and the first multi-scale feature map.

[0176] The target detection head unit 240 is used to determine the detection box result data according to the set of fusion object query quantities and the reference coordinates.

[0177] According to the exemplary embodiment, the remote sensing target detection dataset, the preset hierarchical feature map, the preset feature sequence, the domain generalization object query quantity, the first multi-scale feature map, the fusion object query quantity set, the reference coordinates corresponding to the fusion object query quantity, and the detection box result data have been described in the above construction method 1000, and thus will not be elaborated here.

[0178] According to an aspect of the present application, the present application provides a detection method 3000 for a detection model based on domain generalization remote sensing images, and the detection model is constructed by the above construction method 1000. The detection method 3000 can be executed by a detection system. Exemplarily, the detection system can be a host (or server) with data processing capabilities.

[0179] See Figure 13 , the detection method 3000 can include step S310 - step S340.

[0180] In step S310, the detection system extracts the remote sensing target dataset to be detected to determine the preset hierarchical feature map.

[0181] According to the exemplary embodiment, the remote sensing target dataset to be detected can be a remote sensing target dataset in a complex environment such as fog interference and low light interference.

[0182] In step S320, the detection system determines the domain generalization object query quantity and the first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and the preset feature sequence.

[0183] In step S330, the detection system determines the fusion object query quantity set and the reference coordinates corresponding to the fusion object query quantity according to the domain generalization object query quantity and the first multi-scale feature map.

[0184] In step S340, the detection system determines the detection box result data according to the fusion object query quantity set and the reference coordinates.

[0185] According to the exemplary embodiment, the preset hierarchical feature map, the preset feature sequence, the domain generalization object query quantity, the first multi-scale feature map, the fusion object query quantity set, the reference coordinates corresponding to the fusion object query quantity, and the detection box result data have been described in the above construction method 1000, and thus will not be elaborated here.

[0186] The process by which the detection system determines the detection box result data is similar to the process of determining the detection box result data in the above construction method 1000, and thus will not be elaborated here.

[0187] The detection method provided by the present application can cope with cross-domain distribution differences and achieve high-performance detection of remote sensing targets in complex scenarios without a large amount of target domain data.

[0188] According to one aspect of the present application, the present application provides a detection system 400 for a detection model of domain generalization remote sensing images. The detection model is constructed by the above-mentioned construction method 1000. Refer to Figure 14 , the detection system 400 includes an image input unit 410 and a detection model 200. The detection model 200 includes a pre-trained feature extraction unit 210, a domain generalization feature extraction unit 220, an encoder-decoder unit 230, and an object detection head unit 240.

[0189] The image input unit 410 acquires a remote sensing target data set to be detected.

[0190] The pre-trained feature extraction unit 210 extracts the remote sensing target data set to be detected to determine a preset hierarchical feature map.

[0191] The domain generalization feature extraction unit 220 determines a domain generalization object query quantity and a first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and a preset feature sequence.

[0192] The encoder-decoder unit 230 has a multi-scale deformable attention mechanism. The encoder-decoder unit 230 determines a set of fused object query quantities and reference coordinates corresponding to the fused object query quantities according to the domain generalization object query quantity and the first multi-scale feature map.

[0193] The object detection head unit 240 determines detection box result data according to the set of fused object query quantities and the reference coordinates.

[0194] According to the exemplary embodiment, the preset hierarchical feature map, the preset feature sequence, the domain generalization object query quantity, the first multi-scale feature map, the set of fused object query quantities, the reference coordinates corresponding to the object query quantity, and the detection box result data have been described in the above-mentioned construction method 1000, and thus will not be elaborated herein.

[0195] According to another aspect of the present application, the present application further provides a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the construction method of the detection model of domain generalization remote sensing images as described above.

[0196] According to another aspect of the present application, the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors can implement the construction method of the detection model of domain generalization remote sensing images as described above.

[0197] According to another aspect of the present application, the present application further provides a computer program product, including: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, when the program instructions are executed by a computer, enabling the computer to execute the method for constructing the detection model of the domain-generalized remote sensing image as described above.

[0198] According to another aspect of the present application, the present application further provides a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it can implement the detection method of the detection model based on the domain-generalized remote sensing image as described above.

[0199] According to another aspect of the present application, the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the detection method of the detection model based on the domain-generalized remote sensing image as described above.

[0200] According to another aspect of the present application, the present application further provides a computer program product, including: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, when the program instructions are executed by a computer, enabling the computer to execute the detection method of the detection model based on the domain-generalized remote sensing image as described above.

[0201] Finally, it should be noted that the above are only the preferred embodiments of the present application and are not used to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions of the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for constructing a detection model of domain-generalized remote sensing images, characterized in that, The detection model is used to detect the target objects in remote sensing images, and the construction method includes: Determine the remote sensing target detection dataset through the obtained remote sensing image dataset and environmental algorithms; Determine the preset hierarchical feature map according to the remote sensing target detection dataset; Determine the domain generalization object query quantity and the first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and the preset feature sequence corresponding to the preset hierarchical feature map; Determine the fusion object query quantity set and the reference coordinates corresponding to the fusion object query quantity set according to the domain generalization object query quantity and the first multi-scale feature map; Determine the detection box result data according to the fusion object query quantity set and the reference coordinates to construct the detection model.

2. The construction method according to claim 1, wherein The step of determining the domain generalization object query quantity and the first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and the preset feature sequence corresponding to the preset hierarchical feature map includes: Determine the preset hierarchical similarity sequence according to the preset hierarchical feature map and the preset feature sequence corresponding to the preset hierarchical feature map; Based on a multi-layer perceptron, determine the preset hierarchical object query quantity according to the preset feature sequence; Determine the preset hierarchical correction deviation quantity according to the preset hierarchical similarity sequence and the preset hierarchical object query quantity; Determine the preset hierarchical enhanced feature map according to the preset hierarchical correction deviation quantity and the preset hierarchical feature map; Determine the first multi-scale feature map according to the preset hierarchical enhanced feature map; Determine the domain generalization object query quantity according to the preset hierarchical object query quantity.

3. The construction method according to claim 1, wherein The step of determining the fusion object query quantity set and the reference coordinates corresponding to the fusion object query quantity set according to the domain generalization object query quantity and the first multi-scale feature map includes: Determine the multi-scale position information according to the position information corresponding to the first multi-scale feature map, and the multi-scale position information is a set of position information corresponding to the first-scale feature map; Determine the second multi-scale feature map according to the first multi-scale feature map and the multi-scale position information; Based on the encoder group of the multi-scale deformable attention mechanism, determine the third multi-scale feature map according to the second multi-scale feature map; Determine the fusion object query quantity according to the initial object query quantity and the domain generalization object query quantity; Based on the decoder group of the scale variability attention mechanism, determine the fusion object query quantity set and the reference coordinates according to the third multi-scale feature map and the fusion object query quantity.

4. The construction method according to claim 3, characterized in that, The multi-scale deformable attention mechanism includes a multi-scale deformable attention sub-mechanism and a feed-forward sub-mechanism; Based on the encoder group of the multi-scale deformable attention mechanism, the step of determining the third multi-scale feature map according to the second multi-scale feature map includes: Execute the feed-forward multi-scale feature map determination step at least twice in a loop, and the feed-forward multi-scale feature map determination step includes: Based on the multi-scale deformable attention sub-mechanism, determine the enhanced multi-scale feature map according to the second multi-scale feature map; Based on the feed-forward sub-mechanism, determine the feed-forward multi-scale feature map according to the enhanced multi-scale feature map; Determine the third multi-scale feature map according to the feed-forward multi-scale feature map.

5. The construction method according to claim 3, characterized in that The multi-scale deformable attention mechanism includes a self-attention sub-mechanism, a multi-scale deformable attention sub-mechanism, and a feed-forward sub-mechanism; The decoder group based on the scale-variable attention mechanism determines the set of fused object query amounts and the reference coordinates according to the third multi-scale feature map and the fused object query amount, including: At least loop and execute the feed-forward feature information determination step of the fused object query amount twice. The feed-forward feature information determination step of the fused object query amount includes: Based on the self-attention sub-mechanism, determine the self-attention object query amount according to the fused object query amount; Based on the multi-scale deformable attention sub-mechanism, determine the feature information of the fused object query amount according to the third multi-scale feature map output by the encoder group and the self-attention object query amount; Based on the feed-forward sub-mechanism, determine the feed-forward feature information of the fused object query amount according to the feature information of the fused object query amount; Determine the set of fused object query amounts and the reference coordinates according to the feed-forward feature information of the fused object query amount.

6. The construction method according to claim 1, characterized in that, The determining the detection box result data according to the set of fused object query amounts and the reference coordinates to construct the detection model includes: According to the preset parameters of the target detection unit, determine the loss function of the target detection head unit. The loss function includes a bounding box regression loss function and a classification loss function; Determine the optimization factor of the target detection head unit according to the loss function; Through the target detection head unit, determine the detection box result data according to the optimization factor of the target detection head unit, the set of fused object query amounts, and the reference coordinates to construct the detection model.

7. A detection model for domain-generalized remote sensing images, characterized in that, The detection model is constructed by the construction method according to any one of claims 1-6. The detection model includes: A pre-trained feature extraction unit for extracting a remote sensing target detection data set to determine a preset hierarchical feature map, where the remote sensing target detection data set is determined by the acquired remote sensing image and environmental algorithm; A domain generalization feature extraction unit for determining a domain generalization object query amount and a first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and a preset feature sequence; An encoder-decoder unit with a multi-scale deformable attention mechanism for determining a set of fused object query amounts and reference coordinates corresponding to the fused object query amounts according to the domain generalization object query amount and the first multi-scale feature map; A target detection head unit for determining detection box result data according to the set of fused object query amounts and the reference coordinates.

8. A detection method for a detection model of domain generalization remote sensing images, characterized in that, The detection model is constructed by the construction method according to any one of claims 1-6. The detection method includes: Extract the remote sensing target data set to be detected to determine a preset hierarchical feature map; Determine a domain generalization object query amount and a first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and a preset feature sequence; Determine a set of fused object query quantities and corresponding reference coordinates based on the domain generalization object query quantity and the first multi-scale feature map; Determine the detection box result data based on the set of fused object query quantities and the reference coordinates.

9. A detection system for a detection model of domain generalization remote sensing images, characterized in that, The detection model is constructed by the construction method according to any one of claims 1-6, and the detection system includes: An image input unit for acquiring a remote sensing target data set to be detected; A pre-trained feature extraction unit for extracting the remote sensing target data set to be detected to determine a preset hierarchical feature map; A domain generalization feature extraction unit for determining a domain generalization object query quantity and a first multi-scale feature map of the preset hierarchical feature map according to the preset hierarchical feature map and a preset feature sequence; An encoder-decoder unit with a multi-scale deformable attention mechanism for determining a set of fused object query quantities and corresponding reference coordinates according to the domain generalization object query quantity and the first multi-scale feature map; A target detection head unit for determining the detection box result data according to the set of fused object query quantities and the reference coordinates.

10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the construction method of the detection model for domain generalization remote sensing images according to any one of claims 1-6.

11. An electronic device, characterized in that, Including: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the construction method of the detection model for domain generalization remote sensing images according to any one of claims 1-6.

12. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the detection method of the detection model based on domain generalization remote sensing images according to claim 8.

13. An electronic device, characterized in that, Including: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the detection method of the detection model based on domain generalization remote sensing images according to claim 8.

Citation Information

Cited By

  • Construction method, device and equipment of model for remote sensing image target detection

    CN121789067A