Target detection method and device for multi-scale remote sensing image

By integrating multi-scale remote sensing image datasets and adjusting the backbone network to construct a candidate backbone network, the problems of insufficient remote sensing image datasets and single scale are solved, and the accuracy and generalization performance of the target detection model are improved.

CN120747751APending Publication Date: 2025-10-03JIMEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510916512.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing remote sensing image datasets for marine environment monitoring suffer from insufficient data volume and single scale, resulting in poor accuracy of target detection models.

Method used

By integrating multiple remote sensing image datasets of different types, a target detection model with a codec structure is constructed, and the backbone network is adjusted using the state space model and replaced with a candidate backbone network to form an updated detection model.

Benefits of technology

The generalization performance and accuracy of the target detection model are improved, and the detection capability of remote sensing image targets is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747751A_ABST
    Figure CN120747751A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, in particular to a multi-scale remote sensing image target detection method and device, and the method comprises the steps: integrating a plurality of different types of remote sensing image data sets, and obtaining an integrated data set; constructing a target detection model of a coding and decoding structure based on the integrated data set; performing backbone network adjustment on the target detection model based on a state space model, and constructing a candidate backbone network; and replacing the initial backbone network in the target detection model with the candidate backbone network to obtain an updated detection model for performing target detection on the to-be-detected remote sensing image. According to the method, the target detection network is constructed through the rich integrated data set, and the initial backbone network is replaced by the candidate backbone network with relatively high precision, so that the generalization performance of the updated detection model can be improved, and the precision of the updated detection model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular to a method and device for detecting targets in multi-scale remote sensing images in the field of machine learning technology. Background Art

[0002] In recent years, with the advancement of science and technology and the need for development, people have become increasingly interested in marine environmental monitoring, marine resource exploration, and various marine disasters. Based on this, a full-ocean dataset has been proposed. Given the complex and ever-changing marine environment, the selection of image types for the dataset is crucial. Synthetic Aperture Radar (SAR) remote sensing offers unique advantages in marine environmental monitoring, as it provides high-resolution two-dimensional images unaffected by sunlight, cloud cover, and weather conditions. The proposed full-ocean dataset uses SAR images as its dataset. Specifically, the full-ocean dataset can be divided into marine and underwater datasets. Current SAR marine datasets have relatively few images and lack large datasets, leading to overfitting in small datasets. Underwater SAR datasets also face similar challenges. Furthermore, underwater SAR datasets are of a single scale and lack multi-scale datasets, resulting in poor accuracy in object detection models trained using these datasets. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and device for detecting targets in multi-scale remote sensing images. The technical solutions adopted are as follows:

[0004] In a first aspect, an embodiment of the present invention provides a method for detecting objects in multi-scale remote sensing images, the method comprising:

[0005] Integrate multiple remote sensing image datasets of different types to obtain an integrated dataset;

[0006] Based on the integrated data set, construct an object detection model with a codec structure;

[0007] Adjusting the backbone network of the target detection model based on the state space model to construct a candidate backbone network;

[0008] The initial backbone network in the target detection model is replaced by the candidate backbone network to obtain an updated detection model for performing target detection on the remote sensing image to be detected.

[0009] In a second aspect, an embodiment of the present invention provides a device for detecting an object in a multi-scale remote sensing image, the device comprising:

[0010] An integration module is used to integrate multiple remote sensing image datasets of different types to obtain an integrated dataset;

[0011] A first construction module is used to construct an object detection model of a codec structure based on the integrated data set;

[0012] A second construction module is used to adjust the backbone network of the target detection model based on the state space model to construct a candidate backbone network;

[0013] A replacement module is used to replace the initial backbone network in the target detection model with the candidate backbone network to obtain an updated detection model for performing target detection on the remote sensing image to be detected.

[0014] In a third aspect, a computer program product is provided, comprising: a computer program code, which, when executed on a computer, enables the computer to execute the method in the first aspect or any one of the possible implementations of the first aspect.

[0015] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code runs on a computer, the computer executes the method in the first aspect or any possible implementation method described in the first aspect.

[0016] The present invention has the following beneficial effects: by integrating multiple remote sensing image data sets of different types, a relatively rich integrated data set can be obtained; then, based on the integrated data set, a target detection model with a codec structure is constructed, and the backbone network of the target detection model is adjusted using a state-space model to construct a candidate backbone network; thus, the accuracy of the candidate backbone network can be improved; finally, the initial backbone network in the target detection model is replaced with the candidate backbone network, thereby obtaining an updated detection model for target detection in the remote sensing image to be detected. In this way, by constructing a target detection network using the rich integrated data set and replacing the initial backbone network with the candidate backbone network with higher accuracy, the generalization performance of the updated detection model can be improved, thereby improving the accuracy of the updated detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is a schematic diagram of an implementation flow of a method for detecting a target in a multi-scale remote sensing image provided by an embodiment of the present invention;

[0019] Figure 2 This is another implementation flowchart of a method for detecting targets in multi-scale remote sensing images provided by an embodiment of the present invention;

[0020] Figure 3 is a schematic diagram of the composition structure of a candidate backbone network provided by an embodiment of the present invention;

[0021] Figure 4 1 is a schematic diagram of the composition structure of the updated detection model provided by an embodiment of the present invention;

[0022] Figure 5 Schematic diagram of the structure of the hybrid module provided by an embodiment of the present invention;

[0023] Figure 6 2 is a schematic diagram of the composition structure of the residual module provided in an embodiment of the present invention;

[0024] Figure 7 is a schematic diagram of the detection results of the updated detection model provided by an embodiment of the present invention;

[0025] Figure 8 1 is a schematic diagram of the structure of a multi-scale remote sensing image target detection device provided by an embodiment of the present invention;

[0026] Figure 9 It is a structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] To further illustrate the technical means and effectiveness of the present invention in achieving its intended objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of a multi-scale remote sensing image target detection method proposed by the present invention. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0028] In the description of the embodiments of the present invention, unless otherwise specified, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present invention, "multiple" refers to two or more than two.

[0029] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.

[0030] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0031] The following describes in detail a method for detecting targets in multi-scale remote sensing images provided by the present invention in conjunction with the accompanying drawings. Figure 1 , which shows a schematic diagram of an implementation flow of a method for target detection in a multi-scale remote sensing image provided by an embodiment of the present invention, the method comprising:

[0032] 101, integrating multiple remote sensing image datasets of different types to obtain an integrated dataset.

[0033] Here, the datasets of different scenarios are integrated to form a large dataset, namely, the integrated dataset. Specifically, the integrated dataset includes: a marine SAR dataset and an underwater SAR dataset, and its type is: ship.

[0034] 102. Construct an object detection model of a codec structure based on the integrated data set.

[0035] Here, the initial detection model is trained using the integrated dataset to obtain the target detection model.

[0036] In some possible implementations, the above step 102 can be performed by Figure 2 The steps shown achieve:

[0037] 201, based on the initial backbone network, encoder-decoder and prediction head network, build the initial detection model.

[0038] Here, an initial detection model based on the integrated dataset is established, which mainly consists of three parts: the initial backbone network, the encoder-decoder, and the prediction head.

[0039] 202. Train the initial detection model based on the integrated data set to obtain the target detection model.

[0040] In some possible implementations, the initial backbone network, codec, and prediction head network in the initial detection model are trained using the integrated dataset to obtain an object detection model. That is, the above step 202 can be implemented by the following steps 221 to 225 (not shown):

[0041] 221 , using the initial backbone network of the initial detection model to perform feature extraction on the integrated data set to obtain feature information.

[0042] Here, the initial backbone network extracts feature information from the input image, and the Encoder-Decoder performs higher-dimensional feature extraction on the output feature information of the initial backbone network and passes the predicted image feature information to the prediction head for image output. Specifically, the initial backbone network is responsible for extracting features from the input image to obtain feature information.

[0043] 222. Perform embedding processing on the feature information to obtain processed features.

[0044] Here, for the feature information output by the initial backbone network, an embedding operation is first performed, that is, the feature information output by the initial backbone network is mapped into the form of a vector. The embedding operation can be divided into input embedding and position embedding.

[0045] In some possible implementations, the feature information is first subjected to input embedding to obtain a continuous vector; for example, input embedding maps the symbols of the input sequence to a continuous vector representation. Next, the feature information is subjected to position embedding to obtain position information corresponding to the continuous vector; for example, position embedding is a set of vectors of the same dimension as the input vector after the embedding operation, used to provide position information. Finally, the continuous vector is fused with the position information to obtain the processed feature. For example, the position embedding and the input embedding are added together to obtain the processed feature, which is then used as the input of the encoder.

[0046] 223 , use the codec of the initial detection model to encode and decode the processed features to obtain multi-scale sample decoding data.

[0047] Here, the processed features are used as the input to the encoder in the codec, and the output of the encoder is used as the input to the decoder, thereby obtaining multi-scale sample decoded data output by the decoder. The encoder in the codec includes multiple encoding modules, each of which includes: a multi-head attention mechanism, a residual layer, a normalization layer, and a feedforward network layer; correspondingly, the decoder in the codec includes multiple decoding modules, each of which includes: a multi-head attention mechanism, a residual layer, a normalization layer, a multi-head cross attention mechanism, and a feedforward network.

[0048] 224 , using the prediction head network of the initial detection model to perform target prediction on the multi-scale sample decoded data to obtain an initial prediction result.

[0049] Here, the multi-scale sample decoded data output by the decoder is used as the input of the prediction head network to perform target prediction, thereby obtaining the initial prediction result.

[0050] 225. Based on the initial prediction result and the true value corresponding to the integrated data set, the network parameters of the initial backbone network, the codec, and the prediction head network are adjusted to obtain the target detection model.

[0051] Here, the network parameters such as the weights and learning rates of the initial backbone network, the codec, and the prediction head network are iteratively trained by the difference between the initial prediction result and the true value corresponding to the integrated data set, thereby obtaining a target detection model that meets the convergence conditions.

[0052] 103. Based on the state space model, the backbone network of the target detection model is adjusted to construct a candidate backbone network.

[0053] Here, the network parameters of the initial backbone network of the target detection model are adjusted by introducing a state space model to obtain a candidate backbone network.

[0054] In some possible implementations, the above step 103 may be implemented by the following steps 131 and 132 (not shown):

[0055] 131, determine the network parameters of the initial backbone network of the target detection model.

[0056] Here, the network parameters of the initial backbone network include weight coefficients, learning rates, and the like.

[0057] 132. Adjust the network parameters based on the state space model to obtain the candidate backbone network.

[0058] Here, first, continuous parameters of the state-space model are determined; for example, continuous weight factors of the state-space model. Then, network parameters of the initial backbone network of the object detection model are adjusted based on the continuous parameters to obtain the candidate backbone network. For example, the continuous parameters are discretized and the candidate backbone network is constructed using the state-space model with the discretized parameters.

[0059] In some possible implementations, first, the continuous parameters are discretized based on a preset time scaling parameter to obtain discrete parameters; then, the network parameters of the initial backbone network of the target detection model are adjusted based on the discrete parameters to obtain the candidate backbone network. For example, based on the Kalman filter, the state space model (SSM) can be regarded as a linear time-invariant system, mapping the stimulus u(t) to the response (t) through the hidden layer h(t). The SSM can be formulated as shown in formulas (1) and (2):

[0060] h′(t)=Ah(t)+Bu(t) (1);

[0061] (t) = Ch(t) (2);

[0062] Among them, A, B, and C are continuous weight factors.

[0063] In order to improve the calculation efficiency, the parameters A, B, and C in the above formula need to be converted into discrete parameters. Specifically, assuming the time interval is △, the zero-order hold principle can be used to obtain the corresponding discrete parameters. The corresponding discrete parameters can be expressed as shown in formulas (3), (4) and (5):

[0064]

[0065]

[0066] The discretization is achieved by introducing a time scaling parameter △; exp(·) represents an exponential function, and the superscript -1 represents an inverse operation.

[0067] Formulas (1) and (2) can be expressed using discrete parameters as shown in formulas (6) and (7):

[0068]

[0069] In some possible implementations, first, the output feature map of the state space model is adjusted based on the discrete parameters through discrete parameters to obtain an adjusted feature map; second, the adjusted feature map is compressed and hidden to obtain a target feature map; finally, the network parameters of the initial backbone network of the target detection model are adjusted based on the target feature map to obtain the candidate backbone network.

[0070] For example, based on SSM, it is introduced into the visual task and two-dimensional selective scanning is performed. The image is expanded in four directions to create four separate sequences, which are processed separately by SSM. Finally, the obtained features are combined to generate a complete two-dimensional feature map. Assuming the input feature map x, the output feature map of the two-dimensional selective scanning is As shown in formulas (8), (9) and (10):

[0071] x v =expand(x,v) (8);

[0072]

[0073] where ν∈{1,2,3,4} represents four different scan directions. Furthermore, expand(·) and merge(·) represent scan expansion and scan merging operations, and the S6 function is used as the core operator of the state-space model block, which promotes the interaction between each element in the one-dimensional array and any previously scanned sample by compressing the hidden state.

[0074] In some possible implementations, the network structure of the candidate backbone network is as follows: Figure 3 As shown, the candidate backbone network includes: linear layer 301 (Linear), convolutional layer 302 (Conv1), activation function 303, SSM 304, linear layer 305 (Linear), linear layer 306 (Linear), convolutional layer 307 (Conv1), and activation function 308. After the integrated dataset is input into linear layer 301, it is processed by linear layer 301, convolutional layer 302, activation function 303, SSM 304, linear layer 305, linear layer 306, convolutional layer 307, and activation function 308, and a multi-scale feature map is output from linear layer 305.

[0075] 104 , replacing the initial backbone network in the target detection model with the candidate backbone network to obtain an updated detection model for performing target detection on the remote sensing image to be detected.

[0076] Here, the initial backbone network of the established baseline Encoder-Decoder target detection model is replaced with a candidate backbone network based on feature extraction of the state space model to form a new feature extraction model. This is the updated detection model. The composition structure of the updated detection model is as follows: Figure 4As shown in FIG, the updated detection model includes: Backbone (SSM) based on SSM (candidate backbone network), Encoder (encoder), Decoder (decoder) and prediction results 42; wherein, the encoder includes: PosEmb (position encoding), MHSA (multi-head attention mechanism), RepBlock (multiple residual modules), Fusion (multiple hybrid modules); the decoder includes: multi-scale encoded data 43, Query Selection (query selection mechanism), DecoderLayers & Head (decoding layer and prediction head network), target category and bounding box prediction 44. In the encoder, the composition structure of the hybrid module is as follows Figure 5 As shown, the hybrid module includes: Concat (concatenation layer), Conv1 (convolution layer) and RepBlock (residual module); the composition structure of the residual module is as follows Figure 6 As shown in the figure, the residual module includes: Split (slicing), Conv1 (1×1 convolution layer), 3×3Conv1 and Concat (splicing layer).

[0077] Among them, image 41 is the input image of the candidate backbone network. The candidate backbone network extracts features from image 41 and obtains three layers of output feature information (i.e., M3, M4, and M5). The input of the multi-scale cross encoder comes from the three layers of output feature information of the backbone network, M4 comes from the downsampled feature information of M3, and M5 comes from the downsampled feature information of M4. This strategy helps to capture more global and local features to achieve more stable performance. The downsampling layer performs semantic feature fusion in a shallow to deep order, while the upsampling layer performs semantic feature fusion in a deep to shallow order. The feature fusion of the multi-scale cross encoder provides reliable instance features to the decoder. The multi-scale cross encoder performs convolution and attention operations on the obtained three layers of semantic features of different depths to obtain higher-level semantic features. The position encoding information of the high-level features and multi-head attention are used to allow the model to capture the connection between conceptual entities, which facilitates the subsequent modules to locate and identify objects.

[0078] In the encoder-decoder architecture, object queries are initialized with zero-valued learned embeddings and iteratively optimized by the decoder. However, since there is no direct correlation between the object query values ​​and the image features obtained by the encoder, this setup results in redundant training time and hinders the decoder's ability to find an optimal solution to the query. To alleviate the challenges associated with optimizing object queries in the encoder-decoder, the object query is initialized by selecting the top K features from the encoder using confidence scores. However, a high classification score for a given object does not necessarily ensure accurate detection, as high classification scores can coexist with low intersection-over-union (IoU) values. To address this issue, constraints are imposed on the decoder's queries during model training. Specifically, the model is trained to produce high classification scores for features with high IoU values ​​and low classification scores for features with low IoU values, thereby encouraging the model to prioritize features exhibiting both high classification and high IoU. The query selection decoder comprises multiple components, including multi-head attention, deformable attention, cross attention, and a feed-forward network module. In the decoder, each object is represented by an object query, and the query vector interacts with the image features output by the encoder to generate a category and bounding box prediction for the object. Cross attention promotes the interaction between the object query and the encoder feature map, helping the model learn the spatial location of the target. The entire decoder adopts a multi-level structure, and each layer provides further refinement of the output of the previous layer to improve the detection accuracy. In order to further improve the prediction accuracy and ensure alignment with the ground truth, a fine-grained localization loss is introduced, as shown in formulas (11) and (12):

[0079]

[0080] Among them, Pr l (n) k Represents the probability distribution corresponding to the k-th prediction. φ is the relative offset, calculated as φ=(d GT -d 0 ) / (H,H,W,W);d GT represents the ground truth edge distance, n ← and n → is the bin index adjacent to φ. With weight ω ← and ω → The cross entropy (CE) loss ensures that the interpolated values ​​are precisely aligned with the ground truth offsets. By combining the intersection-over-union (IoU)-based weighting, the fine-grained localization loss encourages distributions with lower uncertainty to become more concentrated, resulting in more accurate and reliable bounding box regression.

[0081] In some embodiments, after training the target detection model with a multi-scale joystick dataset to obtain an updated detection model, the candidate backbone network in the updated detection model is used to extract multi-scale features of the remote sensing image to be detected to obtain multi-scale feature information; and the codec of the updated detection model is used to cross-fuse the multi-scale feature information to obtain multi-scale fusion features; finally, the prediction head network of the updated detection model is used to perform target detection on the multi-scale fusion features to obtain the target detection result of the remote sensing image to be detected.

[0082] like Figure 4 As shown in the figure, the feature extraction backbone network (i.e., the candidate backbone network) extracts features at three different depths: M3, M4, and M5. M4 is derived from the downsampled feature information of M3, and M5 is derived from the downsampled feature information of M4. This strategy helps capture more global and local features, resulting in more stable performance. M3, M4, and M5 are then passed to the encoder for further feature information processing. The encoder performs convolution and attention operations on the three layers of semantic features obtained at different depths to obtain higher-level semantic features. The positional encoding information of the high-level features and multi-head attention enable the model to capture the connections between conceptual entities, facilitating object localization and recognition in subsequent modules. The encoder output is transmitted to the decoder. After receiving the encoder output, the decoder uses a query selection mechanism to find the optimal solution. Specifically, during training, the model is constrained to produce high classification scores for features with high IoU scores and low classification scores for features with high IoU scores, guiding the model to prioritize features with both high classification and high IoU scores. In the decoder, each target is represented by an object query. The query vector interacts with the image features output by the encoder to generate the target category and bounding box prediction, thereby obtaining the detection result, such as Figure 7 As shown, the detection frames 71 and 72 are the detection results of the marked ships for target detection in the remote sensing image to be detected.

[0083] To facilitate the effective deployment of the proposed detection model, embodiments of the present invention have implemented a lightweight design for the existing detection model. State-space models can model long-range feature dependencies, which is crucial for understanding the global context in remote sensing imagery. Models based on state-space models can employ a multi-scan strategy to ensure that every part of the image is connected to every other part. However, this multi-scan strategy significantly increases feature redundancy in the state-space model. By reducing the number of trained selective state-space model layers, the overall computational requirements can be significantly reduced. Embodiments of the present invention minimize the impact of the number of layers in the variable state-space module on prediction results while significantly reducing the number of parameters. Therefore, the layer count in the variable state-space module is reduced, and depthwise separable convolutions are introduced in the multi-scale cross-sensor fusion and reparameterization block modules to further reduce the parameter count. Comparative experimental results show that after lightweighting, the lightweight model has a 35% reduction in parameters compared to the non-lightweighted model, while only experiencing a 0.4% drop in average precision.

[0084] The embodiment of the present invention provides a multi-scale remote sensing image target detection device, see Figure 8 , which shows a schematic diagram of the structure of a multi-scale remote sensing image target detection device provided by one embodiment of the present invention. The device 800 includes:

[0085] An integration module 801 is used to integrate multiple remote sensing image datasets of different types to obtain an integrated dataset;

[0086] A first construction module 802 is configured to construct an object detection model of a codec structure based on the integrated data set;

[0087] A second construction module 803 is configured to adjust the backbone network of the target detection model based on the state space model to construct a candidate backbone network;

[0088] The replacement module 804 is configured to replace the initial backbone network in the target detection model with the candidate backbone network to obtain an updated detection model for performing target detection on the remote sensing image to be detected.

[0089] In some possible implementations, the second construction module 803 is further configured to determine network parameters of an initial backbone network of the target detection model; and adjust the network parameters based on the state space model to obtain the candidate backbone network.

[0090] In some possible implementations, the second construction module 803 is further used to determine continuous parameters of the state space model; and adjust network parameters of the initial backbone network of the target detection model based on the continuous parameters to obtain the candidate backbone network.

[0091] In some possible implementations, the second construction module 803 is further used to discretize the continuous parameters based on preset time scaling parameters to obtain discrete parameters; and adjust the network parameters of the initial backbone network of the target detection model based on the discrete parameters to obtain the candidate backbone network.

[0092] In some possible implementations, the second construction module 803 is further used to adjust the output feature map of the state space model based on the discrete parameters to obtain an adjusted feature map; compress and hide the adjusted feature map to obtain a target feature map; and adjust the network parameters of the initial backbone network of the target detection model based on the target feature map to obtain the candidate backbone network.

[0093] In some possible implementations, the first construction module 802 is further used to construct an initial detection model based on the initial backbone network, codec and prediction head network; and train the initial detection model based on the integrated data set to obtain the target detection model.

[0094] In some possible implementations, the first construction module 802 is also used to use the initial backbone network of the initial detection model to extract features from the integrated data set to obtain feature information; embed the feature information to obtain processed features; use the codec of the initial detection model to encode and decode the processed features to obtain multi-scale sample decoding data; wherein, the encoder in the codec includes multiple encoding modules, each encoding module includes: a multi-head attention mechanism, a residual layer, a normalization layer, and a feedforward network layer; correspondingly, the decoder in the codec includes multiple decoding modules, each decoding module includes: a multi-head attention mechanism, a residual layer, a normalization layer, a multi-head cross attention mechanism, and a feedforward network; use the prediction head network of the initial detection model to perform target prediction on the multi-scale sample decoding data to obtain an initial prediction result; based on the initial prediction result and the true value corresponding to the integrated data set, the network parameters of the initial backbone network, the codec and the prediction head network are adjusted to obtain the target detection model.

[0095] In some possible implementations, the first construction module 802 is further used to perform input embedding processing on the feature information to obtain a continuous vector; perform position embedding processing on the feature information to obtain position information corresponding to the continuous vector; and fuse the continuous vector with the position information to obtain the processed feature.

[0096] In some possible implementations, the replacement module 804 is further used to use the candidate backbone network in the updated detection model to extract multi-scale features of the remote sensing image to be detected to obtain multi-scale feature information; use the codec of the updated detection model to cross-fuse the multi-scale feature information to obtain multi-scale fusion features; use the prediction head network of the updated detection model to perform target detection on the multi-scale fusion features to obtain the target detection result of the remote sensing image to be detected.

[0097] Optionally, the transmission medium can be a wired link (for example, but not limited to, coaxial cable, optical fiber and digital subscriber line (DSL)) or a wireless link (for example, but not limited to, wireless Fidelity (WIFI), Bluetooth and mobile device network). It should be noted that the device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the method embodiments provided in the above embodiments belong to the same concept. The specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0098] Figure 9 FIG. 1 is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention. For example, Figure 9 As shown, the computer device 900 includes: a memory 901, a processor 902, and a computer program 903 stored in the memory 901 and running on the processor 902, wherein when the processor 902 executes the computer program 903, the computer device can execute any of the multi-scale remote sensing image target detection methods described above.

[0099] In addition, an embodiment of the present invention also protects a system, which may include a memory and a processor, wherein the memory stores an executable program code, and the processor is used to call and execute the executable program code to perform a target detection method for a multi-scale remote sensing image provided by an embodiment of the present invention. This embodiment can divide the system into functional modules according to the above method example. For example, it can correspond to each functional module, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical functional division. There may be other division methods in actual implementation. It should be noted that all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.

[0100] It should be understood that the device provided in this embodiment is used to perform the above-mentioned target detection method for multi-scale remote sensing images, and therefore can achieve the same effect as the above-mentioned implementation method. In the case of an integrated unit, the device may include a processing module and a storage module. Among them, when the device is applied to a device, the processing module can be used to control and manage the actions of the device. The storage module can be used to support the device to execute mutual program codes, etc. Among them, the processing module can be a processor or a controller, which can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the contents disclosed in the present invention. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module can be a memory.

[0101] In addition, the apparatus provided in the embodiments of the present invention may be specifically a chip, component, or module. The chip may include a connected processor and memory; the memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute the multi-scale remote sensing image target detection method provided in the above embodiment. This embodiment also provides a computer-readable storage medium storing computer program code. When the computer program code is executed on a computer, the computer executes the above-mentioned method steps to implement the multi-scale remote sensing image target detection method provided in the above embodiment.

[0102] This embodiment also provides a computer program product. When the computer program product is run on a computer, it causes the computer to execute the above-mentioned steps to implement a method for detecting targets in a multi-scale remote sensing image provided by the above embodiment. The device, computer-readable storage medium, computer program product, or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding method provided above, and will not be repeated here. Through the description of the above embodiments, those skilled in the art can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual application, the above-mentioned functions can be distributed to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not performed. On the other hand, the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or unit, which may be electrical, mechanical or other forms.

[0103] It should be noted that the above-mentioned order of the embodiments of the present invention is for description only and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous. The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. The above content is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered within the scope of protection of the present invention.

Claims

1. A method for target detection in multi-scale remote sensing images, characterized in that: The target detection method of the multi-scale remote sensing image comprises: Integrate multiple remote sensing image datasets of different types to obtain an integrated dataset; Based on the integrated data set, construct an object detection model with a codec structure; Adjusting the backbone network of the target detection model based on the state space model to construct a candidate backbone network; The initial backbone network in the target detection model is replaced by the candidate backbone network to obtain an updated detection model for performing target detection on the remote sensing image to be detected.

2. The target detection method for multi-scale remote sensing images according to claim 1, characterized in that: The adjusting the backbone network of the target detection model based on the state space model to construct a candidate backbone network includes: Determining network parameters of an initial backbone network of the target detection model; The network parameters are adjusted based on the state space model to obtain the candidate backbone network.

3. The target detection method for multi-scale remote sensing images according to claim 2, characterized in that: The adjusting the network parameters based on the state space model to obtain the candidate backbone network includes: determining continuous parameters of the state-space model; The network parameters of the initial backbone network of the target detection model are adjusted based on the continuous parameters to obtain the candidate backbone network.

4. The target detection method for multi-scale remote sensing images according to claim 3, characterized in that: The adjusting the network parameters of the initial backbone network of the target detection model based on the continuous parameters to obtain the candidate backbone network includes: discretize the continuous parameter based on a preset time scaling parameter to obtain a discrete parameter; The network parameters of the initial backbone network of the target detection model are adjusted based on the discrete parameters to obtain the candidate backbone network.

5. The target detection method for multi-scale remote sensing images according to claim 4, characterized in that: The adjusting the network parameters of the initial backbone network of the target detection model based on the discrete parameters to obtain the candidate backbone network includes: Adjusting the output characteristic graph of the state-space model based on the discrete parameters to obtain an adjusted characteristic graph; Compressing and hiding the adjusted feature map to obtain a target feature map; Based on the target feature map, the network parameters of the initial backbone network of the target detection model are adjusted to obtain the candidate backbone network.

6. The method for target detection in multi-scale remote sensing images according to claim 1, wherein: The target detection model of the encoding and decoding structure is constructed based on the integrated data set, including: Build an initial detection model based on the initial backbone network, encoder-decoder, and prediction head network; The initial detection model is trained based on the integrated data set to obtain the target detection model.

7. The method for target detection in multi-scale remote sensing images according to claim 6, characterized in that: The training of the initial detection model based on the integrated data set to obtain the target detection model includes: Using the initial backbone network of the initial detection model to perform feature extraction on the integrated data set to obtain feature information; Embedding the feature information to obtain processed features; The processed features are encoded and decoded using the codec of the initial detection model to obtain multi-scale sample decoding data; wherein the encoder in the codec includes a plurality of encoding modules, each encoding module includes: a multi-head attention mechanism, a residual layer, a normalization layer, and a feedforward network layer; correspondingly, the decoder in the codec includes a plurality of decoding modules, each decoding module includes: a multi-head attention mechanism, a residual layer, a normalization layer, a multi-head cross attention mechanism, and a feedforward network; Using the prediction head network of the initial detection model to perform target prediction on the multi-scale sample decoded data to obtain an initial prediction result; Based on the initial prediction results and the true values ​​corresponding to the integrated data set, the network parameters of the initial backbone network, the codec and the prediction head network are adjusted to obtain the target detection model.

8. The method for target detection in multi-scale remote sensing images according to claim 7, characterized in that: The embedding process of the feature information to obtain processed features includes: Performing input embedding processing on the feature information to obtain a continuous vector; Performing position embedding processing on the feature information to obtain position information corresponding to the continuous vector; The continuous vector is fused with the position information to obtain the processed feature.

9. The method for target detection in multi-scale remote sensing images according to claim 1, wherein: The method further comprises: Extracting multi-scale features of the remote sensing image to be detected using the candidate backbone network in the updated detection model to obtain multi-scale feature information; Cross-fusing the multi-scale feature information using the codec of the updated detection model to obtain multi-scale fused features; The prediction head network of the updated detection model is used to perform target detection on the multi-scale fusion features to obtain a target detection result of the remote sensing image to be detected.

10. A target detection device for multi-scale remote sensing images, characterized in that: The device comprises: An integration module is used to integrate multiple remote sensing image datasets of different types to obtain an integrated dataset; A first construction module is used to construct an object detection model of a codec structure based on the integrated data set; A second construction module is used to adjust the backbone network of the target detection model based on the state space model to construct a candidate backbone network; A replacement module is used to replace the initial backbone network in the target detection model with the candidate backbone network to obtain an updated detection model for performing target detection on the remote sensing image to be detected.