Remote sensing image target detection method and device fusing ground object scene relation, terminal equipment and storage medium
By extracting multi-scale features and fusing global scene features, we generate scene-ground object relationship features, which solves the problem of insufficient utilization of scene information in traditional remote sensing target detection methods, improves detection efficiency and accuracy, and enhances feature expression capabilities.
Patent Information
- Application Number
- CN202510516184.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional remote sensing target detection methods ignore scene information, resulting in a high false positive rate of the model. In addition, the feature spaces of scene classification and target detection models are inconsistent, and manual setting of correction values is required, which cannot cover all complex situations.
Through multi-scale feature extraction, global average pooling, 1*1 convolution mapping and Kronecker product processing, scene features are fused into scene ground object features and global scene features to generate scene ground object relationship features and achieve adaptive fusion.
It improves the efficiency and accuracy of target detection, adapts to various complex scenarios, enhances the comprehensive expression ability of features, and reduces dependence on manual correction values.
Smart Images

Figure CN120673019A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image processing, and in particular to a remote sensing image target detection method, device, terminal equipment and storage medium integrating ground object scene relationship. Background Art
[0002] With the advancement of satellite technology, the resolution of remote sensing images has been significantly improved. Now, they can clearly depict various geospatial objects, such as ships, vehicles, and aircraft. These high-resolution images not only capture more intricate details, but also more accurately represent the geometric structure of the objects.
[0003] In remote sensing applications, target detection is a core task, aiming to effectively detect, accurately identify, and precisely locate various objects in remote sensing images. Traditional target detection methods primarily rely on high-resolution optical images and employ candidate bounding box detection networks for object recognition. However, remote sensing observations are often affected by complex scene conditions, such as lighting variations and occlusions between objects. Traditional methods neglect the use of scene information, resulting in high false positive rates. To address this challenge, existing techniques have proposed modifying target detection algorithms based on prior information about scene categories, thereby improving their accuracy and stability.
[0004] Specifically, existing technologies use classification algorithms to obtain prior information about scene categories and combine it with object detection. This involves adjusting the target confidence using a preset correction value based on the proportion of different scene categories within the scene frame. Because the feature spaces of scene classification and object detection models are inconsistent, meaning that the features extracted by different models may not be semantically aligned, fusion is difficult. Therefore, correction values often need to be manually set based on experimental observations, which cannot cover all complex situations. Summary of the Invention
[0005] The embodiments of the present invention provide a remote sensing image target detection method, apparatus, terminal device and storage medium that integrate the relationship between ground objects and scenes. The method can adaptively integrate scene information and ground object information in various complex scenes, thereby improving the efficiency and accuracy of target detection, and further solving the problem in the prior art that the correction value needs to be manually set due to the inconsistency of the feature space between the scene classification and the target detection model.
[0006] An embodiment of the present invention provides a remote sensing image target detection method integrating ground object and scene relationships, comprising: Acquire remote sensing images to be detected; Input the remote sensing image to be detected into the trained target detection model, so that the target detection model can extract multi-scale features from the remote sensing image to obtain multi-scale ground feature features; Perform global average pooling and convolution processing on the deepest feature in the multi-scale feature layer to obtain the global scene feature; Perform 1*1 convolution mapping on each scale of the multi-scale feature to obtain enhanced feature with the same semantic space as the global scene feature. Perform Kronecker product processing on the enhanced feature and global scene features to obtain the scene feature relationship feature; The scene-ground object relationship features are integrated into the multi-scale object features to obtain the scene-ground object mutual information; The mutual information of scene objects is detected to obtain the target detection result.
[0007] Furthermore, multi-scale feature extraction is performed on the remote sensing image to be detected to obtain multi-scale ground feature features, including: The first feature extractor is used to extract features from the remote sensing image to be detected, thereby obtaining image features; The built-in feature pyramid encoder is used to extract multi-scale features from image features to obtain multi-scale ground feature features.
[0008] Furthermore, the scene-ground object relationship features are integrated into the multi-scale object features to obtain the scene-ground object mutual information, including: After normalizing the scene-ground object relationship features, they are multiplied by the multi-scale ground object features to obtain the normalized weighted scene-ground object features; The normalized weighted scene object features are progressively downsampled and residually connected to obtain the scene object mutual information.
[0009] Furthermore, the mutual information between the scene and the ground objects is detected to obtain the target detection results, including: The local feature extraction branch is used to perform convolution processing on the mutual information of scene objects to obtain the local perception scene object feature map; The regression branch detects the position of the target frame on the local perception scene object feature map and generates the position of each detected target frame; Classify each detection target frame through the classification branch to generate the category and confidence of each detection target frame; The detection target frame corresponding to the confidence level greater than the preset threshold is regarded as the credible target frame; The positions, categories and confidence levels of all trusted target boxes are taken as target detection results.
[0010] Furthermore, the target detection model is determined in the following way: Acquire a plurality of first training samples; each first training sample includes: a first remote sensing image sample and the position, category and confidence of a corresponding first actual target frame; A number of first training samples are input into the target detection model to be trained, so that the target detection model is trained with the first remote sensing image sample as input and the position, category and confidence of the first predicted target box as output, and during the training process, the loss function is calculated according to the position, category and confidence of the first predicted target box and the position, category and confidence of the corresponding first actual target box; the network parameters of the target detection model are adjusted according to the loss function until the loss function converges, thereby obtaining a trained target detection model.
[0011] Furthermore, the loss function includes: ; in, represents the loss function value, represents the smoothing coefficient, represents the distribution probability of each category, Indicates the confidence level corresponding to each category, Indicates the total number of categories.
[0012] Furthermore, before inputting the first training samples into the target detection model to be trained, initial parameters of the second feature extractor in the target detection model to be trained are determined by: Acquire a plurality of second training samples; each second training sample includes: a second remote sensing image sample and the position, category, and confidence of a corresponding second actual target frame; wherein the second remote sensing image sample includes a plurality of second actual target frames of different categories; Initializing a multi-classification remote sensing target detection model to obtain a multi-classification remote sensing target detection model to be trained; the multi-classification remote sensing target detection model includes: a third feature extractor, a feature enhancement encoder, and a feature decoder based on a multi-classification head; Inputting a number of second training samples into the multi-classification remote sensing target detection model to be trained, so that the third feature extractor to be trained performs feature extraction on the second remote sensing image sample to obtain sample image features; the feature enhancement encoder to be trained performs feature enhancement on the sample image features to obtain sample image enhancement features; the feature decoder to be trained performs target detection on the sample image enhancement features to obtain the position, category and confidence of the second predicted target box; during the training process, calculating the total classification regression loss function according to the position, category and confidence of the second predicted target box and the position, category and confidence of the corresponding second actual target box; adjusting the network parameters of the multi-classification remote sensing target detection model according to the total classification regression loss function until the total classification regression loss function converges, thereby obtaining a trained multi-classification remote sensing target detection model; The parameters of the third feature extractor in the trained multi-classification remote sensing target detection model are used as the initial parameters of the second feature extractor in the target detection model to be trained.
[0013] Based on the above method embodiment, the present invention provides a corresponding device embodiment, including: a remote sensing image acquisition module, a multi-scale feature extraction module, a global feature extraction module, a feature mapping module, a feature association module, a feature fusion module and a target detection module; A remote sensing image acquisition module is used to acquire remote sensing images to be detected; The multi-scale feature extraction module is used to input the remote sensing image to be detected into the trained target detection model so that the target detection model can perform multi-scale feature extraction on the remote sensing image to be detected and obtain multi-scale ground feature; The global feature extraction module is used to perform global average pooling and convolution processing on the deepest feature of multi-scale feature to obtain global scene features; The feature mapping module is used to perform 1*1 convolution mapping on the feature of each scale in the multi-scale feature to obtain the enhanced feature with the same semantic space as the global scene feature; The feature association module is used to perform Kronecker product processing on the enhanced feature and the global scene feature to obtain the scene feature relationship feature; Feature fusion module, used to fuse scene-ground object relationship features into multi-scale object features to obtain scene-ground object mutual information; The target detection module is used to detect the mutual information of scene objects and obtain target detection results.
[0014] Based on the above-mentioned method embodiment, the present invention provides a corresponding terminal device embodiment, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the steps of the remote sensing image target detection method that integrates the relationship between land objects and scenes as described in the present invention.
[0015] Based on the above-mentioned method embodiment, the present invention provides a corresponding computer-readable storage medium embodiment, including: a stored computer program, which, when the computer program is running, controls the device where the computer-readable storage medium is located to execute the steps of the remote sensing image target detection method that integrates the relationship between land objects and scenes as described in the present invention.
[0016] Compared with the prior art, the beneficial effects of the embodiment of this solution are: The present invention obtains a remote sensing image to be detected and inputs the remote sensing image to be detected into a trained target detection model so that the target detection model performs multi-scale feature extraction on the remote sensing image to be detected, thereby obtaining multi-scale land feature. Since the deepest land feature in land feature features of different scales contains high-level semantic information, the deepest land feature in the multi-scale land feature is subjected to global average pooling and convolution processing, and the high-level semantic information is converted into a global representation to obtain a global scene feature. Then, a 1*1 convolution mapping is performed on the land feature of each scale in the multi-scale land feature. Since the 1*1 convolution mapping does not change the spatial dimension of the feature, but realizes the linear transformation of the feature by changing the number of channels, the land feature features of different scales are aligned with the global scene features in the semantic space, thereby obtaining an enhanced land feature. Subsequently, the enhanced object features and global scene features are subjected to the Kronecker product process. The Kronecker product process can capture all possible element combinations between the two feature tensors, thereby generating scene-object relationship features. The scene-object relationship features contain the relationship between the scene and the objects. The scene-object relationship features are fused with the multi-scale object features to obtain scene-object mutual information. This feature not only contains the local details of the objects, but also incorporates the global relationship between the objects and the scene, thereby enhancing the comprehensive expression ability of the features. Finally, the scene-object mutual information is detected. Because the scene-object mutual information integrates the information of the scene and the objects, the detection decoder can perform target detection based on the scene-object mutual information that integrates the object information and the scene information.
[0017] In summary, the present invention realizes the alignment of object features and scene features in the semantic space through 1*1 convolution mapping, and fuses scene information and object information through Kronecker product, thereby realizing the adaptive fusion of the relationship between scene and object. It can adaptively fuse scene information and object information in various complex scenes, thereby improving the efficiency and accuracy of target detection, and further solving the problem in the prior art that the correction value needs to be manually set due to the inconsistency of the feature space of scene classification and target detection model. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 1 is a flow chart of a remote sensing image target detection method integrating ground object and scene relationship provided by one embodiment of the present invention; Figure 2 1 is a schematic diagram of the structure of a global context fuser provided in a target detection model according to an embodiment of the present invention; Figure 3 1 is a schematic diagram of the structure of a detection decoder provided in a target detection model according to an embodiment of the present invention; Figure 4 1 is a flow chart of a target detection model training process provided by one embodiment of the present invention; Figure 51 is a schematic structural diagram of a feature decoder based on a multi-classification head according to an embodiment of the present invention; Figure 6 1 is another flowchart of the target detection model training process provided by one embodiment of the present invention; Figure 7 It is a structural diagram of a remote sensing image target detection device integrating the relationship between ground objects and scenes provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features.
[0021] like Figure 1 As shown, in order to solve the problem in the prior art that the correction value needs to be manually set due to the inconsistency between the feature spaces of the scene classification and the target detection model, an embodiment of the present invention provides a remote sensing image target detection method that integrates the relationship between the ground object and the scene, and the method includes at least the following steps: Step S1: Acquire the remote sensing image to be detected; For step S1, a high spatial resolution remote sensing image is obtained from the remote sensing platform , as the remote sensing image to be detected, where It means a sheet with 3 channels and width , height is Remote sensing images. Remote sensing images refer to image data of the earth's surface or atmosphere obtained through remote sensing technology.
[0022] Step S2: inputting the remote sensing image to be detected into the trained target detection model, so that the target detection model performs multi-scale feature extraction on the remote sensing image to be detected to obtain multi-scale ground feature; For step S2, the remote sensing image to be detected obtained in step S1 is , and input it into the trained target detection model. The target detection model will perform multi-scale feature extraction on the remote sensing image to obtain a multi-scale feature representation containing rich ground object information.
[0023] It is understandable that multi-scale feature representation can provide richer information, enabling the model to better understand the scene and context in the image, helping the model to identify more details and features in complex remote sensing imagery. At the same time, multi-scale feature extraction enables the model to adapt to remote sensing images of different resolutions and scales, thereby improving the model's generalization and adaptability.
[0024] In a preferred embodiment, multi-scale feature extraction is performed on the remote sensing image to be detected to obtain multi-scale ground feature, including: The first feature extractor is used to extract features from the remote sensing image to be detected, thereby obtaining image features; The built-in feature pyramid encoder is used to extract multi-scale features from image features to obtain multi-scale ground feature features.
[0025] In one embodiment of the present invention, the trained object detection model includes a first feature extractor, a feature pyramid encoder, a global context fuser, and a detection decoder.
[0026] In this trained target detection model, the remote sensing image to be detected First, the built-in first feature extractor performs preliminary feature extraction on the remote sensing image to be detected, outputting image features. These features contain essential information from the remote sensing image, such as edges, texture, and color. Raw remote sensing images typically contain a large amount of data. This preliminary feature extraction effectively reduces the data dimension, reduces computational effort, and improves processing efficiency. This preliminary feature extraction also captures key information from the remote sensing image, thereby improving the accuracy and robustness of target detection.
[0027] Then, the built-in feature pyramid encoder is used to extract multi-scale features from the image features output by the first feature extractor. The feature pyramid encoder can capture the features of objects at different scales. , Represents the number of scales, thereby more comprehensively describing the information in remote sensing images.
[0028] Step S3: Perform global average pooling and convolution processing on the deepest feature in the multi-scale feature to obtain the global scene feature; For step S3, the deepest feature in the multi-scale feature is integrated through the built-in global context fuser. , as potential scene information, is processed by global average pooling, which can compress the spatial dimensions of the feature map into a scalar, thereby aggregating the global information of the entire feature map. Through global average pooling, the model can better understand the overall scene in the image, thereby enhancing the expressiveness of features. In addition, global average pooling can significantly reduce the dimensionality of the feature map and reduce the amount of computation, while retaining important global information, which is beneficial for the subsequent 1*1 convolution processing step.
[0029] After global average pooling, the global aggregated feature vector is obtained. Then, the globally aggregated feature vector is processed by 1*1 convolution. 1*1 convolution can linearly transform and fuse the features without changing the spatial dimension of the feature map, and finally obtain the global scene feature. .
[0030] Step S4: Perform 1*1 convolution mapping on the feature of each scale in the multi-scale feature to obtain the enhanced feature with the same semantic space as the global scene feature; For step S4, Figure 2 As shown, the built-in global context fusion device is used to integrate multi-scale ground feature , linear transformation is performed through 1*1 convolution, changing the representation of features so that they are projected onto the same features as the global scene features The same semantic space to obtain enhanced ground features with the same semantics as the global scene features , .
[0031] Through 1*1 convolution mapping, features of objects at different scales are converted into the same semantic space as the global scene features, so that these features can be compared and fused at the same semantic level to achieve semantic alignment.
[0032] Step S5: Perform Kronecker product processing on the enhanced object features and the global scene features to obtain scene object relationship features; For step S5, the enhanced features after alignment are calculated using the following formula: and global scene features The similarity between them is to perform Kronecker product processing: in, Represents the relationship characteristics of the scene and objects. Indicates enhanced terrain features. represents the Kronecker product operation, represents the global scene features, represents the index of the scale, , Indicates the number of scales.
[0033] It should be noted that the Kronecker product is a matrix operation. In this embodiment, it is used to combine the matrix of enhanced feature and the matrix of global scene features to generate a new feature matrix containing more information, namely, the scene feature relationship feature. In this new matrix, the elements of the original feature matrix are multiplied by specific rules, thereby capturing the complex relationship between scene features and object features, and realizing the interaction between scene features and object features, namely, scene-object relationship features Used to represent the relationship between scene information and ground object information.
[0034] Step S6: Fusing the scene-ground object relationship features into the multi-scale object features to obtain the scene-ground object mutual information; In a preferred embodiment, scene-ground-object relationship features are fused into multi-scale object features to obtain scene-ground-object mutual information, including: After normalizing the scene-ground object relationship features, they are multiplied by the multi-scale ground object features to obtain the normalized weighted scene-ground object features; The normalized weighted scene object features are progressively downsampled and residually connected to obtain the scene object mutual information.
[0035] For step S6, the scene-ground object relationship feature is obtained through step S5. After that, the scene-ground object relationship features are normalized using the following formula and multiplied by the multi-scale ground object features to obtain the normalized weighted scene-ground object features: in, represents the normalized weighted scene feature, Represents the relationship characteristics of the scene and objects. In this embodiment, the sigmoid activation function is used for normalization to enhance the nonlinear expression capability. As a weight, multi-scale features Weighting is performed to achieve deep fusion of scene information and ground object information.
[0036] Then, the normalized weighted scene object features are progressively downsampled and residually connected using the following formula to obtain the scene object mutual information: in, represents the mutual information of scene objects, Represents the normalized weighted scene feature Perform downsampling operation to make its size with the i-1th feature map and perform residual connection. In the feature pyramid, since the feature map sizes of different levels may be different, downsampling operation is required to match the size. This process is actually the aggregation of multi-scale features. Through downsampling and residual connection, the features of different scales are fused together, thereby enhancing the expressive power of the features. The final mutual information of the ground object scene is obtained. , has integrated multi-scale ground feature features, global scene features and the relationships between them, and has high expressive power and rich contextual information.
[0037] Step S7: Detect the scene object mutual information to obtain the target detection result.
[0038] In step S7, the scene object mutual information obtained in step S6 is passed through the detection decoder built into the trained object detection model to perform object detection to obtain object detection results. In this embodiment, the object detection results are the position, category and confidence of each object box.
[0039] In a preferred embodiment, detecting the scene-ground object mutual information to obtain the target detection result includes: The local feature extraction branch is used to perform convolution processing on the mutual information of scene objects to obtain the local perception scene object feature map; The regression branch detects the position of the target frame on the local perception scene object feature map and generates the position of each detected target frame; Classify each detection target frame through the classification branch to generate the category and confidence of each detection target frame; The detection target frame corresponding to the confidence level greater than the preset threshold is regarded as the credible target frame; The positions, categories and confidence levels of all trusted target boxes are taken as target detection results.
[0040] In one embodiment of the present invention, Figure 3 As shown in the figure, the detection decoder built into the trained target detection model includes: local feature extraction branch, regression branch and classification branch.
[0041] The specific process of target detection includes: first, a local feature extraction branch performs a 3x3 convolution on the scene object mutual information, performing layer-by-layer convolution on the scene object mutual information to capture local features in the image and obtain a local perception scene object feature map. Next, the regression branch regresses the extracted local perception scene object feature map, and a regression algorithm (such as linear regression or support vector regression) is used to predict the location of the target object in the image. In this embodiment, this is achieved by generating a bounding box that defines the approximate range of the target object in the image, thereby obtaining the location of each detected target box. Subsequently, the classification branch classifies each detected target box, and a classification algorithm (such as a softmax classifier or support vector machine) is used to classify the image area within the bounding box generated by the regression branch to determine the category of the target object within each bounding box. At the same time, the classification algorithm also outputs a confidence score indicating the model's confidence in the classification result.
[0042] To filter out low-confidence detection results and improve detection accuracy, the detection target boxes corresponding to confidence levels greater than a preset threshold are considered trusted target boxes. The filtered detection results are then integrated, with the positions, categories, and confidence levels of all trusted target boxes being the final target detection results.
[0043] It should be noted that when dealing with the relationship between scenes and objects, traditional methods usually require manual configuration of confidence correction values based on experience. However, manually configured correction values are often based on limited experience and subjective judgment, making it difficult to accurately reflect the complex and changeable relationship between scenes and objects. Once configured, these correction values can usually only be applied to specific models and scenes, and cannot be flexibly adapted to other unconfigured models or new scenes. In contrast, the present invention extracts multi-scale object features from remote sensing images through a feature pyramid encoder, so that it covers multi-scale and multi-level object detail features. The deepest object features in the multi-scale object features are globally averaged and pooled through a global context fuser to aggregate global information. A 1*1 convolution operation is used to align the semantic space. The complex relationship between scene features and object features is captured through a Kronecker product operation, realizing the interaction between scene features and object features. Finally, the interactive information is fused with the multi-scale object features, so that the original multi-scale object features not only contain the detailed information of the object itself, but also make full use of the mutual information between scene objects, thereby enhancing the feature expression ability and improving the accuracy of target detection. Therefore, compared with traditional methods, the present invention can more accurately capture the relationship between scenes and objects, has stronger generalization ability and higher processing efficiency, and is suitable for a wider range of remote sensing image analysis and target detection tasks.
[0044] like Figure 4As shown in Figure 2, the training process of the target detection model includes the following steps: Step S21: Acquire a plurality of first training samples; each first training sample includes: a first remote sensing image sample and the position, category and confidence of a corresponding first actual target frame; In step S21, a number of first training samples are obtained from the database. These samples will be used to train the target detection model. Each first training sample includes a first remote sensing image sample and the position, category and confidence of the corresponding first actual target frame. These first remote sensing image samples have been annotated in advance. In this embodiment, the first remote sensing image sample is a remote sensing image, and the position of the first actual target frame is the actual position of the target object in the remote sensing image, usually in the form of a bounding box. Given in the form of, is the center coordinate of the bounding box (i.e., the target box), and They represent the width and height of the bounding box respectively. The category is the category of the target object, such as a vehicle, building, or water body. The confidence level indicates the degree of certainty about the location and category of the target box, which is a value between 0 and 1.
[0045] The first training sample is used as the training set. Similarly, a validation set and a test set are obtained. The training set is used to train the object detection model, the validation set is used to evaluate the performance of the object detection model and adjust its parameters during the training process, and the test set is used to evaluate the final performance of the object detection model after training is completed.
[0046] It should be noted that the category of the first actual target frame corresponding to the first remote sensing image sample of the present invention only contains a single category (for example, only vehicles or only buildings), so that the target detection model can focus more on learning the features of this specific category, thereby improving the detection accuracy of this category.
[0047] Step S22: Input a number of first training samples into the target detection model to be trained, so that the target detection model is trained with the first remote sensing image sample as input and the position, category and confidence of the first predicted target box as output, and during the training process, calculate the loss function according to the position, category and confidence of the first predicted target box, and the position, category and confidence of the corresponding first actual target box; adjust the network parameters of the target detection model according to the loss function until the loss function converges to obtain a trained target detection model.
[0048] In step S22, several first training samples are input into the object detection model to be trained to optimize the model's parameters so that it can accurately predict the location, category, and confidence of the target box. Specifically, the object detection model performs forward propagation based on the input training samples, outputs the location, category, and confidence of the first predicted target box, and calculates the difference between the predicted result and the true label using a loss function.
[0049] Preferably, the loss function includes: ; in, represents the loss function value, represents the smoothing coefficient, represents the distribution probability of each category, Indicates the confidence level corresponding to each category, Represents the total number of categories. The loss function takes into account the confidence distribution q of each category and the confidence p of the model prediction, and adjusts the smoothness of the confidence through the smoothing coefficient ϵ.
[0050] Next, the gradient of the loss function with respect to the model parameters is calculated using the backpropagation algorithm. An optimization algorithm (such as stochastic gradient descent or Adam) is then used to update the model parameters to reduce the loss function. This process is repeated until the loss function converges or the preset number of training rounds is reached. In each iteration, the model adjusts its parameters based on the current loss function value, gradually improving the accuracy of the prediction.
[0051] After the target detection model is obtained through training, the model is further evaluated and verified through the validation set and test set to ensure its effectiveness and reliability in practical applications.
[0052] In a preferred embodiment, before inputting a plurality of first training samples into the target detection model to be trained, initial parameters of the second feature extractor in the target detection model to be trained are determined by: Acquire a plurality of second training samples; each second training sample includes: a second remote sensing image sample and the position, category, and confidence of a corresponding second actual target frame; wherein the second remote sensing image sample includes a plurality of second actual target frames of different categories; Initializing a multi-classification remote sensing target detection model to obtain a multi-classification remote sensing target detection model to be trained; the multi-classification remote sensing target detection model includes: a third feature extractor, a feature enhancement encoder, and a feature decoder based on a multi-classification head; Inputting a number of second training samples into the multi-classification remote sensing target detection model to be trained, so that the third feature extractor to be trained performs feature extraction on the second remote sensing image sample to obtain sample image features; the feature enhancement encoder to be trained performs feature enhancement on the sample image features to obtain sample image enhancement features; the feature decoder to be trained performs target detection on the sample image enhancement features to obtain the position, category and confidence of the second predicted target box; during the training process, calculating the total classification regression loss function according to the position, category and confidence of the second predicted target box and the position, category and confidence of the corresponding second actual target box; adjusting the network parameters of the multi-classification remote sensing target detection model according to the total classification regression loss function until the total classification regression loss function converges, thereby obtaining a trained multi-classification remote sensing target detection model; The parameters of the third feature extractor in the trained multi-classification remote sensing target detection model are used as the initial parameters of the second feature extractor in the target detection model to be trained.
[0053] In one embodiment of the present invention, before training the target detection model, the feature extractor is pre-trained using multi-classification samples, and the parameters of the feature extractor obtained by pre-training are used as the initial parameters of the second feature extractor in the target detection model to be trained, so that the target detection model has a certain generalization ability to adapt to a wider range of application scenarios.
[0054] First, we obtain multiple remote sensing target detection datasets of different classification systems from the database, such as DIOR, HRRSD, and LEVIR. These datasets represent a remote sensing image containing multiple types of targets. In this embodiment, the batch sizes of the DIOR, HRRSD, and LEVIR datasets are respectively 、 and , we can get The sizes of DIOR, HRRSD and LEVIR datasets are 、 and , then the corresponding batch size of each data set is Can be calculated as: Next, the remote sensing images are preprocessed. Specifically, they are randomly cropped to different sizes and aspect ratios, covering at least 80% of the original image area, while maintaining a maximum of 80%. The cropped images are then scaled to a preset size of 224*224. This preprocessing increases the diversity of the training data, thereby improving the generalization ability of the model. The preprocessed remote sensing images are annotated with information such as the location and category of the objects in the image. The annotated images and their corresponding labels are used as the second training sample for pretraining the feature extractor.
[0055] Subsequently, the second training sample is input into the multi-classification remote sensing target detection model to be trained, and the multi-classification remote sensing target detection model includes a third feature extractor, a feature enhancement encoder, and a feature decoder based on a multi-classification head. In this embodiment, the third feature extractor uses the residual network ResNet50, which contains 50 layers, including 49 convolutional layers and 1 fully connected layer. Its core lies in the residual module. By introducing a skip connection mechanism, the input can bypass several convolutional layers and pass directly to the subsequent layers, effectively alleviating the gradient vanishing and gradient exploding problems in deep networks and promoting network training. The feature enhancement encoder adopts a 12-layer Transformer Block network structure with a multi-head attention mechanism. Figure 5 As shown in FIG, the feature decoder based on multiple classification heads includes three feature enhancement encoding modules, a feature decoding module and three classification heads.
[0056] The second remote sensing image sample in the second training sample is processed by the third feature extractor to obtain the ground feature. Then, the feature enhancement encoder adjusts the input size of the feature enhancement encoder network to 3*224*224. The input image passes through a convolution layer with a convolution kernel size of 16*16, a stride of 16, and an output channel number of 768 to generate a 14*14*768 feature map. Then, the feature map is flattened into a 196×768 feature vector in the spatial dimension using a flattening operation, i.e., the feature enhancement code, as the output of the feature enhancement encoder. Finally, the feature enhancement code is passed through a feature decoder based on a multi-classification head to output the corresponding classification loss and regression loss. The classification loss of each data set is and regression loss Calculated as follows: in, 、 and Represent the weights of the three classification heads respectively, 、 and Represents the classification loss of the three classification heads output, 、 and Represent the regression losses of the three classification head outputs respectively.
[0057] The total classification regression loss is calculated as the weighted sum of classification loss and regression loss, expressed as: The network parameters of the multi-class remote sensing target detection model are adjusted based on the total classification regression loss function until the total classification regression loss function converges. When the total classification regression loss function converges, the model's performance on the training set reaches a relatively stable state. At this point, it can be considered that the model has learned the characteristics and location information of the targets in the remote sensing images, resulting in a trained multi-class remote sensing target detection model.
[0058] Finally, the parameters of the third feature extractor in the trained multi-class remote sensing target detection model are used as the initial parameters for the second feature extractor in the target detection model to be trained. Through this parameter transfer, the target detection model can use the parameters of the already trained feature extractor as the initial parameters, thus avoiding the process of training the feature extractor from scratch and enabling the target detection model to converge to the optimal solution more quickly. In addition, the feature extractor is pre-trained using multi-class samples, which means that during the pre-training phase, the feature extractor will be exposed to samples from multiple categories, thereby learning a broader and more general feature representation.
[0059] like Figure 6 As shown, it is another flow chart of the target detection model training process of the present invention, which uses multiple remote sensing target detection datasets containing different classifications for pre-training. These datasets cover a rich range of land object types and scenes, which helps the model learn a more comprehensive feature representation. Then, according to the parameters of the feature extractor obtained by pre-training, the weights of the parameters of the feature extractor in the target detection model are initialized. Then, a single-class remote sensing target detection dataset is used for comprehensive training to extract features from the input remote sensing image, and the extracted features are encoded using a feature pyramid structure to capture information at different scales. The global context information is integrated into the feature representation to enhance the model's ability to understand the scene, and the fused features are decoded to predict the location, category and confidence of the target.
[0060] like Figure 7 As shown, based on the above method embodiment, a corresponding device embodiment is provided; An embodiment of the present invention provides a remote sensing image target detection device that integrates the relationship between ground objects and scenes, including: a remote sensing image acquisition module, a multi-scale feature extraction module, a global feature extraction module, a feature mapping module, a feature association module, a feature fusion module, and a target detection module; A remote sensing image acquisition module is used to acquire remote sensing images to be detected; The multi-scale feature extraction module is used to input the remote sensing image to be detected into the trained target detection model so that the target detection model can perform multi-scale feature extraction on the remote sensing image to be detected and obtain multi-scale ground feature; The global feature extraction module is used to perform global average pooling and convolution processing on the deepest feature of multi-scale feature to obtain global scene features; The feature mapping module is used to perform 1*1 convolution mapping on the feature of each scale in the multi-scale feature to obtain the enhanced feature with the same semantic space as the global scene feature; The feature association module is used to perform Kronecker product processing on the enhanced feature and the global scene feature to obtain the scene feature relationship feature; Feature fusion module, used to fuse scene-ground object relationship features into multi-scale object features to obtain scene-ground object mutual information; The target detection module is used to detect the mutual information of scene objects and obtain target detection results.
[0061] It can be understood that the above-mentioned device embodiment corresponds to the method embodiment of the present invention, which can implement the remote sensing image target detection method that integrates the relationship between ground objects and scenes provided by any of the above-mentioned method embodiments of the present invention.
[0062] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Furthermore, in the drawings of the device embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which may be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement the present invention without inventive effort.
[0063] Based on the above-mentioned embodiment of the remote sensing image target detection method integrating the relationship between land objects and scenes, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the remote sensing image target detection method integrating the relationship between land objects and scenes of any embodiment of the present invention is implemented.
[0064] For example, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to implement the present invention. The one or more module elements may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.
[0065] The terminal device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0066] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device using various interfaces and lines.
[0067] Based on the above method embodiment, another embodiment is provided: another embodiment of the present invention provides a computer-readable storage medium, including a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the remote sensing image target detection method that integrates the relationship between land objects and scenes as described in any one of the above method embodiments of the present invention.
[0068] The module / unit integrated into the remote sensing image target detection device / terminal device that integrates the relationship between ground objects and scenes, if implemented as a software functional unit and sold or used as a standalone product, can be stored in a computer-readable storage medium. Based on this understanding, the present invention can also implement all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a removable hard drive, a magnetic disk, an optical disk, computer memory, read-only memory (ROM), random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium.
[0069] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A remote sensing image target detection method integrating the relationship between ground objects and scenes, characterized in that: include: Acquire remote sensing images to be detected; Inputting the remote sensing image to be detected into the trained target detection model so that the target detection model performs multi-scale feature extraction on the remote sensing image to be detected to obtain multi-scale ground object features; Perform global average pooling and convolution processing on the deepest feature in the multi-scale feature layer to obtain the global scene feature; Perform 1*1 convolution mapping on each scale of the multi-scale feature to obtain enhanced feature with the same semantic space as the global scene feature. Performing Kronecker product processing on the enhanced feature and the global scene feature to obtain a scene feature relationship feature; Fusing the scene-ground object relationship features into multi-scale ground object features to obtain scene-ground object mutual information; The mutual information of scene objects is detected to obtain the target detection result.
2. The remote sensing image target detection method integrating the relationship between ground objects and scenes according to claim 1 is characterized in that: Performing multi-scale feature extraction on the remote sensing image to be detected to obtain multi-scale ground feature features, including: Extracting features from the remote sensing image to be detected by a built-in first feature extractor to obtain image features; Multi-scale feature extraction is performed on the image features through a built-in feature pyramid encoder to obtain multi-scale ground feature features.
3. The remote sensing image target detection method integrating the relationship between ground objects and scenes according to claim 2 is characterized in that: The scene-ground object relationship features are integrated into the multi-scale ground object features to obtain the scene-ground object mutual information, including: After normalizing the scene-ground object relationship feature, multiplying it with the multi-scale ground object feature to obtain a normalized weighted scene-ground object feature; The normalized weighted scene object features are progressively downsampled and residually connected to obtain scene object mutual information.
4. The remote sensing image target detection method integrating the relationship between ground objects and scenes according to claim 3 is characterized in that: The detecting of the scene object mutual information to obtain the target detection result includes: The local feature extraction branch is used to perform convolution processing on the mutual information of scene objects to obtain the local perception scene object feature map; Performing target frame position detection on the local perception scene object feature map through a regression branch to generate the position of each detected target frame; Classify each detection target frame through the classification branch to generate the category and confidence of each detection target frame; The detection target frame corresponding to the confidence level greater than the preset threshold is regarded as the credible target frame; The positions, categories and confidence levels of all trusted target boxes are taken as target detection results.
5. The remote sensing image target detection method integrating ground object and scene relationship according to claim 4 is characterized in that: The target detection model is determined by: Obtaining a number of first training samples; Each first training sample includes: a first remote sensing image sample and the position, category and confidence of a corresponding first actual target frame; A number of first training samples are input into the target detection model to be trained, so that the target detection model is trained with the first remote sensing image sample as input and the position, category and confidence of the first predicted target box as output, and during the training process, the loss function is calculated according to the position, category and confidence of the first predicted target box and the position, category and confidence of the corresponding first actual target box; the network parameters of the target detection model are adjusted according to the loss function until the loss function converges, thereby obtaining a trained target detection model.
6. The remote sensing image target detection method integrating the relationship between ground objects and scenes according to claim 5 is characterized in that: The loss function includes: in, Represents the loss function value, ∈ represents the smoothing coefficient, q represents the distribution probability of each category, p represents the confidence level corresponding to each category, and K represents the total number of categories.
7. The remote sensing image target detection method integrating ground object and scene relationship according to claim 6, characterized in that: Before inputting a plurality of first training samples into the target detection model to be trained, initial parameters of a second feature extractor in the target detection model to be trained are determined by: Obtaining a number of second training samples; Each second training sample includes: a second remote sensing image sample and the position, category and confidence of a corresponding second actual target frame; wherein the second remote sensing image sample includes a plurality of second actual target frames of different categories; Initializing a multi-classification remote sensing target detection model to obtain a multi-classification remote sensing target detection model to be trained; the multi-classification remote sensing target detection model includes: a third feature extractor, a feature enhancement encoder, and a feature decoder based on a multi-classification head; Inputting a number of second training samples into the multi-classification remote sensing target detection model to be trained, so that the third feature extractor to be trained performs feature extraction on the second remote sensing image sample to obtain sample image features; the feature enhancement encoder to be trained performs feature enhancement on the sample image features to obtain sample image enhancement features; the feature decoder to be trained performs target detection on the sample image enhancement features to obtain the position, category and confidence of the second predicted target box; during the training process, calculating the total classification regression loss function according to the position, category and confidence of the second predicted target box and the position, category and confidence of the corresponding second actual target box; adjusting the network parameters of the multi-classification remote sensing target detection model according to the total classification regression loss function until the total classification regression loss function converges, thereby obtaining a trained multi-classification remote sensing target detection model; The parameters of the third feature extractor in the trained multi-classification remote sensing target detection model are used as the initial parameters of the second feature extractor in the target detection model to be trained.
8. A remote sensing image target detection device integrating the relationship between ground objects and scenes, characterized in that: include: Remote sensing image acquisition module, multi-scale feature extraction module, global feature extraction module, feature mapping module, feature association module, feature fusion module and target detection module; The remote sensing image acquisition module is used to acquire the remote sensing image to be detected; The multi-scale feature extraction module is used to input the remote sensing image to be detected into the trained target detection model, so that the target detection model performs multi-scale feature extraction on the remote sensing image to be detected to obtain multi-scale ground object features; The global feature extraction module is used to perform global average pooling and convolution processing on the deepest feature in the multi-scale feature layer to obtain the global scene feature; The feature mapping module is used to perform 1*1 convolution mapping on the feature of each scale in the multi-scale feature to obtain the enhanced feature with the same semantic space as the global scene feature; The feature association module is used to perform Kronecker product processing on the enhanced feature and the global scene feature to obtain a scene feature relationship feature; The feature fusion module is used to fuse the scene-ground object relationship features into the multi-scale ground object features to obtain the scene-ground object mutual information; The target detection module is used to detect the mutual information of scene objects and obtain target detection results.
9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for remote sensing image target detection integrating the relationship between ground objects and scenes as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the remote sensing image target detection method integrating the relationship between ground objects and scenes as described in any one of claims 1 to 7.
Citation Information
Cited By
Remote sensing image out-of-distribution detection method and device, electronic equipment and storage medium
CN122176542A
Remote sensing image distribution outside detection method and device, electronic equipment and storage medium
CN122176542B