End-to-end target object frame extraction method and device, equipment and medium

By using an end-to-end target object bounding box extraction method and adjusting the model with a total loss function, the problems of vertex redundancy and boundary artifacts in remote sensing images are solved, achieving high-precision building boundary extraction, especially accurate characterization of small-scale buildings.

CN121746902APending Publication Date: 2026-03-27STATE GRID INFORMATION & TELECOMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing building extraction methods suffer from vertex redundancy, boundary artifacts, and insufficient characterization of small-scale buildings in remote sensing images, and are particularly unstable in complex urban scenes.

Method used

An end-to-end target object bounding box extraction method is adopted. By acquiring sample data of real bounding box markings, a target object bounding box extraction model is constructed. The model is then adjusted using the total loss function for training and testing. Finally, bounding box prediction is performed on the remote sensing image to be extracted, and accurate target object bounding boxes are output.

Benefits of technology

It improves the accuracy of building extraction and the geometric consistency of boundaries, reduces vertex redundancy, enhances the ability to depict small-scale buildings, and improves the detection accuracy of target object boundaries in remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746902A_ABST
    Figure CN121746902A_ABST
Patent Text Reader

Abstract

The invention provides an end-to-end target object frame extraction method and device, equipment and a medium. The method comprises the following steps: marking an original remote sensing image by using a real frame to obtain sample data, and training a constructed target object frame extraction model by using the sample data to obtain a trained target object frame extraction model; then, testing the trained target object frame extraction model to obtain a final target object frame extraction model; and finally, processing a remote sensing image to be extracted by using the final target object extraction model, and outputting a corresponding target object frame prediction result, so that an accurate target object contour in the remote sensing image to be extracted can be obtained, and the obtained target object frame prediction result is subjected to geographic data marking or analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, device and medium for end-to-end target object bounding box extraction. Background Technology

[0002] Automatic building detection and high-precision vector polygon generation are important issues that have long been of concern in the field of remote sensing and geographic information science. The core objective is to extract building boundaries from aerial or satellite remote sensing images and represent them in the form of vectorized polygons.

[0003] In recent years, methods have focused on directly learning polygon vertex sequences to avoid complex post-processing. However, these methods still suffer from problems in practical applications, such as vertex redundancy, boundary artifacts, and insufficient depiction of small-scale buildings. Summary of the Invention

[0004] In view of this, the purpose of this application is to propose an end-to-end target object bounding box extraction method, including: Acquire the original remote sensing image, mark the target objects in the original remote sensing image with real bounding boxes, and obtain sample data; A target object bounding box extraction model is constructed, and the sample data is trained using the target object bounding box extraction model. Based on the difference between the training output result and the real bounding box, a total loss function is determined. The target object bounding box extraction model is adjusted based on the loss value determined by the total loss function to obtain the trained target object bounding box extraction model. The trained target object bounding box extraction model is tested, and after the trained target object bounding box extraction model passes the test, it is used as the final target object bounding box extraction model. The remote sensing image to be extracted is acquired, and the remote sensing image to be extracted is input into the final target object bounding box extraction model for bounding box prediction, and the target object bounding box prediction result is output.

[0005] To achieve the above objectives, this application provides an end-to-end target object bounding box extraction device, comprising: The sample acquisition unit is configured to acquire the original remote sensing image, mark the target objects in the original remote sensing image with real bounding boxes, and obtain sample data. The model training unit is configured to construct a target object bounding box extraction model, use the target object bounding box extraction model to train the sample data, determine the total loss function based on the difference between the training output and the real bounding box, and adjust the target object bounding box extraction model based on the loss value determined by the total loss function to obtain the trained target object bounding box extraction model. The testing unit is configured to test the trained target object bounding box extraction model, and after the trained target object bounding box extraction model passes the test, the trained target object bounding box extraction model is used as the final target object bounding box extraction model. The application unit is configured to acquire the remote sensing image to be extracted, input the remote sensing image to be extracted into the final target object bounding box extraction model for bounding box prediction, and output the target object bounding box prediction result.

[0006] Based on the same inventive concept, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.

[0007] Based on the same inventive concept, this disclosure also provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to perform the method described above.

[0008] As can be seen from the above, the end-to-end target object bounding box extraction method, apparatus, device, and medium provided in this application can obtain sample data by marking the original remote sensing image with real bounding boxes. This sample data is then used to train the constructed target object bounding box extraction model. The model is adjusted based on the loss value of the total loss function determined by the difference between the training results and the real bounding boxes, thus completing the training of the target object bounding box extraction model and obtaining the trained model. To improve the accuracy of the trained model, it is tested. Only if the trained model meets the accuracy requirements of the test will the test pass, resulting in the final target object bounding box extraction model. Finally, the final model can be used to process the remote sensing image to be extracted, outputting the corresponding target object bounding box prediction results. This allows for the accurate outline of the target object in the remote sensing image to be extracted, and the obtained target object bounding box prediction results can be used for geographic data labeling or analysis. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1This is a flowchart illustrating the end-to-end target object bounding box extraction method according to an embodiment of this application. Figure 2 This is a schematic diagram illustrating the training process of the target object bounding box extraction model according to an embodiment of this application; Figure 3 This is a schematic diagram illustrating the joint training process of the target object bounding box extraction model in an embodiment of this application; Figure 4 This is a schematic diagram of the end-to-end target object border extraction device according to an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0012] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0013] Definitions: CNN: Convolutional Neural Network.

[0014] U-Net: An image segmentation model based on convolutional neural networks (CNN), consisting of an encoder (extracting features) and a decoder (reconstructing the image). It achieves pixel-level classification by fusing multi-scale features through skip connections.

[0015] DeepLab: A series of semantic segmentation deep learning models developed by the Google Research team. It uses atrous convolution and spatial pyramid pooling (ASPP) as core technologies and achieves high-precision pixel-level classification through an Encoder-Decoder architecture.

[0016] Mask R-CNN: Mask Region-based Convolutional Neural Network, is an improved version of Faster R-CNN that achieves object detection and instance segmentation by adding a mask prediction branch.

[0017] Polygon-RNN: An image segmentation model that combines convolutional neural networks (CNN) and recurrent neural networks (RNN), primarily used for automatically annotating the contours of objects in images.

[0018] PolyMapper: A novel deep learning model for automatically extracting building footprints and road network topology maps from aerial images.

[0019] Curve-GCN is a fast, interactive object annotation framework based on PyTorch.

[0020] PolyWorld is a neural network model that extracts polygonal objects from images in an end-to-end manner.

[0021] HiSup: Introducing hierarchical supervision signals (vertices, boundaries, and masks), HiSup uses a convolutional neural network to achieve high-precision polygon mapping of buildings.

[0022] RoI: Region of Interest, a feature model primarily used in object detection tasks. Its core principle is to map the region of interest in the original image to a feature map and generate a fixed-size feature representation through pooling operations. RoIAlign is a key technique used in object detection algorithms (such as Mask R-CNN) to improve region feature extraction.

[0023] Logit: Logit model, discrete choice model.

[0024] MobileViT: A lightweight, general-purpose vision Transformer model designed to combine the spatial inductive bias of convolutional neural networks (CNNs) with the global dependency modeling capabilities of Transformers.

[0025] ResNet Block: Residual Module.

[0026] MLP: Multilayer Perceptron, a type of feedforward artificial neural network model.

[0027] ReLU: Rectified Linear Unit, is a commonly used activation function in artificial neural networks.

[0028] Det2Poly: End-to-End Vectorized Building Outline Extraction Guided by Object Detection, the target object bounding box extraction model obtained in this application.

[0029] LN: Layer Normalization, a deep learning normalization technique that makes network training more stable by standardizing all features of a single sample.

[0030] In related technologies, building extraction methods have evolved from traditional methods based on manual features to deep learning-driven methods. Early research mainly relied on features such as spectrum, texture, shadow, or geometric rules, achieving building detection and contour extraction through rule setting and heuristic processing. However, in complex urban scenes, the large differences in building appearances and complex environmental backgrounds make such methods unstable. With the widespread application of deep convolutional neural networks and high-resolution remote sensing imagery, the research focus has gradually shifted to learning-driven automated methods. In this framework, a typical multi-step paradigm first obtains a building mask through semantic segmentation or instance segmentation, then uses heuristic algorithms to convert the mask into vector polygons, and finally performs vertex simplification or boundary regularization. However, multi-step methods are prone to error accumulation and have significant shortcomings in boundary geometric accuracy. To overcome these limitations, a series of end-to-end methods have emerged in recent years, attempting to directly learn polygon vertex sequences to avoid complex post-processing. However, these methods still suffer from vertex redundancy, boundary artifacts, and insufficient characterization of small-scale buildings in practical applications.

[0031] With the widespread application of Convolutional Neural Networks (CNNs), building extraction technology has entered a new stage driven by deep learning. Representative works include U-Net and its derivatives, the DeepLab series, and Mask R-CNN, which can learn multi-level representations of images on large-scale datasets, thus significantly improving the robustness and completeness of building detection. In particular, the introduction of instance segmentation frameworks enables building extraction to distinguish individual buildings in dense urban scenes. However, these methods typically output rasterized masks, which, while performing well at the semantic level, often suffer from problems such as blurred boundaries, irregular lines, and unclear corners at the geometric level.

[0032] To overcome the limitations of masks in geometric representation, research has begun to explore end-to-end methods for directly generating vector polygons. On one hand, models based on recurrent neural networks (such as Polygon-RNN and PolyMapper) predict polygon vertex sequences point-by-point through autoregressive mechanisms, enabling the modeling of complex shapes. On the other hand, models based on graph convolutional networks (such as Curve-GCN and PolyWorld) reconstruct building topology by explicitly learning vertex relationships. Furthermore, with the development of the Transformer architecture, some methods transform polygon prediction into an ensemble prediction problem, generating vertex sets in a single inference. These methods significantly improve boundary accuracy and geometric consistency, but still face challenges such as redundant vertices, boundary artifacts, and unstable localization in small-to-medium scale buildings and complex terrain scenes.

[0033] The embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0034] The end-to-end target object bounding box extraction method proposed in this application includes: buildings, industrial parts, medical detection targets (e.g., tumors and / or various organs).

[0035] like Figure 1 As shown, the method includes: Step 101: Obtain the original remote sensing image, mark the target objects in the original remote sensing image with real bounding boxes, and obtain sample data.

[0036] In practice, the high-resolution original remote sensing image is marked with a real bounding box. The marking content includes the target object's rotated quadrilateral bounding box and the polygon vertex coordinate sequence, so as to simultaneously meet the supervision requirements of target detection and vector boundary prediction.

[0037] Data augmentation was performed on the labeled original remote sensing images using operations such as random rotation (±30°), scaling (0.5–2.0), color jitter, and random cropping to simulate the changes of the target object under different scales, viewpoints, and lighting conditions, thereby improving the model's generalization ability.

[0038] The augmented dataset is divided into a training set, a validation set, and a test set in an 8:1:1 ratio. The training set contains sample data used for training. The validation set is used to validate the trained object bounding box extraction model, and the test set is used to test the trained model.

[0039] Step 102: Construct a target object bounding box extraction model, train the sample data using the target object bounding box extraction model, determine the total loss function based on the difference between the training output and the real bounding box, and adjust the target object bounding box extraction model according to the loss value determined by the total loss function to obtain the trained target object bounding box extraction model.

[0040] In practice, the target object bounding box extraction model includes a feature extraction backbone network, a detection query module, a ROI feature bridging module, and a polygon vertex query module.

[0041] Step 103: Test the trained target object bounding box extraction model. After the trained target object bounding box extraction model passes the test, use the trained target object bounding box extraction model as the final target object bounding box extraction model.

[0042] In practice, the trained target object bounding box extraction model is validated using the validation set obtained above. After the validation is confirmed to be successful, the trained target object bounding box extraction model is tested using the test set. If the test fails, the trained target object bounding box extraction model will be retrained. If the test passes, the trained target object bounding box extraction model will be used as the final target object bounding box extraction model.

[0043] Step 104: Obtain the remote sensing image to be extracted, input the remote sensing image to be extracted into the final target object bounding box extraction model for bounding box prediction, and output the target object bounding box prediction result.

[0044] The above scheme allows for the generation of sample data from raw remote sensing images by marking them with real bounding boxes. This sample data is then used to train a target object bounding box extraction model. The model is adjusted based on the loss value of the total loss function, determined by the difference between the training results and the real bounding boxes, thus completing the training process and obtaining a trained target object bounding box extraction model. To improve the accuracy of this model, it is tested. Only when the trained model meets the accuracy requirements of the test is the test passed, resulting in the final target object bounding box extraction model. Finally, this final model can be used to process the remote sensing image to be extracted, outputting the corresponding predicted target object bounding boxes. This allows for the accurate identification of the target object outlines in the remote sensing image, and the predicted bounding boxes can then be used for geographic data labeling or analysis.

[0045] In some embodiments, such as Figure 2 As shown, step 102 includes: Step 1021: Construct a target object bounding box extraction model that includes a feature extraction backbone network, a detection query module, an ROI feature bridging module, and a polygon vertex query module.

[0046] In some embodiments, the feature extraction backbone network is used to perform multi-scale feature extraction processing on the original remote sensing images in the sample data to obtain a multi-scale feature map. The detection query module is used to determine candidate boxes based on the multi-scale feature map; The ROI feature bridging module is used to perform ROI feature bridging processing on the candidate box to obtain bridging features. The polygon vertex query module is used to determine multiple vertices based on the bridging features and to determine the bounding box based on the multiple vertices.

[0047] The above approach enables the construction of an accurate initial target object extraction model, ensuring its effectiveness in extracting the bounding boxes of target objects in remote sensing images.

[0048] Step 1022: Input the sample data into the feature extraction backbone network and the detection query module for the first stage of training. Train and adjust the detection query module to obtain the feature extraction backbone network and the detection query module after the first stage of training.

[0049] In some embodiments, the original remote sensing image is divided according to height into: low layers with height less than or equal to a first height value, mid-low layers with height greater than the first height but less than or equal to the second height, mid-high layers with height greater than the second height but less than or equal to the third height, and high layers with height greater than the third height.

[0050] The feature extraction backbone network (e.g., MobileViT) consists of five layers; First layer: For the original remote sensing image in the sample data, the texture features of the lower layer are extracted using the first convolution kernel combined with the first residual convolution module.

[0051] In practice, the original remote sensing image is an image of a predetermined size (e.g., 512×512). A downsampling operation is performed on the original remote sensing image corresponding to the first convolutional kernel (e.g., a 3×3 convolutional kernel with a stride of 2). Then, combined with the first residual convolutional module (e.g., a ResNet Block), preliminary feature encoding is performed to extract low-level texture features (e.g., local information such as edges, colors, and brightness) from the original remote sensing image, outputting a feature map of low-level texture features of a predetermined size (e.g., 256×256×64).

[0052] The second layer: For the original remote sensing images in the sample data, the first semantic features of the shape contour and basic structure information of local targets in the middle and low layers are extracted by using the second convolution kernel combined with the second residual convolution module.

[0053] In practice, the second convolutional kernel (e.g., with a stride of 2) performs a downsampling operation and, combined with the second residual convolutional module (e.g., a ResNet Block), extracts the first semantic features at the mid-to-low level of the original remote sensing image. These low-level semantic features can capture the shape contours and basic structural information of local targets while reducing spatial resolution, and output a feature map of the first semantic features (e.g., with a size of 128×128×128).

[0054] The third layer: For the original remote sensing images in the sample data, the structural features of the target object are extracted by using the third convolutional kernel in combination with the ResNet module and the attention mechanism.

[0055] In practice, the third convolutional kernel (e.g., with a stride of 2) performs downsampling, combining with the ResNet module and attention mechanism to enhance the response to salient target regions such as the target object. The third layer focuses on extracting the structural features of the target object from the original remote sensing image, namely the boundary shape, roof outline, texture consistency, and directional features of the target object. The attention mechanism is used to weight the feature channels and spatial dimensions, highlighting the features of the target object region and suppressing background interference, outputting a feature map of the target object's structural features (e.g., with a size of 64×64×256).

[0056] Fourth layer: For the original remote sensing images in the sample data, the fourth convolution kernel combined with the fourth residual convolution module is used to extract the second semantic features of scene relationships and object combinations in the middle and high layers.

[0057] In practice, the fourth convolutional kernel (e.g., with a stride of 2) performs downsampling and is combined with the fourth residual convolutional module (e.g., ResNet Block) to extract mid-to-high-level semantic features. The second semantic features obtained from the fourth layer have stronger semantic expressive power, can identify more complex scene relationships and object combination features, provide contextual information support for subsequent detection and segmentation tasks, and output a feature map of the second semantic features (e.g., with a size of 32×32×512).

[0058] Fifth layer: For the original remote sensing images in the sample data, the fifth convolutional kernel is used in combination with the ResNet module and the feature pyramid network to extract the third semantic features that distinguish the overall semantics of the scene from the scene category.

[0059] In practice, the fifth convolutional kernel (e.g., with a stride of 2) performs downsampling operations, and in conjunction with the ResNet module and the Feature Pyramid Network (FPN), high-level third semantic features are extracted. The third semantic features mainly characterize the overall semantics and category discrimination ability of the scene, such as distinguishing target objects from different types of land features such as roads, water bodies, and vegetation, and output the feature map of the third semantic features (e.g., with a size of 16×16×1024).

[0060] The feature extraction backbone network combines the first semantic feature, the target object structural feature, the second semantic feature, and the third semantic feature to form a multi-scale feature map output.

[0061] In practice, the corresponding multi-scale feature maps are {C2, C3, C4, C5}. Here, C2, C3, C4, and C5 represent the feature maps output from the second, third, fourth, and fifth layers, respectively.

[0062] The above method can obtain a multi-scale feature map that accurately represents various features of the target object, which facilitates subsequent analysis based on the multi-scale feature map to determine the corresponding target object bounding box.

[0063] In some embodiments, step 1022 includes: Step 10221: Input the sample data into the feature extraction backbone network, and perform multi-scale feature extraction processing on the original remote sensing image in the sample data to obtain a multi-scale feature map.

[0064] In practice, the process of determining the multi-scale feature map {C2, C3, C4, C5} is as described in the above embodiment, and will not be repeated here.

[0065] Step 10222: Determine multiple initial candidate boxes in the detection query module, and map the coordinates of each initial candidate box to the corresponding level of the multi-scale feature map according to the step size of the multi-scale feature map to obtain the mapped candidate box.

[0066] In practice, the detection and query module uses a fixed number (N) of initial candidate boxes, and the coordinates of each initial candidate box are... and the corresponding feature representation vector ,in, , w and h represent the coordinates of the center point and the width and height of the bounding box after normalization, respectively, and R represents a real number.

[0067] Initialize candidate boxes N×4 (normalizing the center point, width, and height to [0,1]). Next, based on the coordinates of each initialized candidate box, map them onto the corresponding level feature map on the multi-scale feature map {C2,C3,C4,C5} according to the feature map stride to obtain the mapped candidate box.

[0068] (1); in, The levels corresponding to the multi-scale feature maps. and These correspond to the dimensions of the baseline level and the baseline level, respectively, with B representing the candidate box size.

[0069] Step 10223, execute using the detection and query module: Step 102231: Determine the offset of the mapped candidate box, update the mapped candidate box according to the offset to obtain the updated candidate box, and normalize the updated candidate box to obtain the normalized candidate box.

[0070] In practice, RoIAlign processing is performed on each mapped candidate box to obtain S×S ROI features. Then, through a dynamic convolution mechanism, adaptive convolution weights are generated for each mapped candidate box to adjust the ROI features. Instance-level augmentation and iterative bounding box optimization are performed. The augmented features are input into the detection heads (classification and regression heads) for class determination and bounding box coordinate offset prediction to obtain the offset. .

[0071] The classification head consists of two layers of MLP and ReLU, outputting class probabilities: (2).

[0072] The regression head consists of two MLP layers, which output the offset of the candidate box after the current mapping. Applying the offset to the mapped candidate boxes yields the updated candidate boxes. (x',y',w',h'): (3).

[0073] Step 102232: Determine the first loss function by combining the normalized candidate boxes with the true bounding boxes labeled in the sample data. Train and adjust the detection query module based on the first loss value determined by the first loss function to obtain the feature extraction backbone network and the detection query module after the first stage of training.

[0074] In practice, the loss function of the detection branch consists of the classification loss (focal loss), the bounding box regression loss (L1 loss), and the GIoU loss. The first loss function... It is expressed as follows: (4); in, These are the center point and width / height of the updated candidate boxes after normalization. The center point and width / height of the normalized true border are used as monitoring signals. Focal loss for classification weights, for The weight, for The weight, The loss determined for GIoU Loss.

[0075] After calculating the first loss value using the first loss function, the adjustment amount corresponding to each layer of the detection and query module can be determined based on the first loss value, and the detection and query module can be trained and adjusted according to the adjustment amount. After the first stage of training, the feature extraction backbone network and the detection and query module after the first stage of training can be obtained.

[0076] The above approach involves first training the feature extraction backbone network and the detection query module, which ensures that the feature extraction backbone network and the detection query module can achieve better prediction results.

[0077] Step 1023: Input the sample data into the feature extraction backbone network and detection query module, the ROI feature bridging module and the polygon vertex query module after the first stage of training, and perform the second stage of training processing to train and adjust the polygon vertex query module to obtain the polygon vertex query module after the second stage of training.

[0078] In some embodiments, step 1023 includes: Step 10231: Input the sample data into the feature extraction backbone network trained in the first stage, and perform multi-scale feature extraction processing on the original remote sensing images in the sample data to obtain multi-scale feature maps (such as...). Figure 3 (As shown).

[0079] Step 10232: Using the detection query module trained in the first stage, determine candidate boxes (e.g., based on the multi-scale feature map). Figure 3 As shown, candidate boxes are predicted bounding boxes.

[0080] In practice, during the second stage of training, the detection and query module trained in the first stage will be fixed. After the sample data passes through the feature extraction backbone network and the detection and query module trained in the first stage, the candidate boxes of the target objects can be obtained.

[0081] Step 10233: Use the ROI feature bridging module to process the candidate box for ROI features, obtain ROI features, flatten the ROI features and input them into the encoder for encoding to obtain bridging features.

[0082] In practice, the ROI feature bridging module first selects an appropriate level based on the candidate box b obtained by the detection query module on the multi-scale features {C2,C3,C4,C5} according to equation (1). Then, RoIAlign is performed on each candidate box to obtain S×S ROI features. Then, Flattened , and Let represent the spatial dimensions of the ROI features, and C represent the number of channels. Then, these are fed into a two-layer Transformer encoder consisting of a multi-head self-attention network and a feedforward network to obtain the bridging features for each detection instance i. .

[0083] Step 10234: Execute using the polygon vertex query module: Step 102341: Determine the coordinate reference point and the embedding vector, and determine the position embedding of each vertex for the coordinate reference point.

[0084] In practical implementation, the vector corresponding to the candidate box (such as...) Figure 3 The learnable query vector shown includes the coordinate reference point of the initial geometric position and the embedding vector. and embedding vector Where R is a real number and d = 256. Coordinate reference point. As the initial geometric position prior of the vertices within the RoI, a uniform distribution is used for initialization, and the position is dynamically updated layer by layer during decoding through a coordinate offset module to achieve adaptive adjustment of the vertex position. Embedded vector and learnable logit scalar Random initialization is performed at the start of training, and automatic updates are made through backpropagation during model training.

[0085] Position embedding , where sigmoid is the sigmoid normalization function, which ensures that the vertex coordinates are normalized to the [0,1] interval, PE is the position code, which converts the coordinate reference point into a high-dimensional vector, and MLP and LN represent multilayer perceptron layer normalization and LN layer normalization, respectively.

[0086] Step 102342: Determine the scalar representing whether a vertex is valid, process the scalar using a multilayer perceptron, and then convert it into adaptively normalized modulation parameters.

[0087] In practice, (5); in, and Modulation parameters used to achieve vertex-level adaptive normalization. It is a multilayer perceptron, a scalar (logit scalar). , Used to store the learnable logit scalar for each vertex The modulation parameters are converted into adaptive normalized (AdaLN) modulation parameters.

[0088] Step 102343: Determine the initial vertex query vector by combining the modulation parameters, the position embedding, and the embedding amount.

[0089] In practical implementation, the embedding vector will be Location embedding and logit scalar As the initial vertex query vector for vertex queries: (6); Among them, classification confidence As a confidence modulation weight, it retains a strong response to the vertex features of high-confidence targets, while attenuating the vertex activation strength of low-confidence targets (regions where the detection query module is uncertain).

[0090] Step 102344: Perform cross-attention processing based on the bridging features and the vertex query vector obtained in the previous step to obtain the updated vertex query vector for the current iteration.

[0091] Step 102345: Input the updated vertex query vector into the decoder to obtain the coordinate offset and validity probability, and obtain the updated target object vertex coordinates by normalizing the updated vertex query vector.

[0092] In specific implementation, the first l The next update uses the k-th reference point in the i-th detection instance. For sampling locations, bridging features in each detection instance i Perform cross attention on the vertex to obtain updated vertex query features. Then, the updated features The input is fed into a decoder consisting of two fully connected layers and ReLU activation, and the output is the coordinate offset and vertex validity probability. (7); where MLP() is to use the MLP layer for processing, and ReLU() is to use the ReLU layer for processing.

[0093] Updated vertex coordinates of the target object using Sigmoid constraints: (8), among which, For the last time l -1 is the updated vertex coordinates of the target object.

[0094] Step 102346: Determine a second loss function based on the difference between the updated target object vertex coordinates and the real vertices corresponding to the real border markers; adjust the parameters of the polygon vertex query module based on the second loss value determined by the second loss function to obtain the polygon vertex query module after the second stage of training.

[0095] In specific implementation, the second loss function of the polygon vertex query module Regression loss based on vertex coordinates Vertex validity classification loss Composition, represented as follows: (9); Where N and M represent the number of polygons and the number of vertices corresponding to each polygon, respectively. and These represent the updated vertex coordinates of the target object and the real vertices corresponding to the ground bounding box markers, respectively. and These represent the validity probabilities of the updated target object vertex coordinates and the real vertices corresponding to the real bounding box markers, respectively. and These represent the weights of the corresponding losses, and the vertex validity classification loss, respectively. The Focal loss function is used.

[0096] Using the above scheme, after the detection and query module is trained in the first stage, the polygon vertex query module will be trained separately in the second stage, thereby obtaining an accurate polygon vertex query module trained in the second stage.

[0097] Step 1024: Input the sample data into the feature extraction backbone network and detection query module trained in the first stage, the ROI feature bridging module and the polygon vertex query module trained in the second stage, determine the total loss function, and perform joint training processing in the third stage. Jointly train and optimize the detection query module trained in the first stage and the polygon vertex query module trained in the second stage to obtain the trained target object bounding box extraction model.

[0098] In some embodiments, step 1024 includes: Step 10241: Input the sample data into the feature extraction backbone network and detection query module after the first stage of training, and determine the first loss function based on the output of the detection query module after the first stage of training.

[0099] Step 10242: The output of the detection query module after the first stage of training is processed by the ROI feature bridging module and the polygon vertex query module after the second stage of training, and the second loss function is determined based on the output of the polygon vertex query module after the second stage of training.

[0100] Step 10243: Add a second loss function with a set ratio to the first loss function to obtain the total loss function. Use the total loss function to jointly train and optimize the detection query module after the first stage of training and the polygon vertex query module after the second stage of training to obtain the detection query module after the third stage of training and the polygon vertex query module after the third stage of training. Combine the feature extraction backbone network and the ROI feature bridging module to obtain the trained target object bounding box extraction model.

[0101] In practical implementation, the first loss function is obtained. The process is the same as in the above embodiments, and the second loss function is obtained. The process is the same as in the above embodiments, and will not be repeated here.

[0102] Determine the total loss function: (10); among which, To set the ratio, the formula (10) is minimized, and then the detection query module after the first stage of training and the polygon vertex query module after the second stage of training are jointly trained in the third stage. In this way, the detection query module and the polygon vertex query module after the third stage of training can be obtained, and the target object bounding box extraction model is obtained by combining the feature extraction backbone network and the ROI feature bridging module.

[0103] During training, the outermost quadrilateral boundary corresponding to the candidate bounding boxes obtained by the detection query module and the multi-vertex set generated by the polygon vertex query module is fused to obtain the target object bounding box result, which can be represented as: (11); in, and They represent the first l Second and third l -1 times the candidate box of the detection query module This is the border correction amount obtained from the polygon vertex query module. To integrate weights, the influence of the correction suggestions from the polygon vertex query module on the candidate boxes obtained by the current detection query module is controlled.

[0104] In a preferred embodiment, the specific testing process in step 103 is as follows: The remote sensing images from the test set are input into the trained target object bounding box extraction model (i.e., the trained Det2Poly model). This model can complete the entire process from candidate region generation in the remote sensing image to target object vector boundary prediction in an end-to-end manner. Specifically, the trained model first automatically identifies the target object region through a detection branch, and then further parses the boundary geometry using a polygon vertex query module. Finally, it outputs the target object polygon vector bounding box (i.e., the target object bounding box prediction result) corresponding to the input remote sensing image. Evaluation of the output results on the test set effectively tests the detection accuracy and contour representation ability of the target object bounding box prediction result, demonstrating its good generalization performance and practical application value. After passing the test, the trained target object bounding box extraction model is used as the final target object bounding box extraction model.

[0105] The technical effects of this invention are as follows: (1) The detection query module is used for detection query. Dynamic convolution is used to realize efficient adaptive updating of candidate boxes, which reduces redundant candidates and maintains high accuracy and efficiency of detection in large-scale remote sensing images.

[0106] (2) Design an ROI feature bridging module to deeply integrate the candidate boxes of the detection query module with the local region features, ensuring the consistency and fine granularity of the input feature representation, and providing accurate support for subsequent polygon prediction.

[0107] (3) A polygon vertex query module is introduced, which combines the bridging feature constraint attention output by the ROI feature bridging module with Logit to effectively suppress redundant vertices and improve boundary geometry consistency and vertex positioning accuracy.

[0108] (4) Construct a target object bounding box extraction model (e.g., Det2Poly model) to achieve integrated processing of target object detection and vector boundary prediction in remote sensing images, avoid the fragmentation of traditional pipelines, and significantly improve the automation and reliability of building boundary extraction in remote sensing scenarios.

[0109] Det2Poly is an end-to-end deep learning-based object bounding box extraction model, primarily used for efficient object detection and high-precision polygon vector boundary extraction in remote sensing images. This Det2Poly innovatively integrates a dynamic object detection query mechanism with vertex sequence prediction capabilities: In the detection phase, learnable candidate boxes and dynamic instance interaction heads are used to adaptively optimize candidate bounding boxes; a ROI feature bridging module (integrating dynamic RoI assignment and RoIAlign operation) ensures accurate alignment of detection results with multi-scale features, providing fine-grained semantic information for boundary modeling; in the polygon prediction phase, an RoI-constrained attention mechanism is introduced to limit the computational range and reduce complexity, and a Logit embedding fusion strategy is used to dynamically suppress redundant vertex generation based on detection confidence; finally, a joint loss function is used to achieve end-to-end joint optimization of detection and boundary prediction, significantly improving geometric regularity and localization accuracy, and effectively avoiding error accumulation in multi-stage methods.

[0110] Furthermore, the technical framework adopted by Det2Poly has the potential for cross-domain applications, and can be extended to scenarios such as vector boundary extraction of road facilities and obstacles in autonomous driving environments, instance segmentation of cells or organs in medical images (such as fine extraction of tumor boundaries), and part size measurement and defect identification in industrial inspection (such as polygon annotation of surface cracks), providing high-precision and high-efficiency visual perception solutions for multiple industries.

[0111] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.

[0112] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0113] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides an end-to-end target object border extraction device.

[0114] refer to Figure 4 The device includes: The sample acquisition unit 201 is configured to acquire the original remote sensing image, mark the target objects in the original remote sensing image with real bounding boxes, and obtain sample data. The model training unit 202 is configured to construct a target object bounding box extraction model, use the target object bounding box extraction model to train the sample data, determine the total loss function based on the difference between the training output result and the real bounding box, and adjust the target object bounding box extraction model based on the loss value determined by the total loss function to obtain the trained target object bounding box extraction model. The test unit 203 is configured to test the trained target object bounding box extraction model, and after the trained target object bounding box extraction model passes the test, the trained target object bounding box extraction model is used as the final target object bounding box extraction model. Application unit 204 is configured to acquire the remote sensing image to be extracted, input the remote sensing image to be extracted into the final target object bounding box extraction model for bounding box prediction, and output the target object bounding box prediction result.

[0115] In some embodiments, the model training unit 202 is specifically configured as follows: Construct a target object bounding box extraction model that includes a feature extraction backbone network, a detection query module, a ROI feature bridging module, and a polygon vertex query module; The sample data is input into the feature extraction backbone network and the detection query module for the first stage of training. The detection query module is then trained and adjusted to obtain the feature extraction backbone network and the detection query module after the first stage of training. The sample data is input into the feature extraction backbone network and detection query module, the ROI feature bridging module and the polygon vertex query module after the first stage of training, and the second stage of training is performed to train and adjust the polygon vertex query module to obtain the polygon vertex query module after the second stage of training. The sample data is input into the feature extraction backbone network and detection query module trained in the first stage, the ROI feature bridging module and the polygon vertex query module trained in the second stage, the total loss function is determined, and the third stage of joint training is performed to jointly train and optimize the detection query module trained in the first stage and the polygon vertex query module trained in the second stage to obtain the trained target object bounding box extraction model.

[0116] In some embodiments, the model training unit 202 is further configured to: The feature extraction backbone network is used to perform multi-scale feature extraction processing on the original remote sensing images in the sample data to obtain multi-scale feature maps. The detection query module is used to determine candidate boxes based on the multi-scale feature map; The ROI feature bridging module is used to perform ROI feature bridging processing on the candidate box to obtain bridging features. The polygon vertex query module is used to determine multiple vertices based on the bridging features and to determine the bounding box based on the multiple vertices.

[0117] In some embodiments, the original remote sensing image is divided according to height into: low layer with height less than or equal to a first height value, mid-low layer with height greater than the first height but less than or equal to the second height, mid-high layer with height greater than the second height but less than or equal to the third height, and high layer with height greater than the third height. The feature extraction backbone network consists of five layers; First layer: For the original remote sensing image in the sample data, the first convolution kernel combined with the first residual convolution module is used to extract the texture features of the lower layer; The second layer: For the original remote sensing images in the sample data, the first semantic features of the shape contour and basic structure information of local targets in the middle and low layers are extracted by using the second convolution kernel combined with the second residual convolution module. The third layer: For the original remote sensing images in the sample data, the third convolutional kernel is used in combination with the ResNet module and attention mechanism to extract the structural features of the target object; Fourth layer: For the original remote sensing images in the sample data, the fourth convolution kernel combined with the fourth residual convolution module is used to extract the second semantic features of scene relationships and object combinations in the middle and upper layers; Fifth layer: For the original remote sensing images in the sample data, the fifth convolutional kernel is used in combination with the ResNet module and the feature pyramid network to extract the third semantic features that distinguish the overall semantics of the scene from the scene category. The feature extraction backbone network combines the first semantic feature, the target object structural feature, the second semantic feature, and the third semantic feature to form a multi-scale feature map output.

[0118] In some embodiments, the model training unit 202 is further configured to: The sample data is input into the feature extraction backbone network, and multi-scale feature extraction processing is performed on the original remote sensing image in the sample data to obtain a multi-scale feature map. Multiple initial candidate boxes are determined in the detection query module. Based on the stride of the multi-scale feature map, the coordinates of each initial candidate box are mapped to the corresponding level of the multi-scale feature map to obtain the mapped candidate box. Execute using the aforementioned detection and query module: Determine the offset of the mapped candidate box, update the mapped candidate box according to the offset to obtain the updated candidate box, and normalize the updated candidate box to obtain the normalized candidate box. The normalized candidate boxes and the true bounding boxes labeled in the sample data are used to determine the first loss function. The detection query module is then trained and adjusted based on the first loss value determined by the first loss function to obtain the feature extraction backbone network and the detection query module after the first stage of training.

[0119] In some embodiments, the model training unit 202 is further configured to: The sample data is input into the feature extraction backbone network trained in the first stage, and multi-scale feature extraction processing is performed on the original remote sensing image in the sample data to obtain a multi-scale feature map. The detection query module, trained in the first stage, determines candidate boxes based on the multi-scale feature map. The ROI feature bridging module is used to process the ROI features of the candidate box to obtain ROI features. The ROI features are then flattened and input into the encoder for encoding to obtain bridging features. Execute using the polygon vertex query module: Determine the coordinate reference point and the embedding vector, and determine the position embedding of each vertex for the coordinate reference point; A scalar representing whether a vertex is valid is determined, and the scalar is processed by a multilayer perceptron and then converted into adaptively normalized modulation parameters. The modulation parameters, the position embedding, and the embedding amount are used to determine the initial vertex query vector; Based on the bridging features and the previously obtained vertex query vector, cross-attention processing is performed to obtain the updated vertex query vector for the current iteration. The updated vertex query vector is input into the decoder to obtain the coordinate offset and validity probability. The updated vertex query vector is then normalized to obtain the updated target object vertex coordinates. A second loss function is determined based on the difference between the updated target object vertex coordinates and the real vertices corresponding to the real bounding box markers. The parameters of the polygon vertex query module are adjusted based on the second loss value determined by the second loss function to obtain the polygon vertex query module after the second stage of training.

[0120] In some embodiments, the model training unit 202 is further configured to: The sample data is input into the feature extraction backbone network and detection query module after the first stage of training, and the first loss function is determined based on the output of the detection query module after the first stage of training. The output of the detection query module after the first stage of training is processed by the ROI feature bridging module and the polygon vertex query module after the second stage of training, and the second loss function is determined based on the output of the polygon vertex query module after the second stage of training. The first loss function is added to a second loss function with a set ratio to obtain the total loss function. The total loss function is then used to jointly train and optimize the detection query module after the first stage of training and the polygon vertex query module after the second stage of training to obtain the detection query module and the polygon vertex query module after the third stage of training. Combined with the feature extraction backbone network and the ROI feature bridging module, the trained target object bounding box extraction model is obtained.

[0121] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.

[0122] The apparatus of the above embodiments is used to implement the corresponding method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0123] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the methods described in any of the above embodiments.

[0124] Figure 5This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0125] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0126] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0127] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0128] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0129] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0130] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0131] The electronic devices described above are used to implement the corresponding methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0132] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the methods described in any of the above embodiments.

[0133] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0134] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the methods described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0135] Based on the same concept, corresponding to any of the above embodiments, this application also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to perform the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0136] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.

[0137] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.

[0138] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0139] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0140] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application is limited to these examples; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0141] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0142] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0143] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the claims of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.

Claims

1. A method for extracting the bounding box of an end-to-end target object, characterized in that, include: Acquire the original remote sensing image, mark the target objects in the original remote sensing image with real bounding boxes, and obtain sample data; A target object bounding box extraction model is constructed, and the sample data is trained using the target object bounding box extraction model. Based on the difference between the training output result and the real bounding box, a total loss function is determined. The target object bounding box extraction model is adjusted based on the loss value determined by the total loss function to obtain the trained target object bounding box extraction model. The trained target object bounding box extraction model is tested, and after the trained target object bounding box extraction model passes the test, it is used as the final target object bounding box extraction model. The remote sensing image to be extracted is acquired, and the remote sensing image to be extracted is input into the final target object bounding box extraction model for bounding box prediction, and the target object bounding box prediction result is output.

2. The method according to claim 1, characterized in that, The process of constructing a target object bounding box extraction model, training the sample data using the target object bounding box extraction model, determining a total loss function based on the difference between the training output and the real bounding box, and adjusting the target object bounding box extraction model based on the loss value determined by the total loss function to obtain the trained target object bounding box extraction model includes: Construct a target object bounding box extraction model that includes a feature extraction backbone network, a detection query module, a ROI feature bridging module, and a polygon vertex query module; The sample data is input into the feature extraction backbone network and the detection query module for the first stage of training. The detection query module is then trained and adjusted to obtain the feature extraction backbone network and the detection query module after the first stage of training. The sample data is input into the feature extraction backbone network and detection query module, the ROI feature bridging module and the polygon vertex query module after the first stage of training, and the second stage of training is performed to train and adjust the polygon vertex query module to obtain the polygon vertex query module after the second stage of training. The sample data is input into the feature extraction backbone network and detection query module trained in the first stage, the ROI feature bridging module and the polygon vertex query module trained in the second stage, the total loss function is determined, and the third stage of joint training is performed to jointly train and optimize the detection query module trained in the first stage and the polygon vertex query module trained in the second stage to obtain the trained target object bounding box extraction model.

3. The method according to claim 2, characterized in that, The feature extraction backbone network is used to perform multi-scale feature extraction processing on the original remote sensing images in the sample data to obtain multi-scale feature maps. The detection query module is used to determine candidate boxes based on the multi-scale feature map; The ROI feature bridging module is used to perform ROI feature bridging processing on the candidate box to obtain bridging features. The polygon vertex query module is used to determine multiple vertices based on the bridging features and to determine the bounding box based on the multiple vertices.

4. The method according to claim 3, characterized in that, The original remote sensing image is divided into the following layers based on height: low layer with height less than or equal to the first height value, middle and low layer with height greater than the first height but less than or equal to the second height, middle and high layer with height greater than the second height but less than or equal to the third height, and high layer with height greater than the third height. The feature extraction backbone network consists of five layers; First layer: For the original remote sensing image in the sample data, the first convolution kernel combined with the first residual convolution module is used to extract the texture features of the lower layer; The second layer: For the original remote sensing images in the sample data, the first semantic features of the shape contour and basic structure information of local targets in the middle and low layers are extracted by using the second convolution kernel combined with the second residual convolution module. The third layer: For the original remote sensing images in the sample data, the third convolutional kernel is used in combination with the ResNet module and attention mechanism to extract the structural features of the target object; Fourth layer: For the original remote sensing images in the sample data, the fourth convolution kernel combined with the fourth residual convolution module is used to extract the second semantic features of scene relationships and object combinations in the middle and high layers; Fifth layer: For the original remote sensing images in the sample data, the fifth convolutional kernel is used in combination with the ResNet module and the feature pyramid network to extract the third semantic features that distinguish the overall semantics of the scene from the scene category. The feature extraction backbone network combines the first semantic feature, the target object structural feature, the second semantic feature, and the third semantic feature to form a multi-scale feature map output.

5. The method according to claim 3, characterized in that, The process of inputting the sample data into the feature extraction backbone network and the detection query module for a first-stage training process, and training and adjusting the detection query module to obtain the feature extraction backbone network and detection query module after the first-stage training, includes: The sample data is input into the feature extraction backbone network, and multi-scale feature extraction processing is performed on the original remote sensing image in the sample data to obtain a multi-scale feature map. Multiple initial candidate boxes are determined in the detection query module. Based on the stride of the multi-scale feature map, the coordinates of each initial candidate box are mapped to the corresponding level of the multi-scale feature map to obtain the mapped candidate box. Execute using the aforementioned detection and query module: Determine the offset of the mapped candidate box, update the mapped candidate box according to the offset to obtain the updated candidate box, and normalize the updated candidate box to obtain the normalized candidate box. The normalized candidate boxes and the true bounding boxes labeled in the sample data are used to determine the first loss function. The detection query module is then trained and adjusted based on the first loss value determined by the first loss function to obtain the feature extraction backbone network and the detection query module after the first stage of training.

6. The method according to claim 3, characterized in that, The process of inputting the sample data into the feature extraction backbone network and detection query module trained in the first stage, the ROI feature bridging module, and the polygon vertex query module for the second stage of training, and adjusting the polygon vertex query module to obtain the polygon vertex query module trained in the second stage, includes: The sample data is input into the feature extraction backbone network trained in the first stage, and multi-scale feature extraction processing is performed on the original remote sensing image in the sample data to obtain a multi-scale feature map. The detection query module, trained in the first stage, determines candidate boxes based on the multi-scale feature map. The ROI feature bridging module is used to process the ROI features of the candidate box to obtain ROI features. The ROI features are then flattened and input into the encoder for encoding to obtain bridging features. Execute using the polygon vertex query module: Determine the coordinate reference point and the embedding vector, and determine the position embedding of each vertex for the coordinate reference point; A scalar representing whether a vertex is valid is determined, and the scalar is processed by a multilayer perceptron and then converted into adaptively normalized modulation parameters. The modulation parameters, the position embedding, and the embedding amount are used to determine the initial vertex query vector; Based on the bridging features and the previously obtained vertex query vector, cross-attention processing is performed to obtain the updated vertex query vector for the current iteration. The updated vertex query vector is input into the decoder to obtain the coordinate offset and validity probability. The updated vertex query vector is then normalized to obtain the updated target object vertex coordinates. A second loss function is determined based on the difference between the updated target object vertex coordinates and the real vertices corresponding to the real bounding box markers. The parameters of the polygon vertex query module are adjusted based on the second loss value determined by the second loss function to obtain the polygon vertex query module after the second stage of training.

7. The method according to claim 3, characterized in that, The sample data is input into the feature extraction backbone network and detection query module trained in the first stage, the ROI feature bridging module, and the polygon vertex query module trained in the second stage. The total loss function is determined, and a third stage of joint training is performed. The detection query module trained in the first stage and the polygon vertex query module trained in the second stage are jointly trained and optimized to obtain the trained target object bounding box extraction model, including: The sample data is input into the feature extraction backbone network and detection query module after the first stage of training, and the first loss function is determined based on the output of the detection query module after the first stage of training. The output of the detection query module after the first stage of training is processed by the ROI feature bridging module and the polygon vertex query module after the second stage of training, and the second loss function is determined based on the output of the polygon vertex query module after the second stage of training. The first loss function is added to a second loss function with a set ratio to obtain the total loss function. The total loss function is then used to jointly train and optimize the detection query module after the first stage of training and the polygon vertex query module after the second stage of training to obtain the detection query module and the polygon vertex query module after the third stage of training. Combined with the feature extraction backbone network and the ROI feature bridging module, the trained target object bounding box extraction model is obtained.

8. A device for extracting the outline of an end-to-end target object, characterized in that, include: The sample acquisition unit is configured to acquire the original remote sensing image, mark the target objects in the original remote sensing image with real bounding boxes, and obtain sample data. The model training unit is configured to construct a target object bounding box extraction model, use the target object bounding box extraction model to train the sample data, determine the total loss function based on the difference between the training output and the real bounding box, and adjust the target object bounding box extraction model based on the loss value determined by the total loss function to obtain the trained target object bounding box extraction model. The testing unit is configured to test the trained target object bounding box extraction model, and after the trained target object bounding box extraction model passes the test, the trained target object bounding box extraction model is used as the final target object bounding box extraction model. The application unit is configured to acquire the remote sensing image to be extracted, input the remote sensing image to be extracted into the final target object bounding box extraction model for bounding box prediction, and output the target object bounding box prediction result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 7.