Method for Extracting Vector Polygon Contours of Buildings Based on Multimodal Large Models
By applying the prediction-clipping collaborative training strategy of multimodal large models in building extraction, the problems of boundary jagging, topological errors and poor robustness of building extraction in remote sensing images are solved, and efficient and robust vector polygon contour extraction is achieved.
Patent Information
- Application Number
- CN202510360726.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing building extraction methods face the problems of boundary jagging, topological errors and poor robustness when dealing with remote sensing images in complex high-resolution scenarios. The application of multimodal large models in remote sensing images has the problems of sparse characteristics of small buildings and computational redundancy.
The prediction-clipping collaborative training strategy based on multimodal large models is adopted, and the remote sensing image and text instructions of the building are obtained through visual encoder, text instruction encoder, multimodal feature alignment module and large language model, and the remote sensing image and text instructions of the building are generated to generate a sequence of corner points coordinates of the building and connect corner points to form a vector polygon outline.
It realizes efficient and robust extraction of building vector polygonal profiles in complex scenarios, improves the feature capture ability of small-scale buildings, and avoids complex post-processing steps in traditional methods.
Smart Images

Figure CN119888257B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image contour extraction, and in particular to a method, device, storage medium and electronic device for extracting vector polygon contours of buildings based on a multimodal large model. Background Art
[0002] Building extraction is a core task in remote sensing image interpretation and geographic information modeling, and plays a key role in urban planning, land resource management, and smart city construction. The current mainstream building extraction methods mainly follow three technical paths: mask-based segmentation network, contour-based detection model, and geometric modeling method based on primitives.
[0003] However, traditional methods face significant challenges when processing remote sensing images of high-resolution complex scenes: although mask-based methods can achieve pixel-level building area recognition, the generated binary masks need to rely on complex processes such as edge extraction and morphological post-processing to be converted into vector contours, which can easily produce boundary jagged edges and broken polygons; although contour-based methods can directly predict boundary points, they do not adequately model the adjacency relationships of dense building complexes, which can easily lead to contour intersections or topological errors; although primitive-based methods can generate vector vertices, they are less robust to small buildings and occluded scenes, and the vertex positioning accuracy is easily affected by noise.
[0004] On the other hand, with the breakthroughs of multimodal large models (MLLMs) in cross-modal understanding and generation tasks, their ability to combine visual-language representation provides new ideas for building extraction. However, existing research has mostly focused on the field of natural images, and there are significant limitations when directly applying MLLMs to remote sensing images of multiple buildings: the scales of buildings in remote sensing images vary greatly, and densely arranged small buildings are easily submerged by large-scale targets in global features, resulting in feature sparsity. Although some recent studies have attempted to improve the accuracy of small target detection through ROI cropping, the scale adaptation problem between the cropped area and the model input has not yet been solved, and direct scaling will lead to loss of details or computational redundancy. Therefore, studying and implementing end-to-end high-precision vector polygon generation while meeting the efficiency and robustness requirements in complex scenarios is still a technical difficulty that needs to be overcome. Summary of the invention
[0005] The embodiments of the present application provide a method, device, storage medium and electronic device for extracting vector polygon outlines of buildings based on a multimodal large model, which can meet the efficiency and robustness requirements in complex scenarios and improve the ability to capture features of small-scale buildings.
[0006] The embodiment of the present application provides a method for extracting building vector polygon outlines based on a multimodal large model, comprising:
[0007] Obtain the remote sensing image and text instructions of the building;
[0008] Input the remote sensing image and the text instructions into the trained multi-modal large model to obtain a sequence of building polygon corner coordinates, and connect the corners to form a building vector polygon contour;
[0009] Among them, the multi-modal large model is trained through a prediction-cropping collaborative training strategy.
[0010] Furthermore, for the above method for extracting the building vector polygon contour based on the multi-modal large model, where the multi-modal large model includes a visual encoder, a text instruction encoder, a multi-modal feature alignment module, and a large language model;
[0011] The step of inputting the remote sensing image and the text instructions into the trained multi-modal large model to obtain a sequence of building polygon corner coordinates includes:
[0012] Input the remote sensing image into the visual encoder for multi-scale feature extraction to obtain visual features;
[0013] Input the text instructions into the text instruction encoder for deep semantic encoding to obtain text features;
[0014] Input the visual features into the multi-modal feature alignment module to map the visual features to the language space consistent with the large language model;
[0015] Input the mapped visual features and the text features into the large language model to obtain a sequence of building polygon corner coordinates.
[0016] Furthermore, for the above method for extracting the building vector polygon contour based on the multi-modal large model, after the step of inputting the remote sensing image into the visual encoder for multi-scale feature extraction to obtain visual features, it includes:
[0017] Generate position embedding information based on the remote sensing image;
[0018] Concatenate the position embedding information with the visual features to obtain enhanced visual features.
[0019] Furthermore, for the above method for extracting the building vector polygon contour based on the multi-modal large model, where the text instruction encoder includes a sub-word tokenizer and multiple self-attention layers;
[0020] The step of inputting the text instructions into the text instruction encoder for deep semantic encoding to obtain text features includes:
[0021] Input the text instruction into the sub-word tokenizer to obtain a sequence of tokens;
[0022] Input the sequence of tokens into multiple self-attention layers in sequence, and obtain text features after self-attention transformation.
[0023] Further, in the above method for extracting the vector polygon contour of a building based on a multimodal large model, wherein the mapping of the visual features to the language space consistent with the large language model includes:
[0024] Map the visual features to the language space consistent with the large language model through a projector; wherein, the projector is a two-layer multi-layer perceptron;
[0025] The mapping process of the projector is represented by the following formula:
[0026]
[0027] Wherein, 、 、 、 and are all projector parameters, is an activation function.
[0028] Further, in the above method for extracting the vector polygon contour of a building based on a multimodal large model, wherein the large language model includes a tokenizer, and the method further includes:
[0029] Expand the tokenizer, and add pixel feature markers and position feature markers. The pixel feature markers are used to indicate the pixel features of the visual features, and the position feature markers are used to mark the position features of the remote sensing image.
[0030] Further, in the above method for extracting the vector polygon contour of a building based on a multimodal large model, wherein the process of training the multimodal large model through the prediction-cropping co-training strategy includes:
[0031] Annotate the center points and bounding box ranges of each building in the remote sensing image as the first training sample set;
[0032] Input the images in the first training sample set into the multimodal large model. In the first step, predict the center point coordinates of each building on the remote sensing image, and in the second step, predict the bounding box range of each building based on the center point coordinates of the building;
[0033] Based on the ground truth of the bounding box range of each building, crop the remote sensing image to obtain a local area containing only a single building, and form a second training sample set based on the local area containing the single building and the ground truth contour;
[0034] Based on the true contour coordinates corresponding to the second training sample set, freeze the parameters of the visual encoder and the text encoder of the multi-modal large model, and train the remaining modules of the multi-modal large model until convergence.
[0035] An embodiment of the present application further provides a device for extracting a building vector polygon contour based on a multi-modal large model, including:
[0036] An acquisition module, configured to acquire a remote sensing image and a text instruction of a building;
[0037] A vector polygon contour extraction module, configured to input the remote sensing image and the text instruction into the trained multi-modal large model to obtain a sequence of building polygon corner coordinates, and connect the corner points to form a building vector polygon contour;
[0038] Wherein, the multi-modal large model is trained by a prediction-cropping collaborative training strategy.
[0039] An embodiment of the present application further provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded by a processor to execute any one of the above-mentioned methods for extracting a building vector polygon contour based on a multi-modal large model.
[0040] An embodiment of the present application further provides an electronic device, including a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used for the steps in any one of the above-mentioned methods for extracting a building vector polygon contour based on a multi-modal large model.
[0041] The building vector polygon contour extraction method, device, storage medium, and electronic device based on a multimodal large model provided by this application extract the vector polygon contour of a building through the multimodal large model and train the multimodal large model using a prediction-cropping collaborative training strategy. The following are achieved: (1) Without a complex post-processing process, a simple, end-to-end trainable model is used to avoid the morphological post-processing steps required for mask vectorization in traditional methods, complete the generation of the building polygon corner coordinate sequence, and significantly improve efficiency. (2) It has strong semantic understanding ability, is driven by multimodal instructions, combines a visual encoder and a text instruction encoder, and achieves cross-modal fusion through MLP feature alignment, enabling the model to accurately understand semantic instructions such as "sequentially click on the corners" and imitate the human drawing logic. (3) It has scalability. The trained neural network model can be adjusted and applied to other uses, such as land cover change detection based on remote sensing images, extraction of objects of interest based on remote sensing images, etc. (4) It has strong robustness. An adaptive prediction-cropping strategy and a dynamic scaling cropping mechanism are adopted to enhance the model's ability to capture features of small-scale buildings, and good extraction results can be obtained for various ground object elements in remote sensing images of dense building clusters. Brief Description of the Drawings
[0042] The following, in conjunction with the drawings, through a detailed description of the specific embodiments of this application, will make the technical solutions and other beneficial effects of this application obvious.
[0043] Figure 1 It is a flowchart of the building vector polygon contour extraction method based on a multimodal large model provided by an embodiment of this application.
[0044] Figure 2 It is another flowchart of the building vector polygon contour extraction method based on a multimodal large model provided by an embodiment of this application.
[0045] Figure 3 It is the prediction result of the method of this invention provided by an embodiment of this application and the current best P2PFormer method on the WHU dataset.
[0046] Figure 4 It is the prediction result of the building contour extraction method of this invention provided by an embodiment of this application on the WHU and WHU-Mix datasets.
[0047] Figure 5 It is a schematic structural diagram of the building vector polygon contour extraction device based on a multimodal large model provided by an embodiment of this application.
[0048] Figure 6 It is a schematic structural diagram of the electronic device provided by an embodiment of this application. Detailed Description of the Embodiments
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0050] The embodiments of the present application provide a method, device, storage medium, and electronic device for extracting the vector polygon contour of a building based on a multimodal large model. A device for extracting the vector polygon contour of a building based on a multimodal large model provided by the embodiments of the present application can be integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0051] Please refer to Figure 1 and Figure 2 , Figure 1 which is a flowchart of the method for extracting the vector polygon contour of a building based on a multimodal large model provided by the embodiments of the present application. Figure 2 which is another flowchart of the method for extracting the vector polygon contour of a building based on a multimodal large model provided by the embodiments of the present application. It is applied to an electronic device, and the method for extracting the vector polygon contour of a building based on a multimodal large model includes the following steps:
[0052] S1, Obtain the remote sensing image and text instruction of the building.
[0053] S2, Input the remote sensing image and the text instruction into the trained multimodal large model to obtain a sequence of building polygon corner coordinates, and connect the corner points to form the vector polygon contour of the building;
[0054] Among them, the multimodal large model is trained through a prediction-cropping co-training strategy.
[0055] In one embodiment, the multimodal large model includes a visual encoder, a text instruction encoder, a multimodal feature alignment module, and a large language model.
[0056] Step S2 includes the following steps:
[0057] S21, Input the remote sensing image into the visual encoder for multi-scale feature extraction to obtain visual features.
[0058] After step S21, it further includes:
[0059] Generate position embedding information based on the remote sensing image;
[0060] Concatenate the position embedding information with the visual features to obtain enhanced visual features.
[0061] Specifically, the visual encoder uses the RADIO visual encoder based on the Transformer architecture, which is responsible for processing the input remote sensing image and capturing the spatial information and feature representation in the image. Its core is the multi-head self-attention mechanism, which can automatically capture the spatial relationships and high-level semantic features in the remote sensing image. In this process, first, the remote sensing image is resized to a fixed size of 224×224 as the input of the visual encoder. After passing through the visual encoder, visual features , are high-level semantic features with a downsampling factor of 16 (the input remote sensing image is divided into non-overlapping blocks of 16×16 pixels, and multi-scale features are extracted through the multi-head self-attention mechanism, and a high-level semantic feature map with a resolution of 1 / 16 downsampling is output). Subsequently, the learnable position embedding information (positional embeddings) is concatenated with the visual features to obtain an enhanced visual feature representation . This feature combines spatial information and visual features and has the ability to analyze and locate the target area.
[0062] In S22, input the text instruction into the text instruction encoder for deep semantic encoding to obtain text features.
[0063] Step S22 includes the following steps:
[0064] S221, input the text instruction into the sub-word tokenizer to obtain a sequence of word tokens;
[0065] S222, input the sequence of word tokens into multiple self-attention layers in turn, and obtain text features after self-attention transformation.
[0066] The text instruction encoder encodes the input text instruction into high-level features. The encoder used in this module is Qwen2.5 tokenizer --- byte-level byte-pair encoding BBPE, which is constructed based on a pre-trained language model and uses a multi-layer Transformer architecture to perform deep semantic encoding on the task instruction.
[0067] Input the given text instruction X q(such as "extracting the vertex coordinates of building polygons"), generate a sequence of tokens through the sub-word tokenizer Tokenizer(·), and after L layers of self-attention transformation, output text tokens ( ), where d is the hidden layer dimension, and l is the sequence length.
[0068] S23, input the visual features into the multi-modal feature alignment module to map the visual features to the language space consistent with the large language model.
[0069] Specifically, the multi-modal feature alignment module plays a role in connecting the visual encoder and the language model in the model, and is responsible for mapping the visual features to the hidden space consistent with the language model to ensure the effective alignment of multi-modal information.
[0070] Map the visual features to the language space consistent with the large language model through a projector, enabling the large language model to understand image information; among them, the projector is a two-layer multi-layer perceptron (MLP).
[0071] The mapping process of the projector is represented by the following formula:
[0072]
[0073] Among them, , , , and are all projector parameters, and is the activation function.
[0074] S24, input the mapped visual features and text features into the large language model to obtain the sequence of building polygon corner coordinates.
[0075] The large language model is pre-trained based on a large amount of text data and has powerful language understanding, logical reasoning, and instruction-following capabilities. In multi-modal tasks, the large language model can achieve precise parsing of the joint input of images and text through deep fusion with visual features. This model uses Qwen2.5-1.5B as the pre-trained large language model.
[0076] Further, the large language model includes a tokenizer. The tokenizer is extended to add pixel feature markers and position feature markers. The pixel feature markers are used to indicate the pixel features of the visual features, and the position feature markers are used to mark the position features of the remote sensing images.
[0077] Specifically, to enable the large language model to effectively understand and process multimodal inputs from the visual encoder, it is necessary to extend its tokenizer and add a series of specially designed special tokens, mainly including:
[0078] Encode structured text instructions (such as "Extract the vertices of the building polygon in clockwise order"); Add the following special tokens by extending the vocabulary:
[0079] <pixel>: Pixel feature markers, used to indicate image pixel features from the visual encoder, helping the model to recognize and process visual information.
[0080] <x>and <y>: Location feature editing, representing the coordinate axes in the image, used to mark features at different positions in the image, from <x0>To <x(W-1)> represents the abscissa, starting from <y0>To <y(H-1)> represents the ordinate, with a total of H+W markers. The embedding layer of the language model is adjusted to adapt to the newly added markers. After introducing the above special markers, the vocabulary of the model is expanded to ensure the alignment and fusion of multimodal features. In the building polygon extraction task of this invention, a pair of <x> <y>The markers are used to record the central point coordinates and corner point coordinates of the building. By predicting these coordinates in sequence, an accurate building polygon is obtained. The large language model reasons based on the input visual features and text features, generates an output and decodes it into the corresponding answer, including the central point coordinates, the bounding box range, and the building corner point coordinates.
[0081] Furthermore, the prediction-cropping strategy is to solve the problem of insufficient feature expression of small buildings. This strategy is completed in the same framework as the multimodal large model, and we designed a three-step strategy. For a single building in an image, first predict the center point (center) of each building through the multimodal large model. The dialogue is as follows:
[0082] "Please extract the coordinates of the center point of each building in the image."
[0083] " <x31> <y165> <x58> <y67> <x89> <y112> <x24> <y113> <x27> <y67> <x53> <y47> <x29> <y112> <x35> <y161>". Then, based on the predicted center point, further predict the bounding box of each building, determine its width and height, and accurately locate the scope of a single building. The dialogue is as follows:
[0084] "Given the center point of the building, give the x-axis length range and y-axis length range of the area where the building is located: <x31> <y56>。(The two length ranges together can form a building bounding box with the center point as the origin)"
[0085] "x: 24 y: 13". Then, according to the predicted building bounding box, the original image is cropped to extract the local area containing a single building. To enhance the model's perception of small buildings, the cropped area is not fixed to the bounding box range but randomly enlarged by 1.5 - 2 times to ensure that the model can fully capture the target features. Then, the image containing only a single building after cropping is adjusted to a fixed size of 224×224 and fed into the multi-modal model for refined prediction, parsing and outputting the sequence of building polygon corner coordinates from the remote sensing image, and connecting the corner points to form the building vector polygon contour. Finally, based on the bounding box information of the building in the complex image and the corner coordinates in the cropped image, the spatial mapping of the building corner points from the local image to the original image is completed to ensure the accurate positioning of the prediction result in the original image.
[0086] The prediction-cropping co-training strategy realizes the extraction of buildings from coarse to fine through three-step cascaded building center prediction, bounding box prediction, and precise prediction of the contour of a single building. After training, by loading the weights of the trained network model to predict new remote sensing images, the accurate vector polygon contour extraction of multi-building complex remote sensing images can be achieved.
[0087] The process of training a multi-modal large model through the prediction-cropping co-training strategy includes:
[0088] A1, Mark the center point and the bounding box range of each building in the remote sensing image as the first training sample set.
[0089] Specifically, the original large-scale remote sensing image is cropped into the standard size supported by the network (224×224 pixels), and the center point of each building is marked to generate a sample set for training the prediction of the building center point. Then, based on the ground truth of the center point, the local area containing a single building is cropped and scaled (randomly enlarged by 1.5 - 2 times) to generate refined polygon corner coordinates and vector ground truth annotations, forming the bounding box prediction sample set.
[0090] A2, Input the images in the first training sample set into the multi-modal large model. In the first step, predict the center point coordinates of each building on the remote sensing image. In the second step, predict the bounding box range of each building based on the building center point coordinates.
[0091] Input the first training sample set in step A1 (including the center point prediction sample and the bounding box prediction sample) into the multi-modal large model in sequence to train the model to output the center point coordinates and the bounding box range of each building.
[0092] A3. Based on the ground truth of the bounding box range of each building, the remote sensing image is cropped to obtain a local area containing only a single building. A second training sample set is constructed based on the local area containing the single building and the ground truth contour.
[0093] A4. Based on the ground truth contour coordinates corresponding to the second training sample set, the parameters of the visual encoder and the text encoder of the multi-modal large model are frozen, and the remaining modules of the multi-modal large model are trained until convergence.
[0094] Based on the samples generated in step A1, the parameters of the visual encoder (RADIO) and the text encoder (Qwen2.5 - 1.5B) are frozen, and only the feature projector and the center point prediction module are trained. The AdamW optimizer (learning rate 5e-5) is used, and the loss function is the mean squared error (MSE) of the heat map. All linear layers in the language model and the visual encoder are fine-tuned through the LoRA technique to train their output of the building polygon corner coordinates until the overall model converges.
[0095] The following provides a specific embodiment:
[0096] First, construct a building vector polygon contour extraction framework based on the multi-modal large model according to the method of the present invention. Subsequently, obtain the training sample data and use the sample data to train the network model. The sample data used in the embodiment is the WHU and WHU-Mix building extraction data sets. The WHU data set contains 2793 training images, 627 validation images, and 2220 test images, with an image resolution of 1024×1024 pixels; the WHU-Mix data set contains a large-scale aerial and satellite image data set with building boundary annotations in COCO format, sourced from multiple regions around the world, containing 43778 training images, 11675 test images (Test1), and 6011 cross-domain test images (Test2).
[0097] All input images are uniformly adjusted to 224×224 pixels to ensure data consistency and adapt to the input requirements of the visual encoder, and then input into the multi-modal large model for iterative training until the model converges to obtain the optimal weight file. During the training process, the loss output by the language model is used as the loss function to ensure the target consistency of the model in multi-modal learning. After the model training is completed, the test remote sensing image to be predicted is input into the multi-modal large model loaded with the training weights to verify the performance of the model in the building extraction task.
[0098] In order to verify the effectiveness and advancement of the method of the present invention, we compared the proposed method with the current mainstream building extraction algorithms. Including Mask-RCNN, QueryInst, YOLACT, SOLO, E2EC, P2Pformer and Line2Poly algorithms that have outstanding performance in building extraction tasks. All methods are trained on the same hardware environment (four NVIDIAV100 GPUs, 16GB of video memory per card, distributed training) using the same training dataset and evaluated on the WHU test set and WHU-Mix. The prediction results of all methods are based on the mean average precision (mAP), AP at IoU=0.5 (AP 50 ) and AP (AP) when IoU=0.75 75 ) are used as the main evaluation indicators for quantitative evaluation and are recorded in Table 1 (WHU dataset) and Table 2 (WHU-Mix dataset).
[0099] Table 1 Comparison of the accuracy of the proposed method and other advanced building extraction methods in the WHU dataset
[0100]
[0101] Among them, "VECTOR" indicates whether the network can output vector building outlines; "PF" indicates whether post-processing is required.
[0102] Table 2 Comparison of the accuracy of the proposed method and other advanced building extraction methods in the WHU-Mix dataset
[0103]
[0104] Among them, "VECTOR" indicates whether the network can output vector building outlines; "PF" indicates whether post-processing is required.
[0105] From Table 1, we can see that on the WHU test set, the mAP of the proposed method reaches 74.8%, which is 2.1% higher than the second best method P2PFormer (72.7%). 75 The indicator is improved by 1.7%; on the WHU-Mix test set, the overall prediction accuracy of the proposed method is improved by 3 AP on test1 and 0.7 AP on test2 compared with the suboptimal method P2Pformer. In addition, the AP under the high threshold matching standard is 75 In terms of indicators, our method has improved by 12 points on test1 and 13.2 points on test2. Compared with these existing methods, it is proved that the method of the present invention has better robustness and can obtain more accurate building contour vector prediction results. Therefore, the method of the present invention has good engineering practical value.
[0106] Figure 3 This is the prediction result of the method of the present invention provided by the embodiment of the present application and the current best P2PFormer method on the WHU dataset. In Figure 3 the first row is the building contour drawn by our method, and the second row is the drawing result of the P2PFormer method. As can be seen from Figure 3 it, the prediction result of the method of the present invention is more accurate, while the P2PFormer method is more prone to topological errors and building missing detection. Our method is more accurate in predicting the contour of buildings, has accurate topological relationships, and has fewer missing detection cases. Generally speaking, it has more correct prediction results. This proves the innovation and effectiveness of the present invention and has good engineering practical value. Figure 4 This is the prediction result of the building contour extraction method of the present invention provided by the embodiment of the present application on the WHU and WHU-Mix datasets. As can be seen from Figure 4 it, our method has accurate prediction, accurate topological relationships, and few missing detection cases.
[0107] According to the method described in the above embodiment, this embodiment will further describe from the perspective of a building vector polygon contour extraction device based on a multi-modal large model. The building vector polygon contour extraction device based on a multi-modal large model can be specifically implemented as an independent entity or integrated in an electronic device. The electronic device can be a device such as a terminal, a server, etc. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a micro processing box, or other devices, etc.
[0108] Please refer to Figure 5 Figure 5 which specifically describes the building vector polygon contour extraction device based on a multi-modal large model provided by the embodiment of the present application and applied to an electronic device. The building vector polygon contour extraction device based on a multi-modal large model can include:
[0109] An acquisition module, configured to acquire a remote sensing image and a text instruction of a building;
[0110] A vector polygon contour extraction module, configured to input the remote sensing image and the text instruction into a trained multi-modal large model to obtain a sequence of building polygon corner coordinates, and connect the corner points to form a building vector polygon contour;
[0111] Among them, the multi-modal large model is trained by a prediction-cropping collaborative training strategy.
[0112] In specific implementation, each of the above modules and / or units can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above modules and / or units, please refer to the previous method embodiments. For the specific beneficial effects that can be achieved, please also refer to the beneficial effects in the previous method embodiments, which will not be elaborated here.
[0113] In addition, an embodiment of the present application further provides an electronic device, which can be a device such as a computer or a tablet computer. The electronic device can implement the steps in any embodiment of the method for extracting the building vector polygon contour based on the multi-modal large model provided by the embodiments of the present application. Therefore, it can achieve the beneficial effects that can be achieved by any method for extracting the building vector polygon contour based on the multi-modal large model provided by the embodiments of the present invention. For details, please refer to the previous embodiments, which will not be elaborated here.
[0114] Figure 6 The specific structural block diagram of the electronic device provided by the embodiment of the present invention is shown. The electronic device can be used to implement the method for extracting the building vector polygon contour based on the multi-modal large model provided in the above embodiment. The electronic device 500 can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a micro-processing box, or other devices, etc.
[0115] The RF circuit 510 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 510 may include various existing circuit elements for performing these functions. For example, an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, and so on. The RF circuit 510 can communicate with various networks such as the Internet, an enterprise intranet, a wireless network or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network or a metropolitan area network. The above-mentioned wireless network can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short messages, and any other suitable communication protocols, and may even include those protocols that have not been developed yet.
[0116] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, realizes functions such as taking pictures with the front camera, processing the captured images, and switching the display colors of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 520 may further include a memory remotely disposed relative to the processor 580, and these remote memories can be connected to the electronic device 500 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.
[0117] The input unit 530 can be used to receive input digital or character information, and generate a keyboard and a mouse related to user settings and function controls.
[0118] The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode).
[0119] The audio circuit 560, the speaker 561, and the microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the electrical signal converted from the received audio data to the speaker 561, and the speaker 561 converts it into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and then converted into audio data. After the audio data is output to the processor 580 for processing, it is sent to another terminal, such as, through the RF circuit 510, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may also include an earphone jack to provide communication between a peripheral earphone and the electronic device 500.
[0120] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), and it provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it does not belong to the essential components of the electronic device 500 and can be omitted completely within the scope of not changing the essence of the invention according to needs.
[0121] The processor 580 is the control center of the electronic device 500, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and by calling the data stored in the memory 520, it performs various functions of the electronic device 500 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 580 either.
[0122] The electronic device 500 also includes a power supply 590 (such as a battery) for supplying power to each component. In some embodiments, the power supply can be logically connected to the processor 580 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 590 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0123] Although not shown, the electronic device 500 also includes a camera (such as a front camera and a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations:
[0124] Obtain remote sensing images and text instructions of a building;
[0125] Input the remote sensing image and the text instruction into a trained multi-modal large model to obtain a sequence of building polygon corner coordinates, and connect the corner points to form a building vector polygon contour;
[0126] Among them, the multi-modal large model is trained through a prediction-cropping collaborative training strategy.
[0127] Specifically in implementation, the above-mentioned each module can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above-mentioned each module, reference can be made to the previous method embodiments, which will not be elaborated here.
[0128] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present invention provides a storage medium storing multiple instructions that can be loaded by a processor to execute the steps of any one of the embodiments of the method for extracting the vector polygon contour of a building based on a multimodal large model provided by the embodiments of the present invention.
[0129] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0130] Since the instructions stored in the storage medium can execute the steps in any one of the embodiments of the method for extracting the vector polygon contour of a building based on a multimodal large model provided by the embodiments of the present invention, the beneficial effects achievable by any of the methods for extracting the vector polygon contour of a building based on a multimodal large model provided by the embodiments of the present invention can be realized. For details, refer to the previous embodiments and will not be elaborated here.
[0131] The above has introduced in detail a method, device, storage medium, and electronic device for extracting the vector polygon contour of a building based on a multimodal large model provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only for helping to understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application. < / x31> < / x35> < / y112> < / x29> < / y47> < / x53> < / y67> < / x27> < / y113> < / x24> < / y112> < / x89> < / y67> < / x58> < / y165> < / x31> < / y> < / x> < / y> < / x> < / pixel>
Claims
1. A method for extracting building vector polygon outlines based on a multimodal large model, characterized in that: The method comprises: Obtain remote sensing images and text instructions of buildings; Input the remote sensing image and the text instruction into the trained multimodal large model to obtain a coordinate sequence of building polygon corner points, and connect the corner points to form a vector polygon outline of the building; The multimodal large model is trained by a prediction-cropping collaborative training strategy, and the prediction-cropping collaborative training strategy includes: predicting the training samples of the multimodal large model, obtaining the bounding box range of each building in the training samples, and cropping the training samples based on the bounding box range; wherein the multimodal large model includes a visual encoder, a text instruction encoder, a multimodal feature alignment module and a large language model, and the process of training the multimodal large model by the prediction-cropping collaborative training strategy includes: Label the center point and bounding box range of each building in the remote sensing image as the first training sample set; Inputting the images in the first training sample set into the multimodal large model, the first step is to predict the coordinates of the center point of each building on the remote sensing image, and the second step is to predict the range of the bounding box of each building based on the coordinates of the center point of the building; Based on the true value of the bounding box range of each building, the remote sensing image is cropped to obtain a local area containing only a single building, and a second training sample set is formed based on the local area containing the single building and the true value contour; Based on the true value contour coordinates corresponding to the second training sample set, the visual encoder parameters and text encoder parameters of the multimodal large model are frozen, and the remaining modules of the multimodal large model are trained until convergence.
2. The method for extracting building vector polygon outlines based on a multimodal large model according to claim 1, characterized in that: The step of inputting the remote sensing image and the text instruction into a trained multi-modal large model to obtain a coordinate sequence of building polygon corner points includes: Inputting the remote sensing image into the visual encoder to extract multi-scale features to obtain visual features; Inputting the text instruction into the text instruction encoder for deep semantic encoding to obtain text features; Inputting the visual features into the multimodal feature alignment module to map the visual features to a language space consistent with the large language model; The mapped visual features and the text features are input into the large language model to obtain a coordinate sequence of building polygon corner points.
3. The method for extracting building vector polygon outlines based on a multimodal large model according to claim 2, characterized in that: After the step of inputting the remote sensing image into the visual encoder to extract multi-scale features and obtain visual features, the method further comprises: generating location embedding information based on the remote sensing image; The position embedding information is concatenated with the visual feature to obtain an enhanced visual feature.
4. The method for extracting building vector polygon outlines based on a multimodal large model according to claim 2, characterized in that: The text instruction encoder includes a subword segmenter and a plurality of self-attention layers; The step of inputting the text instruction into the text instruction encoder for deep semantic encoding to obtain text features includes: Inputting the text instruction into the subword segmenter to obtain a token sequence; The word sequence is sequentially input into multiple self-attention layers and then transformed by self-attention to obtain text features.
5. The method for extracting building vector polygon outlines based on a multimodal large model according to claim 2, characterized in that: Mapping the visual features to a language space consistent with the large language model includes: Mapping the visual features to a language space consistent with the large language model through a projector; wherein the projector is a two-layer multi-layer perceptron; The mapping process of the projector is expressed by the following formula: in, , , , and are all projector parameters, is the activation function.
6. The method for extracting building vector polygon outlines based on a multimodal large model according to claim 2, characterized in that: The large language model includes a word segmenter, and the method further includes: The word segmenter is expanded to add pixel feature tags and position feature tags, wherein the pixel feature tags are used to indicate the pixel features of the visual features, and the position feature tags are used to mark the position features of the remote sensing image.
7. A device for extracting building vector polygon outlines based on a multimodal large model, the device for extracting building vector polygon outlines based on a multimodal large model is used to implement the method for extracting building vector polygon outlines based on a multimodal large model according to claim 1, characterized in that: include: An acquisition module is used to obtain remote sensing images and text instructions of buildings; A vector polygon outline extraction module is used to input the remote sensing image and the text instruction into the trained multi-modal large model to obtain a coordinate sequence of building polygon corner points, and connect the corner points to form a vector polygon outline of the building; Among them, the multimodal large model is trained through a prediction-cropping collaborative training strategy, and the prediction-cropping collaborative training strategy includes: predicting the training samples of the multimodal large model to obtain the bounding box range of each building in the training samples, cropping the training samples based on the bounding box range, and training the multimodal large model again with the second training sample set obtained after cropping.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the method for extracting building vector polygon outlines based on a multimodal large model as described in any one of claims 1 to 6.
9. An electronic device, characterized in that: It comprises a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the steps in the method for extracting vector polygon outlines of buildings based on a multimodal large model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Low-slow small target detection method, electronic equipment and computer program product
CN118097475A
Generative anaphora segmentation method and device based on implicit structure features
CN118570481A