Remote sensing image building outline and attribute collaborative recognition method based on candidate outline relationship diagram

CN122821131APending Publication Date: 2026-09-25XINJIANG INST OF ECOLOGY & GEOGRAPHY CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611027634.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]目前基于SAM的轮廓提取方案容易出现候选数量过多、过分割或欠分割的问题,且难以直接输出可用于业务数据库入库的结构化属性

Benefits of technology

本申请提供了一种基于候选轮廓关系图的遥感影像地物轮廓与属性协同识别方法,利用通用分割模型从预处理后的遥感影像中生成候选轮廓,并将候选轮廓矢量化为候选对象;基于候选对象之间的重叠、包含、邻接、连通和距离关系构建候选轮廓关系图;利用通用的视觉语言大模型对以候选对象为中心构建的局部图像、轮廓掩膜、上下文区域和结构化辅助特征进行属性识别,得到候选对象的类别、属性和置信度;基于候选轮廓关系图,结合边界质量、属性置信度、光谱纹理一致性、几何先验一致性和拓扑一致性计算综合一致性评分;当综合一致性评分低于预设阈值,或者候选对象的属性识别结果与其几何特征、光谱纹理特征或拓扑关系存在冲突时,根据冲突类型对候选对象执行轮廓合并、轮廓分裂、候选轮廓替换、局部重分割或重新属性识别,直至满足停止条件,得到轮廓与属性一致的矢量识别结果。本申请通过候选轮廓关系图将通用分割模型的轮廓生成能力和通用视觉语言大模型的属性识别能力耦合为一致性反馈闭环,提高复杂遥感场景下的边界质量、属性识别稳定性和矢量成果入库可用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821131A_ABST
    Figure CN122821131A_ABST
Patent Text Reader

Abstract

The application discloses a remote sensing image ground object contour and attribute collaborative identification method based on a candidate contour relationship graph, and relates to the technical fields of remote sensing image multi-modal information processing and geographic information system. A general segmentation model is used to generate a candidate contour and vectorize; a candidate contour relationship graph is constructed; a general visual language model is used for attribute identification; then, a consistency score is calculated; when the consistency score is lower than a preset threshold, or the attribute identification result of the candidate object conflicts with its geometric features, spectral texture features or topological relationship, contour merging, contour splitting, candidate contour replacement, local re-segmentation or re-attribute identification are performed on the candidate object until the stop condition is met, and a vector identification result with consistent contour and attribute is obtained. The application couples contour generation and attribute identification into a consistency feedback closed loop through the candidate contour relationship graph, improves the boundary quality, attribute identification stability and vector result storage availability in complex remote sensing scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of remote sensing image multimodal information processing and geographic information system technology, and in particular to a collaborative recognition of remote sensing image feature contours and attributes based on candidate contour relationship maps. Background Technology

[0002] Currently, remote sensing ground feature identification typically includes two technical approaches: one is the traditional remote sensing interpretation method based on pixels or objects, which uses manually designed spectral, texture, shape and contextual features, combined with a classifier to complete ground feature identification; the other is the semantic segmentation, instance segmentation or object detection method based on deep learning, which trains the network model through a large number of manually labeled samples, thereby directly outputting the ground feature category or target boundary.

[0003] With the development of segmentation-based models, Segment Anything Models (SAMs) have demonstrated strong contour generation capabilities in natural images and some remote sensing scenes. These models can provide a large number of candidate masks with relatively few prior conditions, making them suitable for boundary extraction and target candidate generation in complex remote sensing scenes. However, the results generated by SAMs primarily focus on geometric contours and cannot reliably provide land cover categories, utilization attributes, and business fields.

[0004] On the other hand, multimodal large models or visual language models have strong open-vocabulary semantic understanding capabilities, enabling them to provide category explanations, textual descriptions, and attribute judgments for image regions. For example, visual language models with region image understanding capabilities can combine region images and textual prompts to complete feature category inference, attribute discrimination, and natural language description generation.

[0005] Currently, general segmentation models, visual language models, and GIS vectorization processing are typically treated as independent steps and then linked together. Specifically, general segmentation models can generate candidate masks or contours based on remote sensing imagery, but their output mainly reflects image region boundaries and struggles to consistently provide land cover categories, utilization attributes, quality labels, and structured business fields suitable for database storage. Visual language models can perform semantic understanding and attribute judgment on candidate regions, but they typically only output category text based on local blocks or prompts, unable to reverse-correct candidate contours based on recognition results, and also struggle to handle contour fragmentation, duplicate candidates, oversegmentation, and undersegmentation. Traditional GIS vectorization processing can convert masks or contours into vector patches and perform geometric cleaning, but it relies primarily on geometric rules and spatial operations, lacking a closed-loop feedback optimization mechanism based on semantic attributes, spectral texture, and neighborhood topological relationships. Therefore, while the aforementioned linked approach can separately complete candidate contour generation, attribute recognition, and vector output, it struggles to establish a synergistic constraint between contour quality, attribute reliability, and spatial topological consistency. The main shortcomings are as follows: Currently, supervised segmentation methods rely heavily on large-scale, high-quality samples. When migrating to new regions, seasons, and land surface types, they often require re-labeling and retraining, which is costly.

[0006] Currently, contour extraction schemes based on SAM are prone to problems such as too many candidates, oversegmentation, or undersegmentation, and it is difficult to directly output structured attributes that can be used for data entry into business databases.

[0007] Current multimodal semantic recognition schemes often only make category judgments based on local blocks, without fully combining the multi-scale context, spectral statistics, shape features and neighborhood topology of remote sensing images, which leads to confusion between categories such as roads, water bodies, buildings, farmland and bare land.

[0008] Currently, the serial approach lacks a closed-loop mechanism for "attribute judgment to correct the contour in reverse". When the contour and attributes are inconsistent, there are obvious neighborhood conflicts, or the recognition confidence is low, manual verification and modification are usually still required.

[0009] Currently, the output of this technology often remains at the raster level, making it difficult to simultaneously generate vector results that integrate contour geometry, attribute fields, confidence levels, and quality markers. Summary of the Invention

[0010] The purpose of this application is to provide a method for collaborative recognition of land cover contours and attributes in remote sensing images based on candidate contour relationship maps. By coupling contour generation and attribute recognition into a consistent feedback closed loop through candidate contour relationship maps, the method improves the boundary quality, attribute recognition stability, and the availability of vector results for database entry in complex remote sensing scenarios.

[0011] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for collaborative recognition of land cover contours and attributes in remote sensing images based on candidate contour relationship maps, including: Acquire the remote sensing image to be processed, and preprocess the remote sensing image to be processed to obtain the preprocessed remote sensing image and auxiliary discrimination features associated with the image blocks of the remote sensing image; Based on a general segmentation model, multi-scale candidate contours are generated from the preprocessed remote sensing image to obtain multiple candidate masks or candidate contours. The candidate mask or candidate contour is vectorized into candidate objects, and the geometric features, spectral texture features and spatial position relationships of the candidate objects are extracted. A candidate contour relationship graph is constructed based on the overlap, inclusion, adjacency, connectivity and distance relationships between the candidate objects. An attribute recognition sample is constructed centered on the candidate object, and a general visual language model is used to identify the land cover category and attributes of the attribute recognition sample to obtain the attribute recognition result; the attribute recognition result includes category label, attribute label and confidence information; Based on the candidate contour relationship map, the geometric features, the spectral texture features, and the attribute recognition results, the boundary quality score, attribute recognition score, spectral texture consistency score, geometric prior consistency score, and topological consistency score are calculated respectively, and the consistency score of the candidate object is calculated. When the consistency score is lower than a preset threshold, or when the attribute recognition result conflicts with the geometric features, spectral texture features, or topological relationships in the candidate object's contour relationship graph, contour merging, contour splitting, candidate contour replacement, local re-segmentation, or re-attribute recognition are performed on the candidate object according to the conflict type, and the candidate contour relationship graph is updated. Feedback optimization stops when the consistency score reaches a preset threshold, the change in consistency score between two consecutive iterations is less than a preset change threshold, or the number of iterations reaches the maximum number of iterations. A vector recognition result with consistent contour and attributes is obtained, and a vector result containing contour geometry, category field, attribute field, confidence field, and quality label field is output.

[0012] Secondly, this application provides a device for collaborative recognition of remote sensing image land cover contours and attributes based on candidate contour relationship maps, comprising: The remote sensing image acquisition module is used to acquire remote sensing images to be processed. The preprocessing module is used to preprocess the remote sensing image to be processed to obtain the preprocessed remote sensing image and auxiliary discrimination features associated with image blocks of the remote sensing image. The candidate contour generation module is used to generate multi-scale candidate contours from the preprocessed remote sensing image based on a general segmentation model, so as to obtain multiple candidate masks or candidate contours. The candidate contour relationship graph construction module is used to vectorize the candidate mask or candidate contour into candidate objects, extract the geometric features, spectral texture features and spatial position relationships of the candidate objects, and construct a candidate contour relationship graph based on the overlap relationship, inclusion relationship, adjacency relationship, connectivity relationship and distance relationship between the candidate objects; The attribute recognition module is used to construct attribute recognition samples centered on the candidate objects, and to use a general visual language model to perform land cover category and attribute recognition on the attribute recognition samples to obtain attribute recognition results; the attribute recognition results include category labels, attribute labels and confidence information; The consistency evaluation module is used to calculate the boundary quality score, attribute recognition score, spectral texture consistency score, geometric prior consistency score, and topological consistency score based on the candidate contour relationship map, the geometric features, the spectral texture features, and the attribute recognition results, and to calculate the consistency score of the candidate object. The feedback optimization module is used to perform contour merging, contour splitting, candidate contour replacement, local re-segmentation or re-attribution recognition on the candidate object according to the conflict type when the consistency score is lower than a preset threshold, or when the attribute recognition result conflicts with the geometric features, spectral texture features or topological relationships of the candidate object, and to update the candidate contour relationship map. The vector result output module is used to stop feedback optimization when the consistency score reaches a preset threshold, the change in the consistency score between two consecutive iterations is less than a preset change threshold, or the number of iterations reaches the maximum number of iterations, to obtain a vector recognition result with consistent contour and attributes, and output a vector result containing contour geometry, category field, attribute field, confidence field and quality mark field.

[0013] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for collaborative recognition of remote sensing image feature contours and attributes based on candidate contour relationship maps.

[0014] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned method for collaborative recognition of remote sensing image feature contours and attributes based on candidate contour relationship maps.

[0015] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method for collaborative recognition of remote sensing image feature contours and attributes based on candidate contour relationship maps.

[0016] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a method for collaborative recognition of remote sensing image feature contours and attributes based on candidate contour relationship maps. It utilizes a general segmentation model to generate candidate contours from preprocessed remote sensing images and vectorizes these contours into candidate objects. A candidate contour relationship map is constructed based on the overlap, inclusion, adjacency, connectivity, and distance relationships between candidate objects. A general visual language model is used to perform attribute recognition on local images, contour masks, contextual regions, and structured auxiliary features centered on the candidate objects, obtaining the category, attributes, and confidence level of the candidate objects. Based on the candidate contour relationship map, a comprehensive consistency score is calculated by combining boundary quality, attribute confidence level, spectral texture consistency, geometric prior consistency, and topological consistency. When the comprehensive consistency score is lower than a preset threshold, or when the attribute recognition result of a candidate object conflicts with its geometric features, spectral texture features, or topological relationships, contour merging, contour splitting, candidate contour replacement, local re-segmentation, or re-attribute recognition are performed on the candidate objects according to the conflict type, until a stopping condition is met, resulting in a vector recognition result where the contours and attributes are consistent. This application couples the contour generation capability of the general segmentation model and the attribute recognition capability of the general visual language large model into a consistency feedback closed loop by using a candidate contour relationship graph, thereby improving the boundary quality, attribute recognition stability, and the availability of vector results for database entry in complex remote sensing scenarios. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a remote sensing image feature contour and attribute collaborative recognition method based on candidate contour relationship map; Figure 2 This is a flowchart illustrating the overall closed-loop process of a remote sensing image feature contour and attribute collaborative recognition method based on candidate contour relationship maps. Figure 3 To further refine the module structure diagram; Figure 4 This is a schematic diagram illustrating the application scenario of the remote sensing image feature contour and attribute collaborative recognition method based on candidate contour relationship map; Figure 5 This is a structural diagram of a remote sensing image feature contour and attribute collaborative recognition device based on candidate contour relationship map; Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] This application addresses the problems in the current remote sensing image feature recognition process, where candidate contour generation, attribute recognition, and GIS vectorization are independent. This leads to the inability to automatically correct contours when candidate contours and attribute recognition results are inconsistent, and the inability to effectively reduce duplicate candidates, contour fragmentation, oversegmentation, undersegmentation, and topological conflicts. Furthermore, it makes it difficult to output vector results with controllable quality, traceability, and direct database access. To address these issues, this application proposes a collaborative recognition scheme for remote sensing image feature contours and attributes based on candidate contour relationship maps.

[0021] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] In one exemplary embodiment, such as Figure 1 As shown, a method for collaborative recognition of land cover contours and attributes in remote sensing images based on candidate contour relationship maps is provided, including: Step 100: Acquire the remote sensing image to be processed and preprocess it to obtain the preprocessed remote sensing image and auxiliary discrimination features associated with image patches of the remote sensing image. Specifically, acquire the high-resolution optical remote sensing image to be processed; in other words, the remote sensing image is preferably a high-resolution satellite image.

[0023] The preprocessing of the remote sensing images to be processed includes: The remote sensing images to be processed undergo a unified coordinate transformation, and invalid values, stripe noise, and cloud shadow areas are masked and marked.

[0024] Radiometric correction, geometric correction, and orthorectification are performed on the remote sensing image to be processed to obtain the corrected remote sensing image.

[0025] Using a selected reference band or panchromatic image as a reference, band registration and spatial alignment are performed on multi-band images, and local matching correction is performed when the registration error exceeds a preset threshold.

[0026] The corrected remote sensing image or fused image is divided into multiple image blocks by sliding window segmentation.

[0027] For each image patch, auxiliary discriminant features are calculated. These features include basic spectral statistics, texture features, and exponential features. The exponential features include at least one of NDVI, NDWI, or MNDWI. The basic spectral statistics include at least the mean, standard deviation, range, and band ratio for each band. The texture features include at least the energy, contrast, entropy, homogeneity, and correlation calculated based on the gray-level co-occurrence matrix.

[0028] Figure 2 This is a flowchart illustrating the overall closed-loop process of this application, showing the processing steps from remote sensing image preprocessing, candidate contour generation, geometric feature extraction and vectorization, multimodal attribute recognition, consistency evaluation, closed-loop feedback optimization, to structured result output. Specifically, in... Figure 2 The preprocessing in the overall closed-loop process performs the following steps: Raw data reading and unification. Read multispectral band data, panchromatic band data, imaging time, solar elevation angle, solar azimuth angle, orbital parameters, RPC parameters, and coordinate reference information from remote sensing images; unify the conversion of images from different sources to the preset coordinate reference system, and mask and mark invalid values, stripe noise, and cloud shadow areas.

[0029] Radiometric correction. The original digital quantization value (DN) is converted into a radiance value or apparent reflectance based on the sensor calibration coefficient of the image to reduce the impact of differences in imaging batches, lighting conditions, and sensor response on subsequent identification. Preferably, gain and offset corrections are performed separately for each band, and abnormally bright or abnormally low bright pixels are truncated or normalized.

[0030] Geometric correction and orthorectification. Using RPC parameters, orbital attitude parameters, and digital elevation models, geometric positioning and terrain corrections are performed on remote sensing images to ensure that the corrected images correspond to their true ground locations under a unified map projection. For images with localized geometric distortions, further error correction is performed using ground control points or feature matching points.

[0031] Band registration and spatial alignment. Using a selected reference band or panchromatic image as a reference, the translation, rotation, or scale deviation between each band is estimated through feature point matching, cross-correlation matching, or phase correlation matching, and resampling registration is completed to align multi-band pixels in the same spatial location; when the registration error exceeds a preset threshold, local matching correction is re-executed.

[0032] Panchromatic and multispectral fusion. For remote sensing images that simultaneously contain panchromatic and multispectral bands, the multispectral bands are first resampled to panchromatic band resolution, and then fused using one of the following methods: Brovey transform, IHS transform, principal component transform, or Gram-Schmidt fusion method, to obtain a fused image with both high spatial resolution and spectral information. For images that do not have panchromatic bands, this processing step is skipped.

[0033] Segmentation. Based on the input size requirements of the segmentation model, the corrected or fused image is segmented using a sliding window method to generate multiple fixed-size image blocks. Preferably, the block size is set to 512×512 pixels, 1024×1024 pixels, or equivalent sizes, with 10% to 30% overlap between adjacent image blocks to avoid truncation errors when ground features are located at block boundaries. Simultaneously, the row and column numbers, geographic coordinate range, and resolution information of each image block in the original large image are recorded.

[0034] Auxiliary discriminant feature calculation. For each image patch, basic spectral statistics, texture features, and exponential features are calculated. The basic spectral statistics include at least the mean, standard deviation, range, and band ratio of each band. The texture features include at least the energy, contrast, entropy, homogeneity, and correlation calculated based on the gray-level co-occurrence matrix. The exponential features include at least the vegetation index NDVI and the water index NDWI or MNDWI. The auxiliary discriminant features are then associated and stored with the corresponding image patch as auxiliary inputs for subsequent land cover contour extraction and attribute recognition.

[0035] Step 200: Generate multi-scale candidate contours from the preprocessed remote sensing image based on a general segmentation model to obtain multiple candidate masks or candidate contours.

[0036] Among them, multi-scale candidate contour generation is performed on the preprocessed remote sensing images based on a general segmentation model, including: Multi-scale sliding window processing is performed on the preprocessed remote sensing image using at least two sets of different block sizes and corresponding overlap amounts to obtain multiple image blocks.

[0037] For each image block, the effective pixel ratio and texture intensity index are calculated. When the effective pixel ratio is lower than a preset threshold or the texture intensity index is lower than a preset threshold, the candidate contour generation for that image block is skipped.

[0038] Multiple view variants for segmentation inference are constructed from the filtered image chunks; the view variants include at least one of a visible light true color view, a contrast-enhanced view, or a near-infrared false color view.

[0039] Each view variant is input into a general segmentation model, and multiple candidate masks are generated through dense sampling of point lattices and multi-layer block inference.

[0040] Candidate masks are screened for quality, and those whose area, predicted cross-union ratio (CUI), or stability score does not meet preset quality requirements are deleted. Preset quality requirements include an area not less than a preset threshold, a predicted CUI not lower than a preset threshold, or a stability score not lower than a preset threshold.

[0041] For candidate masks within overlapping cut areas, a preservation area corresponding to the current cut is constructed, and the cut to which the candidate contour belongs is determined based on the location of the representative point of the candidate contour, so as to reduce the repeated preservation of the same feature in adjacent cuts and boundary fragmentation.

[0042] Specifically, regarding candidate contour generation, the image blocks corresponding to the preprocessed remote sensing image are input into a preset segmentation model (general segmentation model) to generate candidate masks; the preset segmentation model is preferably SAM, SAM2, or other general segmentation models capable of outputting candidate masks. To improve the applicability of the general segmentation model in remote sensing scenarios, the candidate contour generation steps include: Multi-scale sliding window processing is performed on the preprocessed remote sensing image using at least two sets of different block sizes and their corresponding overlap amounts, so as to simultaneously take into account the candidate detection of both large-scale and small-scale features; preferably, the block sizes include 1024×1024 pixels and 512×512 pixels, with corresponding overlap amounts of 256 pixels and 128 pixels, respectively.

[0043] For each image block, the effective pixel ratio and texture intensity index are calculated. When the effective pixel ratio is lower than a preset threshold or the grayscale standard deviation is lower than a preset threshold, the candidate generation process for that image block is skipped to reduce the interference of invalid and low-texture areas on the candidate results.

[0044] For the selected image blocks, multiple view variants for segmentation inference are constructed. The view variants include at least a visible light true color view; when contrast enhancement is enabled, a true color enhanced view based on CLAHE processing is also included; when the remote sensing image contains near-infrared bands, a near-infrared false color view and its contrast enhanced view are further constructed to enhance the separability of different types of land features under different spectral combinations.

[0045] Each view variant is input into the segmentation model, and multiple candidate masks are generated through dense sampling of points and multi-layer block inference. The preset segmentation model outputs the corresponding candidate region set on each view, thereby improving the recall ability of target boundaries and weakly textured features in complex remote sensing scenes.

[0046] The candidate masks output by the preset segmentation model are screened for quality. Low-quality candidate masks with an area smaller than a preset threshold, a predicted crossover ratio lower than a preset threshold, or a stability score lower than a preset threshold are deleted, while candidate masks that meet the quality requirements are retained.

[0047] For candidate results within overlapping cut areas, a retention area corresponding to the current cut is constructed, and the cut to which the candidate contour belongs is determined based on the location of the representative point of the candidate contour. When the representative point of the candidate contour falls within the retention area of ​​the current cut, the candidate contour is retained; otherwise, the candidate contour is discarded. This is to avoid the same feature being repeatedly retained in adjacent cuts and to reduce contour fragmentation caused by direct cutting boundaries.

[0048] The candidate contours retained by the current slice are merged to obtain the candidate contour set corresponding to the slice; the candidate contour results under different slices, different scales and different views are further summarized for use in subsequent geometric feature extraction, vectorization and attribute recognition steps.

[0049] Step 300: Vectorize the candidate mask or candidate contour into candidate objects, extract the geometric features, spectral texture features and spatial positional relationships of the candidate objects, and construct a candidate contour relationship graph based on the overlap, containment, adjacency, connectivity and distance relationships between the candidate objects.

[0050] The construction of the candidate contour relationship graph includes: The candidate mask is converted into a binary image, and the outer contour and inner hole contour are obtained by contour tracking.

[0051] Candidate polygons are constructed based on the outer contour and inner hole contour, and then the candidate polygons are restored from the block coordinate system to the whole scene pixel coordinate system, and then transformed to the geographic coordinate system by combining the affine transformation parameters of the remote sensing image.

[0052] Candidate polygons generated from different cuts, scales, and views are globally merged to obtain a set of candidate objects.

[0053] Extract the geometric features of each candidate object in the candidate object set; the geometric features include at least one of the following: area, perimeter, length and width of the circumscribed rectangle, aspect ratio, area of ​​the minimum circumscribed rectangle, rectangularity, compactness, orientation angle, boundary complexity, and centroid coordinates.

[0054] Using candidate objects as graph nodes, and the overlapping, containment, adjacency, connectivity, or distance relationships between candidate objects as graph edges, and taking at least one of the following as graph edge attributes: intersection area ratio, containment area ratio, boundary contact relationship, connectivity relationship, and shortest distance, a candidate contour relationship graph is obtained.

[0055] When morphological correction is used, the binary image is sequentially closed and opened using preset structuring elements to fill small holes inside the contour, remove isolated noise, and smooth local burr boundaries.

[0056] Contour extraction. A contour tracking algorithm is used to extract candidate contours from the binary image, and the outer contour and its corresponding inner hole contour are recorded. Preferably, a contour extraction method that supports hierarchical relationships is adopted. First, the outer contour is extracted as the outer boundary of the polygon, and then its sub-contours are used as the hole boundaries, thereby generating candidate geometric objects containing the outer boundary and the inner hole.

[0057] Polygon construction and geometry cleaning. Candidate polygons are constructed based on the outer contour and hole contour; for self-intersecting, fragmented, or invalid candidate geometric objects, a geometry repair operation is used for topology cleaning; for candidate results composed of multiple discrete parts, they are split into multiple independent polygons, and invalid small patches with an area smaller than a preset threshold are deleted.

[0058] Coordinate restoration and global merging. The candidate polygons in the segmented coordinate system are translated back to the whole-scene pixel coordinate system according to the row and column offset of the corresponding segment in the original image. Then, they are converted into vector contours in the geographic coordinate system by combining the affine transformation parameters of the remote sensing image. The candidate polygons generated by different segments, different scales and different views are globally merged to eliminate duplicate areas and obtain whole-scene candidate contour results.

[0059] Delayed contour simplification. To avoid premature simplification within a single slice that could lead to broken target boundaries or shape distortion, the original contour is preserved during the slice-level processing stage. Topology-preserving boundary simplification is only performed on the final contour after the global merging of candidate contours at the whole scene level is completed.

[0060] Geometric feature extraction. For each final candidate polygon, extract geometric features such as area, perimeter, length and width of the circumscribed rectangle, aspect ratio, area of ​​the minimum circumscribed rectangle, rectangularity, compactness, orientation angle, boundary complexity, and centroid coordinates. Among them, rectangularity can be obtained by the ratio of the candidate polygon's area to the area of ​​its minimum circumscribed rectangle; compactness can be obtained by the combination of the candidate polygon's area and perimeter; orientation angle can be obtained by the direction of the candidate polygon's principal axis or the direction of the long side of the minimum circumscribed rectangle; and boundary complexity can be characterized by the combination of perimeter and area or the number of boundary vertices.

[0061] Candidate Relationship Calculation. Based on the spatial positional relationships between the final candidate polygons, the overlap, containment, adjacency, connectivity, and distance relationships between candidate objects are calculated. Among them, the overlap relationship can be represented by the proportion of intersecting areas; the containment relationship can be represented by the proportion of the area of ​​one candidate object falling inside another candidate object; the adjacency relationship can be represented by boundary contact or buffer zone intersection; the connectivity relationship can be represented by whether the objects are in direct contact or connected through continuous regions; and the distance relationship can be represented by centroid distance or the shortest boundary distance.

[0062] Candidate object relationship graph construction. To ensure that the geometric relationships, attribute recognition results, and subsequent consistency evaluation of candidate contours can be expressed in a unified structure, this application further constructs a candidate object relationship graph after completing the candidate polygon construction and candidate relationship calculation. Specifically, each final candidate polygon is used as a graph node in the candidate object relationship graph, and the overlap, containment, adjacency, connectivity, and distance relationships between candidate objects are used as graph edges. The intersection area ratio, containment area ratio, boundary contact length, buffer intersection result, centroid distance, or shortest boundary distance are used as edge attributes of the corresponding graph edges.

[0063] Each graph node is associated with the node attributes of the candidate object. These node attributes include at least one of the following: spectral statistical features, texture features, exponential features, geometric features, candidate mask quality information, attribute recognition confidence, category label, attribute label, and quality marker. Specifically, spectral statistical features include the mean, standard deviation, range, or band ratio for each band; texture features include gray-level co-occurrence matrix energy, contrast, entropy, homogeneity, or correlation; exponential features include NDVI, NDWI, or MNDWI; and geometric features include area, perimeter, aspect ratio, rectangularity, compactness, orientation angle, boundary complexity, and centroid coordinates.

[0064] The candidate object relationship graph is used to support subsequent attribute identification result fusion, consistency score calculation, and feedback optimization. When calculating the topological consistency score, the system determines whether the current candidate object conforms to the spatial co-occurrence relationship and topological constraints of its corresponding land cover category based on the graph edge type, edge attributes, neighborhood category label, and neighborhood confidence score between the current candidate object and its neighboring candidate objects in the candidate object relationship graph. For example, road candidate objects should exhibit strong directional continuity and connectivity in the candidate object relationship graph; water body candidate objects should exhibit high connectivity and spectral consistency in their neighborhoods; and building candidate objects should exhibit a relatively regular shape distribution and stable adjacency relationships within their local neighborhoods.

[0065] When the attribute identification result of a candidate object is inconsistent with the neighborhood relationship in the candidate object relationship graph, the system treats this inconsistency as a topological conflict and uses it to reduce the topological consistency score or trigger feedback optimization processing. After the feedback optimization processing is completed, the system updates the graph nodes, graph edges, and corresponding attributes in the candidate object relationship graph, so that subsequent consistency scores are calculated based on the updated contour relationships and attribute information.

[0066] Step 400: Construct attribute recognition samples centered on candidate objects, and use a general visual language model to identify land cover categories and attributes in the attribute recognition samples to obtain attribute recognition results. The attribute recognition results include category labels, attribute labels, and confidence information.

[0067] In one embodiment, attribute recognition samples are constructed centered on candidate objects, and a general visual language model is used to identify the land cover category and attributes of the attribute recognition samples to obtain attribute recognition results, including: Centered on the candidate object, construct local clipping blocks, contour mask clipping blocks, and extended clipping blocks containing surrounding context.

[0068] The basic spectral statistics, texture features, exponential features, geometric features, and neighborhood relationships in the candidate contour relationship graph corresponding to the candidate objects are organized into structured auxiliary information.

[0069] By inputting local clipping blocks, contour mask clipping blocks, extended clipping blocks, and structured auxiliary information into a general visual language model, the category labels, attribute labels, semantic description text, and confidence information of candidate objects are obtained.

[0070] When multiple general visual language models or multiple recognition outputs are inconsistent, they are fused based on model confidence, geometric feature matching degree, spectral texture consistency, neighborhood topological consistency, and preset model weights.

[0071] When the fusion confidence is lower than the preset threshold, or when the category label does not match the geometric features, spectral texture features, or neighborhood topology of the candidate object, a category conflict is determined, and the process is triggered to rebuild the prompt words, switch the context range, call the verification model to re-identify, or enter the feedback optimization process.

[0072] Specifically, attribute recognition, which involves performing multi-model collaborative attribute recognition on each candidate contour, includes: Sample construction is performed. Centered on the candidate contour, local clipping blocks, contour mask clipping blocks, and extended clipping blocks containing the surrounding context are constructed respectively. The basic spectral statistics, texture features, and geometric features corresponding to the candidate contour are organized into structured auxiliary information.

[0073] Multi-model division of labor for recognition. Image samples and structured auxiliary information are respectively input into at least two multi-modal models, wherein the first multi-modal model is used to output the main category of land features corresponding to the candidate contour, and the second multi-modal model is used to output the fine-grained utilization attributes or state attributes corresponding to the candidate contour; optionally, a third multi-modal model is used as a verification model to perform semantic consistency verification on the aforementioned recognition results.

[0074] Each model outputs independently. Each multimodal model outputs its corresponding category label, attribute label, semantic description text, and confidence information; among them, the category label and attribute label are uniformly mapped to the preset land cover category system and attribute field system.

[0075] Results fusion. The outputs of multiple multimodal models are fused. When the outputs of multiple models are consistent or similar, the fused result is used as the final recognition result. When the outputs of multiple models are inconsistent, a weighted judgment is made based on the confidence level of each model, the degree of matching with the candidate contour geometric features, the degree of consistency with the spectral texture features, and the preset model weights to obtain the final category and attribute results.

[0076] Conflict verification. If the output differences of multiple multimodal models exceed a preset threshold, or the final confidence score after fusion is lower than a preset threshold, a verification mechanism is triggered. The verification mechanism includes reconstructing prompt words, switching context scope, calling the verification model to re-identify, or marking the candidate contour as an object to be checked for consistency in subsequent steps.

[0077] Output of attribute results. Output the final land cover category, utilization attribute, semantic description, model fusion confidence, and inter-model consistency label for each candidate contour, which will serve as input for subsequent consistency scoring and contour feedback optimization.

[0078] Step 500: Based on the candidate contour relationship map, geometric features, spectral texture features and attribute recognition results, calculate the boundary quality score, attribute recognition score, spectral texture consistency score, geometric prior consistency score and topological consistency score respectively, and calculate the consistency score of the candidate object.

[0079] The consistency score is determined jointly by the category normalization sub-score, graph constraint neighborhood support term, conflict penalty term, uncertainty penalty term, and adaptive weights, specifically including: Regarding the first For each candidate object, obtain the candidate category corresponding to its attribute recognition result. The original indices for boundary quality, attribute recognition, spectral texture, geometric priors, and topological relationships were calculated respectively.

[0080] Based on candidate category For the corresponding category baseline interval, the original boundary quality index, original attribute recognition index, original spectral texture index, original geometric prior index, and original topological relation index are normalized respectively to obtain the boundary quality score. Attribute recognition score Spectral texture consistency score Geometric prior consistency score and topology consistency score ,in , , , and All values ​​are located in the range [0,1], and the larger the value, the higher the consistency.

[0081] Based on the candidate category The adaptive weight vector is generated by considering factors such as area scale, attribute recognition confidence, node degree in the candidate contour relationship graph, neighborhood category complexity, and the number of historical feedback corrections. ,in , , , and Corresponding to , , , and The weighting coefficients are all greater than or equal to 0, and the sum of the weights is 1.

[0082] Based on the candidate contour relationship graph, the first The edge weights between each candidate object and its neighboring candidate objects, the consistency scores of the neighboring candidate objects, and the co-occurrence relationships in the category space are used to compute graph constraints on neighborhood support terms. The graph-constrained neighborhood support term Used to characterize the Whether a candidate object is supported by both the spatial relationship and attribute category of its neighboring candidate objects.

[0083] According to the The degree of mismatch between the attribute recognition results of each candidate object and its geometric features, spectral texture features, and topological relationships is used to calculate a conflict penalty term. The conflict penalty item It includes at least one of geometric conflict penalty, spectral texture conflict penalty, and topological conflict penalty.

[0084] An uncertainty penalty term is calculated based on the output confidence of the general visual language model, the degree of discrepancy between multiple models or multiple recognition results, and the fluctuation of the candidate object boundary quality. .

[0085] Calculate the first The candidate in the first Consistency scoring in round feedback optimization : .

[0086] .

[0087] in, For the first Ontology consistency score of each candidate object; These are graph-constrained neighborhood support terms calculated based on the candidate contour relationship graph from the previous round; The neighborhood support coefficient; This is the conflict penalty coefficient; This represents the uncertainty penalty coefficient. , and All are located in the interval [0,1].

[0088] Step 600: When the consistency score is lower than the preset threshold, or when the attribute recognition result conflicts with the geometric features, spectral texture features or topological relationships in the candidate object's contour relationship graph, perform contour merging, contour splitting, candidate contour replacement, local re-segmentation or re-attribute recognition on the candidate object according to the conflict type, and update the candidate contour relationship graph.

[0089] Step 700: Stop feedback optimization when the consistency score reaches the preset threshold, the change in consistency score between two consecutive iterations is less than the preset change threshold, or the number of iterations reaches the maximum number of iterations. Obtain vector recognition results with consistent contours and attributes, and output vector results containing contour geometry, category field, attribute field, confidence field, and quality label field.

[0090] Specifically, when the consistency score Lower than candidate categories Corresponding dynamic threshold Or conflict penalty items When the conflict level exceeds the preset threshold, the first... One candidate object was identified as the object to be optimized. Dynamic threshold. Based on candidate category The candidate object area scale, neighborhood topological complexity, business quality requirements, and number of historical feedback corrections are determined.

[0091] When the If the change in the consistency score of a candidate object is less than a preset change threshold for two consecutive rounds, or the consistency score... Reaching dynamic threshold Or, when the number of feedback optimizations reaches the maximum number of iterations, stop optimizing the 1st iteration. Feedback optimization for each candidate.

[0092] In an exemplary embodiment, the candidate contour relationship graph is used to describe the spatial constraint relationships between candidate objects. Specifically, each candidate object is treated as a graph node, and the overlap, containment, adjacency, connectivity, and distance relationships between candidate objects are treated as graph edges. The intersection area ratio, containment area ratio, boundary contact length, buffer intersection result, centroid distance, or shortest boundary distance are treated as graph edge attributes. Each graph node is also associated with the candidate object's geometric features, spectral texture features, attribute recognition results, confidence information, and quality markers.

[0093] During the consistency evaluation process, boundary quality score, attribute recognition score, spectral texture consistency score, geometric prior consistency score, and topological consistency score are calculated for each candidate object. Specifically, the boundary quality score characterizes whether the candidate contour boundary is complete, smooth, and stable; the attribute recognition score characterizes the credibility of the output results of the general visual language model; the spectral texture consistency score characterizes whether the spectral statistics, exponential features, and texture features of the candidate object conform to its attribute category; the geometric prior consistency score characterizes whether the area, shape, orientation, and boundary complexity of the candidate object conform to the geometric prior of the corresponding land cover category; and the topological consistency score characterizes whether the spatial relationship between the candidate object and its neighboring objects conforms to the spatial distribution patterns of land cover.

[0094] Feedback optimization is triggered when the overall consistency score (consistency score) falls below a preset threshold, or when the attribute recognition results of a candidate object conflict with its geometric features, spectral texture features, or neighborhood topological relationships. Feedback optimization includes: contour merging for fragmented candidate objects identified as belonging to the same class and being adjacent or connected; contour splitting for candidate objects with significant internal spectral texture differences or inconsistent attribute recognition results; candidate contour replacement for overlapping and competing candidate objects with low confidence; local resegmentation for candidate objects with low boundary quality and significant attribute conflicts; and re-attribute recognition for candidate objects with low attribute confidence or inconsistent model outputs. After each feedback optimization, the geometric features, spectral texture features, and candidate contour relationship graph of the candidate object are recalculated, and the consistency score is recalculated.

[0095] The stopping conditions for feedback optimization include: the consistency score of the candidate object reaches a preset threshold; the change in the consistency score between two consecutive iterations is less than a preset change threshold; or the number of feedback optimization iterations reaches the maximum number of iterations. After stopping, the system outputs vector results containing contour geometry, category field, attribute field, confidence field, and quality label field. The quality label field is used to identify whether the candidate object has been corrected by feedback and whether there are still low confidence or topological conflicts.

[0096] In one exemplary embodiment, the consistency scoring does not employ a simple linear summation with fixed weights, but rather a graph-constrained adaptive scoring mechanism. The system first establishes category baseline intervals based on the attribute categories of candidate objects. For example, road objects prioritize directional continuity and topological connectivity, water objects prioritize water index and spectral texture consistency, building objects prioritize rectangularity, compactness, and directional consistency, and farmland objects prioritize texture homogeneity and neighborhood category continuity. Based on these category baseline intervals, the system uniformly normalizes boundary quality, attribute recognition, spectral texture, geometric priors, and topological relationship indices of different dimensions to the [0,1] interval.

[0097] Furthermore, the system dynamically determines the weights of the five sub-scores based on the candidate object's category, area scale, recognition confidence, node degree in the candidate contour relationship graph, neighborhood category complexity, and number of historical feedback corrections. This allows different evaluation priorities to be applied to candidate objects of different categories and in different spatial environments. For example, for road candidate objects, the weights of topological consistency and boundary continuity are increased; for water body candidate objects, the weights of spectral texture consistency are increased; and for building candidate objects, the weights of geometric prior consistency are increased.

[0098] Meanwhile, the system introduces a graph-constrained neighborhood support term to evaluate whether candidate objects are supported by both neighborhood spatial relationships and neighborhood attribute categories. If the connectivity, spatial co-occurrence, and attribute category combinations between a candidate object and its neighboring objects conform to a preset spatial distribution pattern of land features, the consistency score of the candidate object is increased; if there are abnormal overlaps, abnormal inclusions, mutually exclusive categories, or broken connectivity relationships between a candidate object and its neighboring objects, the consistency score of the candidate object is reduced through a conflict penalty term.

[0099] Furthermore, the system introduces an uncertainty penalty term to handle situations where the confidence level of the general visual language model output is low, the results of multiple models show significant discrepancies, or the quality of candidate boundaries fluctuates significantly. When the uncertainty penalty term is high, even if some sub-scores of a candidate object are high, its overall consistency score will be reduced, triggering a review or feedback optimization. Through the above mechanism, this application can incorporate the candidate object's own characteristics, attribute recognition results, neighborhood topology relationships, and feedback history into the consistency evaluation, thereby improving the stability and interpretability of closed-loop optimization.

[0100] Specifically, the consistency assessment involves calculating a consistency score, including: Boundary quality score calculation. The boundary quality score is calculated based on the predicted cross-union ratio, stability score, boundary smoothness, and local boundary integrity of the candidate mask corresponding to the candidate contour. .

[0101] Attribute recognition score calculation. The attribute recognition score is calculated based on the fusion confidence of multiple multimodal model outputs, inter-model consistency, and the degree of matching between semantic descriptions and candidate contour auxiliary features. .

[0102] Spectral texture consistency score calculation. The spectral texture consistency score is calculated based on the degree of matching between the spectral statistics, exponential features, and texture features of the region corresponding to the candidate contour and its recognized category. .

[0103] Geometric prior consistency score calculation. The geometric prior consistency score is calculated based on the degree of geometric prior matching between the candidate contour's area, perimeter, aspect ratio, rectangularity, compactness, orientation, and boundary complexity and its recognized category. .

[0104] Topological consistency score calculation. Based on the overlap, containment, adjacency, connectivity, and distance relationships between the candidate contour and its neighboring objects, and combined with one or more rules from road continuity, water connectivity, building regularity, farmland texture homogeneity, and land type adjacency rationality, the topological consistency score is calculated. .

[0105] Consistency score fusion. Incorporating boundary quality scores. Attribute recognition score Spectral texture consistency score Geometric prior consistency score and topology consistency score After normalization, the candidate contours are weighted and fused according to preset or adaptive weights to obtain a consistency score. When consistency score If the candidate contour is below a preset threshold, it is marked as an object to be optimized.

[0106] The weighting coefficients can be determined manually or adaptively adjusted based on the statistical results of the validation samples, the candidate object category, or the historical recognition results.

[0107] Closed-loop feedback optimization. When the consistency score is lower than the set threshold, mutual exclusion category conflicts occur, or multiple candidate contours overlap and compete, the closed-loop feedback process is triggered. Based on the attribute recognition results, uncertainty, topological relationship, and candidate quality, the current contour is re-merged, split, replaced, locally re-segmented, candidate optimization, or re-attribute discrimination is performed until the stopping condition is met or the maximum number of iterations is reached.

[0108] Conflict Types and Feedback Optimization Rules. To ensure that closed-loop feedback optimization has clear triggering conditions and processing paths, this application pre-defines multiple conflict types and corresponding feedback optimization rules based on the attribute categories, geometric features, spectral texture features, and neighborhood topological relationships in the candidate object relationship graph. The conflict types include at least road fracture conflicts, water body spectral conflicts, building geometric conflicts, farmland texture conflicts, candidate overlap competition conflicts, and low-confidence attribute conflicts.

[0109] For road-type candidate objects, when a candidate object is identified as a road-type feature and its aspect ratio, orientation angle, or neighborhood orientation continuity meets the road prior conditions, but appears as multiple adjacent or near-neighbor short fragments in the candidate object relationship graph, and the shortest distance between adjacent fragments is less than a preset distance threshold, the orientation angle is less than a preset angle threshold, or the connectivity score is higher than a preset connectivity threshold, it is determined that there is a road breakage conflict, and contour merging, local reconnection, or local resegmentation processing is performed on the corresponding candidate object.

[0110] For water-type candidate objects, when a candidate object is identified as a water-type land cover, but its NDWI or MNDWI is lower than the preset water index threshold, or the number of boundary fragments of the candidate object is higher than the preset fragment threshold, or the number of isolated small patches is higher than the preset isolation threshold, it is determined that there is a water spectral conflict, and the candidate object is subjected to re-attribute identification, candidate contour replacement, local re-segmentation, or low-quality candidate removal processing.

[0111] For building-type candidate objects, when a candidate object is identified as a building-type feature, but its rectangularity, compactness, or orientation consistency is lower than the corresponding building prior threshold, or when there is abnormal inclusion, abnormal overlap, or unstable adjacency relationship with surrounding building candidate objects in the candidate object relationship graph, it is determined that there is a building geometric conflict, and local re-segmentation, boundary correction, candidate outline replacement, or re-attribute recognition processing is performed on the candidate object.

[0112] For candidate objects of the cultivated land category, when a candidate object is identified as a cultivated land feature, but its texture homogeneity is lower than the preset texture threshold, or the category label of its neighboring candidate objects is inconsistent with the prior continuous distribution of cultivated land, or the boundary fragmentation is higher than the preset fragmentation threshold, it is determined that there is a cultivated land texture conflict, and contour splitting, candidate contour replacement, local re-segmentation, or re-attribute recognition processing is performed on the candidate object.

[0113] For candidate overlap competition conflicts, when the proportion of the intersection area between two or more candidate objects is higher than the preset overlap threshold, and the attribute categories corresponding to the candidate objects are mutually exclusive or the confidence difference exceeds the preset confidence difference threshold, the candidate object with the higher consistency score is retained from multiple candidate objects according to the consistency score, attribute recognition confidence and candidate mask quality, and the remaining candidate objects are subjected to candidate contour replacement, merging, deletion or re-identification processing.

[0114] For low-confidence attribute conflicts, when the confidence of the attribute recognition of the candidate object is lower than the preset confidence threshold, or when the degree of category divergence between multiple general visual language models or multiple recognition results is higher than the preset divergence threshold, it is determined that there is a low-confidence attribute conflict, and attribute recognition is re-performed by reconstructing prompt words, adjusting the scope of contextual blocks, calling the verification model, or introducing structured auxiliary features.

[0115] After each feedback optimization process, the system recalculates the geometric features, spectral texture features, attribute recognition results, and candidate object relationship graph of the corrected candidate object and its neighboring candidate objects, and recalculates the consistency score. If the corrected consistency score reaches the corresponding dynamic threshold, or the change in consistency score between two consecutive feedback optimizations is less than the preset change threshold, or the number of feedback optimizations reaches the maximum number of iterations, then the feedback optimization of that candidate object is stopped, and the corresponding quality mark is recorded in the vector results.

[0116] Output results. The output includes structured geographic information results containing vector contours, feature categories, attribute fields, confidence levels, quality markers, and time information. The result format can be Shapefile, GeoJSON, GPKG, database records, or other formats that can be directly called by GIS systems.

[0117] If the attribute is identified as a road feature but the candidate outline has obvious breaks, too many short fragments, or insufficient continuity with the direction of adjacent roads, the priority of reconnection or merging is increased; if the attribute is identified as a water feature but the water index of the corresponding area is low or there are a large number of isolated fragments, the candidate outline is re-screened or re-discriminated.

[0118] The above method is applicable not only to single-phase images but also to multi-temporal images. It can record change markers based on attribute changes in different time phases for dynamic resource monitoring.

[0119] This application differs from current minimal technical combinations of remote sensing image segmentation, attribute recognition, and vectorization concatenation processing schemes in that: after the general segmentation model generates candidate contours, the candidate contours are not directly fed into the attribute recognition model and the results are output. Instead, the candidate contours are first vectorized into candidate objects, and a candidate contour relationship map is constructed based on the overlap, inclusion, adjacency, connectivity, and distance relationships between candidate objects. Then, a general visual language model is used to identify the categories and attributes of the candidate objects. Subsequently, the candidate contour relationship map, the geometric features, spectral texture features, and attribute recognition results of the candidate objects are combined to calculate five sub-scores: boundary quality, attribute recognition, spectral texture consistency, geometric prior consistency, and topological consistency. When the overall consistency score is lower than a preset threshold, or when there is a conflict between the category label and the contour shape, spectral texture, or neighborhood topological relationship of the candidate object, contour merging, contour splitting, candidate contour replacement, local re-segmentation, or re-attribute recognition are triggered according to the conflict type. The candidate contour relationship map is updated during the feedback process until the stopping condition is met, and then a vector result with a quality label is output. Thus, this application forms a closed-loop technical path of "construction of candidate contour relationship graph - attribute recognition - five types of consistency scoring - conflict triggering feedback correction - quality mark vector output".

[0120] Compared with the prior art, this application has at least the following beneficial effects: First, this application elevates contour generation and attribute recognition from a simple sequential process to a collaborative closed-loop process, which can use attribute results to correct contours in reverse, thereby improving boundary quality and semantic stability in complex scenes.

[0121] Second, this application fully integrates the spectrum, texture, shape, and topological relationships of high-resolution remote sensing images, which not only identifies land cover categories but also outputs structured attribute fields that are more suitable for business applications, thereby improving the usability of the results.

[0122] Third, this application can significantly reduce the workload of manual verification and trimming, and is particularly suitable for high-frequency batch tasks such as large-scale natural resource surveys, farmland supervision, and urban feature renewal.

[0123] Fourth, this application outputs vector results that integrate contours and attributes, making it easy to directly access geographic information databases, thematic mapping systems, and statistical analysis processes.

[0124] Fifth, this application adopts a combination of segmentation base model and multimodal large model, which takes into account both geometric boundary acquisition capability and open semantic understanding capability, and has strong generalization potential.

[0125] Example of ground feature identification based on Gaofen-1 satellite imagery: High-resolution optical images acquired by the Gaofen-1 satellite in the study area were selected as input data. Radiometric correction, orthorectification, and block processing were performed on the original images. The block size can be set to 1024 pixels × 1024 pixels, and the overlapping area of ​​adjacent blocks is preserved.

[0126] The blocks are input into the SAM model to generate initial candidate masks. Masks that are too small, have abnormal boundaries, or are obviously repetitive are filtered out, and the remaining masks are converted into vector contours.

[0127] For each vector contour, crop the image of the region inside the contour and the image of the region outside the contour with context, and at the same time calculate its mean spectral value, standard deviation, texture entropy, aspect ratio, rectangularity and adjacency.

[0128] The images and structured features are fed into a visual language model with regional image understanding capabilities, which outputs land cover categories and attribute results, such as buildings, roads, water bodies, forest land, cultivated land, bare land, and their corresponding use status or surface attributes.

[0129] When the attributes output by a visual language model with regional image understanding capabilities conflict with the prior geometry of the contour, such as being identified as a road but having an approximately closed surface contour, or being identified as a body of water but having a significantly low water index in the corresponding region, feedback optimization is triggered to reselect candidate contours, adjust the local resegmentation range, or re-identify the attributes of the contour.

[0130] The final output is a vector result containing contour geometry, category field, attribute field, confidence field, and quality marker field, which is used for subsequent survey mapping and database updates.

[0131] For example, when a candidate object is identified as a road-type feature, but the candidate outline appears as multiple short fragments, and the adjacent fragments in the candidate outline relationship diagram have directional continuity, a boundary distance less than a preset distance threshold, and similar spectral texture features, the multiple short fragments are marked as road breakage conflicts, and outline merging or local reconnection processing is performed.

[0132] When a candidate object is identified as a water body land feature, but its NDWI or MNDWI is lower than the preset water body index threshold, and the candidate outline shows a large number of isolated fragments or abnormal overlap with surrounding non-water body objects, it is marked as a water body spectral conflict, and triggers re-attribute recognition, candidate outline replacement or local re-segmentation processing.

[0133] When a candidate object is identified as a building feature, but its rectangularity, compactness, or orientation consistency is lower than the corresponding building prior threshold, and it has an abnormal inclusion or overlap relationship with adjacent building candidate objects, it is marked as a building geometric conflict, and candidate outline splitting, boundary simplification, or local re-segmentation processing is performed.

[0134] When a candidate object is identified as a farmland feature, but its texture homogeneity is lower than a preset threshold, or when the attribute differences between adjacent objects in the candidate contour relationship map are obvious and the boundaries are broken, it is marked as a farmland texture conflict, and contour splitting, candidate contour replacement, or re-attribute recognition processing is performed.

[0135] After the above conflict resolution is completed, the candidate contour relationship graph is reconstructed and the consistency score is updated. If the updated consistency score reaches the preset threshold, the corrected candidate object is retained; otherwise, the next round of feedback optimization is performed until the stopping condition is met.

[0136] The basic segmentation model is not limited to SAM, and can be replaced by other basic segmentation models that can output candidate masks; without changing the core idea of ​​this application, the replaced model still falls within the protection scope of this application.

[0137] The multimodal large model is not limited to visual language models with region image understanding capabilities; it can also be replaced with other visual language models with region image understanding and attribute discrimination capabilities. When using multiple visual language models or performing multiple recognitions, the system fuses different output results. When multiple output results are consistent or similar, the consistent result is used as the attribute recognition result; when multiple output results are inconsistent, fusion is determined based on the confidence level of each output result, the geometric feature matching degree of the candidate object, the spectral texture consistency degree, the neighborhood topological consistency degree, and the preset model weights. If the confidence level of the fused result is lower than the preset confidence threshold, or the class divergence degree is higher than the preset divergence threshold, the candidate object is marked as a low-confidence object, and the system triggers the reconstruction of prompt words, adjustment of the context range, or re-identification using a verification model.

[0138] Candidate consistency evaluation can be performed using rule-weighted methods, or by using learnable ranking models, graph optimization models, or energy function models.

[0139] This application relates to technologies such as intelligent interpretation of remote sensing images, computer vision, multimodal information processing, and geographic information systems. This application can be applied to scenarios such as land cover surveys, natural resource surveys and monitoring, farmland and forestry and grassland supervision, water and wetland monitoring, urban feature renewal, disaster assessment, and thematic mapping based on high-resolution optical satellite images such as Gaofen-1, aerial images, and other similar remote sensing data. Figure 4 This is a typical scene illustration, showing the comparison of the contours and attribute correction effects of features such as roads, water bodies, buildings, and farmland before and after the loop closure.

[0140] In different business scenarios, attribute fields can be expanded as needed to include land category codes, utilization methods, target status, change types, suspected error markers, etc.

[0141] Figure 3 A module structure diagram for further refining the above method may include a data preprocessing module, a candidate contour generation module, a candidate relationship graph construction module, an attribute recognition module, a consistency evaluation module, a feedback optimization module, and a result output module.

[0142] In one exemplary embodiment, such as Figure 5 As shown, a device for collaborative recognition of land cover contours and attributes in remote sensing images based on candidate contour relationship maps is provided, comprising: The remote sensing image acquisition module is used to acquire remote sensing images to be processed.

[0143] The preprocessing module is used to preprocess the remote sensing image to be processed, and obtain the preprocessed remote sensing image and auxiliary discrimination features associated with the image blocks of the remote sensing image.

[0144] The candidate contour generation module is used to generate multi-scale candidate contours from the preprocessed remote sensing image based on a general segmentation model, thereby obtaining multiple candidate masks or candidate contours.

[0145] The candidate contour relationship graph construction module is used to vectorize the candidate mask or candidate contour into candidate objects, extract the geometric features, spectral texture features and spatial position relationships of the candidate objects, and construct a candidate contour relationship graph based on the overlap relationship, inclusion relationship, adjacency relationship, connectivity relationship and distance relationship between the candidate objects.

[0146] The attribute recognition module is used to construct attribute recognition samples centered on the candidate objects, and to use a general visual language model to perform land cover category and attribute recognition on the attribute recognition samples to obtain attribute recognition results; the attribute recognition results include category labels, attribute labels and confidence information.

[0147] The consistency evaluation module is used to calculate the boundary quality score, attribute recognition score, spectral texture consistency score, geometric prior consistency score, and topological consistency score based on the candidate contour relationship graph, the geometric features, the spectral texture features, and the attribute recognition results, and to calculate the consistency score of the candidate object.

[0148] The feedback optimization module is used to perform contour merging, contour splitting, candidate contour replacement, local re-segmentation, or re-attribution recognition on the candidate object according to the conflict type when the consistency score is lower than a preset threshold, or when the attribute recognition result conflicts with the geometric features, spectral texture features, or topological relationships of the candidate object, and to update the candidate contour relationship map.

[0149] The vector result output module is used to stop feedback optimization when the consistency score reaches a preset threshold, the change in the consistency score between two consecutive iterations is less than a preset change threshold, or the number of iterations reaches the maximum number of iterations, to obtain a vector recognition result with consistent contour and attributes, and output a vector result containing contour geometry, category field, attribute field, confidence field and quality mark field.

[0150] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 6As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores remote sensing image feature contour extraction and attribute recognition data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the remote sensing image feature contour extraction and attribute recognition method.

[0151] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0152] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0153] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0154] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0155] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0156] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0157] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logic devices, etc., and are not limited to these.

[0158] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0159] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for collaborative recognition of land cover contours and attributes in remote sensing images based on candidate contour relationship maps, characterized in that, include: Acquire the remote sensing image to be processed, and preprocess the remote sensing image to be processed to obtain the preprocessed remote sensing image and auxiliary discrimination features associated with the image blocks of the remote sensing image; Based on a general segmentation model, multi-scale candidate contours are generated from the preprocessed remote sensing image to obtain multiple candidate masks or candidate contours. The candidate mask or candidate contour is vectorized into candidate objects, and the geometric features, spectral texture features and spatial position relationships of the candidate objects are extracted. A candidate contour relationship graph is constructed based on the overlap, inclusion, adjacency, connectivity and distance relationships between the candidate objects. An attribute recognition sample is constructed centered on the candidate object, and a general visual language model is used to identify the land cover category and attributes of the attribute recognition sample to obtain the attribute recognition result; the attribute recognition result includes category label, attribute label and confidence information; Based on the candidate contour relationship map, the geometric features, the spectral texture features, and the attribute recognition results, the boundary quality score, attribute recognition score, spectral texture consistency score, geometric prior consistency score, and topological consistency score are calculated respectively, and the consistency score of the candidate object is calculated. When the consistency score is lower than a preset threshold, or when the attribute recognition result conflicts with the geometric features, spectral texture features, or topological relationships in the candidate object's contour relationship graph, contour merging, contour splitting, candidate contour replacement, local re-segmentation, or re-attribute recognition are performed on the candidate object according to the conflict type, and the candidate contour relationship graph is updated. Feedback optimization stops when the consistency score reaches a preset threshold, the change in consistency score between two consecutive iterations is less than a preset change threshold, or the number of iterations reaches the maximum number of iterations. A vector recognition result with consistent contour and attributes is obtained, and a vector result containing contour geometry, category field, attribute field, confidence field, and quality label field is output.

2. The method for collaborative recognition of remote sensing image feature contours and attributes based on candidate contour relationship maps according to claim 1, characterized in that, Preprocessing the remote sensing image to be processed includes: The remote sensing image to be processed is subjected to a unified coordinate transformation, and invalid values, strip noise, and cloud shadow areas are masked and marked. The remote sensing image to be processed is subjected to radiometric correction, geometric correction and orthorectification to obtain the corrected remote sensing image; Using a selected reference band or panchromatic image as a reference, band registration and spatial alignment are performed on multi-band images, and local matching correction is performed when the registration error exceeds a preset threshold. The corrected remote sensing image or fused image is divided into multiple image blocks by a sliding window. For each image block, auxiliary discrimination features are calculated; the auxiliary discrimination features include basic spectral statistics, texture features, and exponential features; the exponential features include at least one of NDVI, NDWI, or MNDWI.

3. The method for collaborative recognition of remote sensing image land cover contours and attributes based on candidate contour relationship maps according to claim 1, characterized in that, Multi-scale candidate contour generation is performed on the preprocessed remote sensing image based on a general segmentation model, including: Multi-scale sliding window processing is performed on the preprocessed remote sensing image using at least two sets of different block sizes and corresponding overlap amounts to obtain multiple image blocks. For each image block, the effective pixel ratio and texture intensity index are calculated. When the effective pixel ratio is lower than a preset threshold or the texture intensity index is lower than a preset threshold, the candidate contour generation of that image block is skipped. Multiple view variants for segmentation inference are constructed from the filtered image chunks; the view variants include at least one of a visible light true color view, a contrast-enhanced view, or a near-infrared false color view; Each view variant is input into the general segmentation model, and multiple candidate masks are generated through dense sampling lattice and multi-layer block inference. The candidate masks are screened for quality, and candidate masks whose area, predicted cross-union ratio, or stability score do not meet the preset quality requirements are deleted. For candidate masks within overlapping cut areas, a preservation area corresponding to the current cut is constructed, and the cut to which the candidate contour belongs is determined based on the location of the representative point of the candidate contour, so as to reduce the repeated preservation of the same feature in adjacent cuts and boundary fragmentation.

4. The method for collaborative recognition of remote sensing image land cover contours and attributes based on candidate contour relationship maps according to claim 1, characterized in that, Constructing the candidate contour relationship graph includes: The candidate mask is converted into a binary image, and the outer contour and inner hole contour are obtained by contour tracking. Candidate polygons are constructed based on the outer contour and inner hole contour, and the candidate polygons are restored from the block coordinate system to the whole scene pixel coordinate system, and then transformed to the geographic coordinate system by combining the affine transformation parameters of the remote sensing image. The candidate polygons generated from different cuts, scales, and views are globally merged to obtain a set of candidate objects. Extract the geometric features of each candidate object in the candidate object set; the geometric features include at least one of area, perimeter, length and width of the circumscribed rectangle, aspect ratio, area of ​​the minimum circumscribed rectangle, rectangularity, compactness, orientation angle, boundary complexity, and centroid coordinates; Using the candidate objects as graph nodes, and the overlapping, inclusion, adjacency, connectivity, or distance relationships between the candidate objects as graph edges, and taking at least one of the following as graph edge attributes: intersection area ratio, inclusion area ratio, boundary contact relationship, connectivity relationship, and shortest distance, a candidate contour relationship graph is obtained.

5. The method for collaborative recognition of remote sensing image feature contours and attributes based on candidate contour relationship maps according to claim 1, characterized in that, An attribute recognition sample is constructed centered on the candidate object, and a general visual language model is used to identify the land cover category and attributes of the attribute recognition sample to obtain the attribute recognition result, including: Centered on the candidate object, construct local clipping blocks, contour mask clipping blocks, and extended clipping blocks containing surrounding context, respectively; The basic spectral statistics, texture features, exponential features, geometric features, and neighborhood relationships in the candidate contour relationship graph corresponding to the candidate objects are organized into structured auxiliary information. The local clipping block, the contour mask clipping block, the extended clipping block, and the structured auxiliary information are input into a general visual language model to obtain the category label, attribute label, semantic description text, and confidence information of the candidate object. When multiple general visual language models or multiple recognition outputs are inconsistent, they are fused based on model confidence, geometric feature matching degree, spectral texture consistency, neighborhood topological consistency and preset model weights; When the fusion confidence is lower than the preset threshold, or when the category label does not match the geometric features, spectral texture features, or neighborhood topology of the candidate object, a category conflict is determined, and the process is triggered to rebuild the prompt words, switch the context range, call the verification model to re-identify, or enter the feedback optimization process.

6. The method for collaborative recognition of remote sensing image land cover contours and attributes based on candidate contour relationship maps according to claim 1, characterized in that, The consistency score is determined jointly by the category normalization sub-score, graph constraint neighborhood support term, conflict penalty term, uncertainty penalty term, and adaptive weights, specifically including: Regarding the first For each candidate object, obtain the candidate category corresponding to its attribute recognition result. And calculate the original indexes of boundary quality, attribute recognition, spectral texture, geometric prior and topological relations respectively; Based on candidate category For the corresponding category benchmark interval, the original boundary quality index, original attribute recognition index, original spectral texture index, original geometric prior index, and original topological relation index are normalized respectively to obtain the boundary quality score. Attribute recognition score Spectral texture consistency score Geometric prior consistency score and topology consistency score ,in , , , and All values ​​are located in the range [0,1], and the larger the value, the higher the consistency. Based on the candidate category The adaptive weight vector is generated by considering factors such as area scale, attribute recognition confidence, node degree in the candidate contour relationship graph, neighborhood category complexity, and the number of historical feedback corrections. ,in , , , and Corresponding to , , , and The weighting coefficients; each weighting coefficient is greater than or equal to 0, and the sum of the weights is 1; Based on the candidate contour relationship graph, the first The edge weights between each candidate object and its neighboring candidate objects, the consistency scores of the neighboring candidate objects, and the co-occurrence relationships in the category space are used to compute graph constraints on neighborhood support terms. The graph-constrained neighborhood support term Used to characterize the Whether a candidate object is supported by both the spatial relationship and attribute category of its neighboring candidate objects; According to the The degree of mismatch between the attribute recognition results of each candidate object and its geometric features, spectral texture features, and topological relationships is used to calculate a conflict penalty term. The conflict penalty item Including at least one of geometric conflict penalty, spectral texture conflict penalty, and topological conflict penalty; An uncertainty penalty term is calculated based on the output confidence of the general visual language model, the degree of discrepancy between multiple models or multiple recognition results, and the fluctuation of the candidate object boundary quality. ; Calculate the first The candidate in the first Consistency scoring in round feedback optimization : ; ; in, For the first Ontology consistency score of each candidate object; These are graph-constrained neighborhood support terms calculated based on the candidate contour relationship graph from the previous round; The neighborhood support coefficient; This is the conflict penalty coefficient; This represents the uncertainty penalty coefficient. , and All are located in the interval [0,1]; When consistency score Lower than candidate categories Corresponding dynamic threshold Or conflict penalty items When the conflict level exceeds the preset threshold, the first... One candidate object was identified as the object to be optimized; The dynamic threshold Based on candidate category The candidate object area scale, neighborhood topological complexity, business quality requirements, and number of historical feedback corrections are determined. When the If the change in the consistency score of a candidate object is less than a preset change threshold for two consecutive rounds, or the consistency score... Reaching dynamic threshold Or, when the number of feedback optimizations reaches the maximum number of iterations, stop optimizing the 1st iteration. Feedback optimization for each candidate.

7. A device for collaborative recognition of land feature contours and attributes in remote sensing images based on candidate contour relationship maps, characterized in that, include: The remote sensing image acquisition module is used to acquire remote sensing images to be processed. The preprocessing module is used to preprocess the remote sensing image to be processed to obtain the preprocessed remote sensing image and auxiliary discrimination features associated with image blocks of the remote sensing image. The candidate contour generation module is used to generate multi-scale candidate contours from the preprocessed remote sensing image based on a general segmentation model, so as to obtain multiple candidate masks or candidate contours. The candidate contour relationship graph construction module is used to vectorize the candidate mask or candidate contour into candidate objects, extract the geometric features, spectral texture features and spatial position relationships of the candidate objects, and construct a candidate contour relationship graph based on the overlap relationship, inclusion relationship, adjacency relationship, connectivity relationship and distance relationship between the candidate objects; The attribute recognition module is used to construct attribute recognition samples centered on the candidate objects, and to use a general visual language model to perform land cover category and attribute recognition on the attribute recognition samples to obtain attribute recognition results; the attribute recognition results include category labels, attribute labels and confidence information; The consistency evaluation module is used to calculate the boundary quality score, attribute recognition score, spectral texture consistency score, geometric prior consistency score, and topological consistency score based on the candidate contour relationship map, the geometric features, the spectral texture features, and the attribute recognition results, and to calculate the consistency score of the candidate object. The feedback optimization module is used to perform contour merging, contour splitting, candidate contour replacement, local re-segmentation or re-attribution recognition on the candidate object according to the conflict type when the consistency score is lower than a preset threshold, or when the attribute recognition result conflicts with the geometric features, spectral texture features or topological relationships of the candidate object, and to update the candidate contour relationship map. The vector result output module is used to stop feedback optimization when the consistency score reaches a preset threshold, the change in the consistency score between two consecutive iterations is less than a preset change threshold, or the number of iterations reaches the maximum number of iterations, to obtain a vector recognition result with consistent contour and attributes, and output a vector result containing contour geometry, category field, attribute field, confidence field and quality mark field.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the remote sensing image feature contour and attribute collaborative recognition method based on candidate contour relationship map as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the remote sensing image feature contour and attribute collaborative recognition method based on candidate contour relationship map as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the remote sensing image feature contour and attribute collaborative recognition method based on candidate contour relationship map as described in any one of claims 1-6.