Complex building area and space intelligent identification method and device based on multi-time-point remote sensing image data
By using an intelligent recognition framework based on high-resolution remote sensing imagery and the Swin Transformer deep learning model, the problems of high labor costs, slow data updates, and strong subjective dependence in existing technologies are solved. This enables high-precision recognition and dynamic monitoring of complex urban building areas, supporting the precision and dynamism of urban renewal strategies.
Patent Information
- Application Number
- CN202511117320.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2026-01-02
AI Technical Summary
Existing technologies for identifying urban building areas suffer from high labor costs, long data update cycles, strong subjective dependence, and identification standards that are easily influenced by the experience of researchers. Furthermore, they are difficult to capture the spatial morphological evolution of urban buildings, and traditional methods cannot adapt to rapid urban renewal.
An intelligent recognition framework based on high-resolution remote sensing imagery and the Swin Transformer deep learning model is adopted. Through multi-time point data processing and deep learning model training, high-precision quantification and dynamic monitoring of building space are achieved. Combined with diversified sample training and adaptive learning rate strategy, the model's ability to learn complex building features is improved.
It enables high-precision identification and dynamic monitoring of complex urban building areas, provides spatial decision support for differentiated urban renewal strategies, improves the accuracy and efficiency of identification, and adapts to rapid urban renewal.
Smart Images

Figure CN121259627A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of target recognition technology, specifically to a method and apparatus for intelligent recognition of complex building areas and spaces based on multi-time-point remote sensing image data. Background Technology
[0002] In the field of building area and spatial identification, existing technologies mainly revolve around manual surveys, remote sensing image analysis, and traditional machine learning. Early methods relied on manual on-site surveys to obtain parameters such as land area and building density, or on visually interpreting features like building density and street layout in remote sensing images to identify urban villages. While these methods provide accurate data at the microscale, they have significant limitations: firstly, high labor costs make it difficult to conduct dynamic monitoring at the city scale; secondly, long data update cycles fail to adapt to the rapid evolution of urban space; and thirdly, strong subjectivity makes identification standards susceptible to the experience of the surveyors. Traditional remote sensing image analysis processes require multi-step preprocessing (OBIA) including image segmentation, object merging, manual feature definition, and rule classification. This necessitates manual identification of basic elements such as buildings and roads, combining domain knowledge to construct feature engineering systems to extract statistical indicators, and reducing image data to single indicators such as building density. This not only leads to the loss of significant spatial information but also over-reliance on segmentation quality. In particular, manually defined features struggle to accurately describe complex roof structures, easily causing blurred building boundaries and over-segmentation problems. Furthermore, existing technologies lack dynamic identification mechanisms, making it difficult to capture the spatial morphological evolution of urban buildings during urban renewal. With the rapid development of cities and the continuous updating and iteration of urban land use, the high-density built environment, irregular building forms, and fragmented ownership characteristics unique to urban villages render traditional methods such as manual surveying inapplicable. There is an urgent need to seek more flexible, convenient, and scientific intelligent methods for identifying building areas and spaces. Summary of the Invention
[0003] This application focuses on the challenge of accurately identifying complex urban building areas and spaces. Addressing the limitations of high-density, heterogeneous building forms (such as irregular layouts, mixed functions, and densely packed urban villages) and traditional identification methods, it proposes an intelligent identification framework based on high-resolution remote sensing imagery and the Swin Transformer deep learning model. The aim is to achieve high-precision quantification of urban building space extent and building area, providing spatial decision support for differentiated urban renewal strategies, and promoting the precision and dynamism of urban spatial governance.
[0004] Specifically, this application provides the following technical solutions: First, this application provides a method for intelligent spatial identification of complex building areas based on multi-time-point high-resolution remote sensing image data, specifically including: Model Training and Optimization: An intelligent deep learning model was used to construct the building space recognition framework. The Microsoft COCO dataset was used as the pre-training foundation, and typical building plots in the area to be identified were selected as the training sample library. Sample labeling strictly followed the "Classification of Urban Land Use and Standard for Planning and Construction Land" (GB50137-2011), covering diverse building types such as high-rise buildings, low-rise buildings, irregularly shaped buildings, and densely spaced buildings, ensuring the representativeness and randomness of the samples. An adaptive learning rate strategy was adopted during training. Initially, a larger learning rate was used to achieve rapid convergence, and later, the model parameters were optimized by fine-tuning the learning rate. Cross-validation was also used to improve the model's ability to learn complex building features and its generalization performance.
[0005] Building space recognition process: Acquire high-resolution remote sensing images of the area to be identified, preprocess them, and input them into the trained deep learning building space recognition model to obtain preliminary building space recognition results for the area to be identified; Based on the recognition results, integrate spatial features to form a complete building area recognition map, and accurately extract the range, boundary contour, and spatial area information of the building area through fusion probability calculation and boundary optimization.
[0006] Multi-time point image preprocessing: For remote sensing images at different time points to be measured, geometric fine correction is performed based on ground control points (GCPs). The spatial resolution of the images is unified by polynomial transformation model and bilinear interpolation method to eliminate spatial deviation of time series data and ensure accurate alignment and comparability of building spatial information at each time point.
[0007] Time-series data integration and analysis: The building space identification results at each time point are correlated and merged in a time series to construct a time-series dataset of building space in the area to be tested; based on this dataset, the dynamic evolution of the building area is tracked, and the changes in building space area at different time points are quantified, providing data support for the study of regional building space evolution patterns.
[0008] Optionally, the remote sensing imagery and building sample dataset used in this application specifically include: It covers remote sensing images of buildings in mountainous areas, plains, high-rise buildings, low-rise buildings, and urban areas with different development intensities. It also includes remote sensing samples of complex building types such as buildings with irregular shapes and dense spatial layouts. By covering a variety of samples, it improves the model's adaptability to complex building scenes.
[0009] Second, this application provides an intelligent identification device for complex building areas and spaces based on multi-time-point high-resolution remote sensing image data, the device comprising: Image preprocessing and initial recognition module. This module is responsible for acquiring and preprocessing high-resolution remote sensing images of the area to be identified: performing geometric fine correction on the original images based on ground control points (GCPs), eliminating spatial distortion through a polynomial transformation model and bilinear interpolation algorithm; adaptively configuring spatial resolution parameters according to recognition accuracy requirements (supporting accuracy adjustment from 10 meters to 1000 meters) to achieve standardized processing of multi-scale images and ensure the geometric consistency of building spatial information. The preprocessed images are input into a complex building space intelligent recognition model based on Swin Transformer (this model is pre-trained on the Microsoft COCO dataset and combined with diverse building samples of the area to be identified—including feature samples such as dense building clusters and irregularly shaped buildings—trained through transfer learning and parameter optimization), outputting the initial building extent and spatial area quantification results of the area to be identified.
[0010] Multi-time point data recognition module. For remote sensing images of the area to be identified at different time points, this module reuses the above preprocessing process (ensuring the spatial reference consistency of the time series data), and sequentially inputs the processed multi-time point images into the trained deep learning model to generate recognition results of the building area boundaries, spatial distribution and area corresponding to each time point, thereby realizing the temporal dynamic capture of building space.
[0011] Spatial matching and accuracy calibration module. This module achieves accurate latitude and longitude matching of identification results at various time points through a geographic coordinate system (such as WGS84), eliminating spatial offset of time-series data; based on pixel-level probability maps and edge features output by deep learning models, it optimizes the accuracy of building boundary extraction by combining morphological operations (such as erosion and dilation), and completes the accurate quantification of building space area through pixel area conversion (combined with preset resolution parameters), ensuring that the result error is controlled within the threshold range.
[0012] The results integration and output module integrates the structured data of building area extent, boundary contours, and area at various time points, supporting multi-format output: it can output building spatial maps of a single time point in vector data (such as Shapefile) or raster data (such as TIFF), or generate panel datasets containing time dimensions (multiple time points) and spatial dimensions (latitude and longitude coordinates), meeting the application needs of dynamic monitoring, spatiotemporal analysis, etc., and is compatible with the secondary development and visualization of mainstream geographic information system (GIS) platforms.
[0013] This application constructs an intelligent recognition technology system suitable for complex urban architectural spaces based on remote sensing imagery and the Swin Transformer deep learning intelligent architecture. Its core technological innovations are reflected in three aspects: First, hierarchical multi-scale feature modeling, which gradually generates multi-resolution feature maps from H / 4×W / 4 to H / 32×W / 32 through the Patch Merging layer, realizing cross-scale information integration from individual building textures (such as roof material and outline details) to the overall layout of the block (such as building density and spatial relationships), breaking through the limitations of traditional methods with single segmentation scale; Second, the shifted window self-attention mechanism (SW-MSA), which reduces computational complexity while efficiently capturing the correlation between local features and global space through window cyclic offset and cross-window information interaction, accurately depicting the spatial topological relationship of high-density and heterogeneous buildings in the city; Third, multi-time point dynamic recognition closed loop, which relies on time-series remote sensing data, and after geometric correction and resolution standardization, combines adaptive learning rate strategy and data augmentation technology to optimize model training, forming a full-process technical framework of "data acquisition - feature extraction - probability fusion - threshold optimization", supporting dynamic monitoring and quantitative evaluation of building spatial evolution. Attached Figure Description
[0014] To clearly illustrate the technical solutions disclosed in this application, the accompanying drawings involved in the embodiments are briefly described below. The following drawings are merely illustrative examples of the technical solutions of this application. Those skilled in the art can derive other related drawings based on these drawings without creative effort, and all such drawings fall within the protection scope of this application.
[0015] Figure 1 is a flowchart illustrating the intelligent identification method for complex building areas and spaces based on multi-time-point high-resolution remote sensing image data proposed in this application. The figure fully demonstrates the entire process logic from multi-source, multi-temporal data acquisition and preprocessing, to model training and optimization, core building space identification, and multi-temporal time-series analysis. It clearly presents the connection and execution order of each step, providing an intuitive reference for understanding the overall implementation path of the method proposed in this application.
[0016] Figure 2 is a schematic diagram illustrating the operational steps of the intelligent identification method for complex building areas and spaces based on multi-time-point high-resolution remote sensing data proposed in this application. This figure further details the specific operational steps of the core identification process, including the technical details of key steps such as image preprocessing parameter configuration, model inference execution, feature fusion calculation, and boundary optimization, aiding in understanding the implementation process of the method in practical applications.
[0017] Figure 3 to Figure 12This figure presents the performance comparison results between the Swin Transformer intelligent model used in this application and the commonly used Yolov8n model in existing technologies. Through visual comparison and quantitative data display, the figure intuitively presents the differences in recognition performance between the two models in complex urban village building scenarios. Specifically, it compares dimensions such as building boundary integrity, recognition rate of small-scale auxiliary buildings, and ability to capture the layout of large-area building clusters, strongly supporting the technical advantages of the model in terms of accuracy and adaptability.
[0018] Figure 13 is a functional module architecture diagram of the intelligent identification device for complex building areas and spaces based on multi-time-point high-resolution remote sensing image data according to this application. The figure clearly shows the composition structure of the device, including the hierarchical relationship and data interaction logic of the image preprocessing and initial identification module, the multi-time-phase data identification module, the spatial matching and accuracy calibration module, and the result integration and output module. It clarifies the functional positioning and coordination mechanism of each module, providing clear guidance for understanding the overall architecture and working principle of the device. Detailed Implementation
[0019] The following will, in conjunction with Figures 1 to 13, provide a detailed explanation of the specific implementation methods and devices for intelligent identification of complex building areas and spaces in this application.
[0020] like Figure 1 The flowchart shown is a flowchart of the intelligent identification method for complex building areas and spaces based on multi-time-point high-resolution remote sensing image data in this application, including:
[0021] Step 1: Construct a deep learning intelligent model framework adapted for complex building area recognition Using the Microsoft COCO dataset as a pre-training foundation, typical building plots in the area to be identified are selected as the training sample library. The samples must be labeled in strict accordance with the "Classification of Urban Land Use and Standard for Planning and Construction Land" (GB50137-2011), covering a variety of building types such as high-rise buildings, low-rise buildings, irregularly shaped buildings, and densely laid-out buildings to ensure the representativeness and randomness of the samples. During the training process, an adaptive learning rate strategy is adopted. In the early stage, a larger learning rate is used to achieve rapid model convergence. In the later stage, the model parameters are optimized by fine-tuning the learning rate. Combined with cross-validation, the model is iterated repeatedly to improve its learning ability and generalization performance for complex building features (such as roof texture, wall boundaries, and dense layout patterns). Finally, a deep learning intelligent model that can be directly used for building space recognition is generated.
[0022] Step 2: Core Identification Process of Architectural Space Accurate identification and quantification of architectural space at a single point in time is achieved based on a trained model. Specifically, this involves: acquiring high-resolution remote sensing images of the area to be identified; improving image quality through preprocessing operations such as noise reduction and enhancement; inputting the preprocessed images into a trained deep learning architectural space recognition model to obtain preliminary architectural space identification results for the area (including pixel-level division of building and non-building areas); integrating spatial features based on the preliminary results, forming a complete architectural area identification map through operations such as merging adjacent building pixels and removing noise interference; and then accurately extracting the range, boundary contours, and spatial area information of the architectural area through probability fusion calculation (combined with the confidence level output by the model) and boundary optimization algorithms (such as morphological erosion and dilation), thus completing the identification and quantification of architectural space at a single point in time.
[0023] Step 3: Multi-time point image preprocessing (optional) This provides a standardized multi-time-point data foundation for time-series change analysis, eliminating spatial biases to ensure temporal comparability. Specifically, this includes: performing geometrical fine-tuning on remote sensing images at different time points based on ground control points (GCPs), correcting geometric distortions through a polynomial transformation model; uniformly calibrating the spatial resolution of all time-point images using bilinear interpolation (which can be set from 10 meters to 1000 meters as needed), eliminating spatial information biases caused by resolution differences; and ensuring accurate geospatial matching of images at different time points through coordinate alignment operations, laying the foundation for subsequent time-series data integration and analysis, and ensuring the comparability and consistency of building spatial information at different time points.
[0024] Step 4: Time-series data integration and dynamic analysis (optional) This study tracks the dynamic evolution of architectural space based on multi-time-point identification results. Specifically, it involves: temporally associating and merging pre-processed architectural space identification results (including scope, boundaries, and area) at various time points to construct a complete time-series dataset of architectural space in the area under test; using this dataset, tracking the expansion, contraction, or morphological changes of the architectural area through spatial overlay analysis; revealing the evolutionary trend of architectural space by quantitatively calculating indicators such as the difference in architectural space area and the proportion of change at different time points; and ultimately providing data support for research on the evolution of regional architectural space and dynamic monitoring of urban renewal, achieving a closed loop from static identification to dynamic evaluation.
[0025] This application example provides an operational flow for an intelligent identification method of complex building areas and spaces based on multi-time-point high-resolution remote sensing data, such as... Figure 2 As shown
[0026] Multi-source, multi-time point data acquisition and preprocessing Data source selection: Taking 179 urban village renewal projects in Hangzhou as the empirical subjects, five high-resolution remote sensing images (source: livingatlas.arcgis.com) were collected in 2017, 2018, 2020, 2022 and 2024, covering the temporal evolution of complex architectural spaces.
[0027] Data preprocessing: Geometric fine correction is performed on the original image based on ground control points (GCPs). The spatial resolution is uniformly calibrated to 500 meters through polynomial transformation and bilinear interpolation to ensure the accuracy of building spatial information. At the same time, data augmentation techniques such as geometric transformation and color dithering are used to expand the dataset and improve the model's ability to learn building features under different lighting and angles.
[0028] Swin Transformer Model Construction and Training Model selection criteria: The Swing Transformer was selected as the core recognition model. Its hierarchical feature modeling and shift window self-attention mechanism (SW-MSA) can effectively adapt to the multi-scale feature requirements of complex urban building complexes. It can capture the geometric details of individual buildings (such as roof texture and wall boundaries) and integrate the overall spatial layout of the block (such as building spacing and road network).
[0029] Model structure design: Hierarchical feature extraction: Multi-resolution feature maps are gradually constructed through the Patch Merging layer (from H / 4×W / 4 in Stage 1 to H / 32×W / 32 in Stage 4) to adapt to the mixed distribution of large building complexes and small illegal buildings, and to achieve cross-scale information integration; Shift window mechanism: Reduce computational complexity through windowing operations, and realize information interaction between adjacent windows through cyclic shift, solve the problems of high computational cost and fragmented local features of traditional visual Transformer global self-attention, and ensure accurate capture of spatial relationships of building clusters.
[0030] Training process optimization: The model was pre-trained using the Microsoft COCO dataset, and 80 typical urban village plots in Hangzhou were selected as training samples. The sample labeling strictly followed the "Classification of Urban Land Use and Standards for Planning and Construction Land" (GB50137-2011). An adaptive learning strategy was adopted during training (large learning in the early stage to achieve rapid convergence, and fine-tuning in the later stage), and the model parameters were optimized through cross-validation.
[0031] Multi-scale feature fusion and architectural spatial recognition process Automatic feature extraction: The model automatically extracts multi-dimensional features of the building space through a multi-level self-attention mechanism, including spectral features (roof material), texture features (building density), morphological features (outline boundaries), and topological features (spatial relationships between buildings), breaking through the limitations of traditional manual feature engineering, which is highly subjective and has a single dimension.
[0032] Building area quantification: Based on the semantic segmentation results output by the model, the building area is accurately calculated through pixel-level precision calibration (average boundary positioning error ≤ 1.2 pixels). This covers both large main buildings and small ancillary buildings such as temporary sheds and arcade ground floors, improving the completeness of recognition.
[0033] Dynamic evaluation and update adaptation mechanism. A dynamic evaluation process of "identification—feedback—optimization" is constructed: based on the identification results, a spatial distribution map of urban villages is generated; combined with urban renewal policy requirements, the identified areas are classified into three levels (core urban villages, peripheral urban villages, and potential urban villages); for areas with identification deviations (such as complex roof structures and shadowed areas), the model is iterated by increasing sample annotations and optimizing attention weights to ensure the adaptability of the technical methods to urban villages in different regions and at different stages of renewal.
[0034] To verify the technical superiority of the Swin Transformer intelligent model used in this application in the spatial recognition task of complex building areas, this application conducts a performance verification experiment with the intelligent model and the commonly used Yolov8n model in the prior art through quantitative analysis and visualization comparison. The specific experimental design and results are as follows. Figures 3 to 12 As shown: This experiment focuses on the spatial identification of buildings in urban villages. High-resolution remote sensing images of typical urban village areas in Hangzhou are selected as the test dataset, covering complex scenes such as high-density building clusters, irregular building forms, and small ancillary facilities (e.g., temporary sheds, and the ground floor of arcade buildings). Comparison objects include: The Swing Transformer model used in this application is constructed based on hierarchical feature modeling and shifted window self-attention mechanism (SW-MSA), and optimized through multi-scale sample training; The commonly used Yolov8n model in existing technologies is a single-stage detection framework that relies on the anchor box mechanism and uses a feature pyramid network (FPN) for feature aggregation. Technical advantages and experimental performance of the Swin Transformer model
[0035] Experimental results show that the Swin Transformer model exhibits significant technical advantages in the task of identifying urban village building spaces, specifically in the following aspects: Hierarchical multi-scale feature fusion capability: By constructing a hierarchical feature extraction network containing 4 Transformer Blocks, this model achieves multi-resolution hierarchical coverage from 1 / 4 to 1 / 32 of the original image size. It can efficiently aggregate semantic information (such as building functional attributes) and spatial details (such as roof texture and wall boundaries) at different scales, breaking through the limitations of single-scale feature representation.
[0036] Cross-scale spatial association capture capability: Relying on the shifted window self-attention mechanism (SW-MSA), the model can accurately establish spatial associations from micro-level individual buildings (5-10 pixel level roof texture) to macro-level area clusters (500-1000 pixel level layout structure), effectively capturing the topological relationships between individual buildings and the spatial syntax features of the cluster as a whole.
[0037] High-precision recognition and quantification capabilities: In building area calculation tasks, this model achieves sub-pixel accuracy with an average boundary positioning error of ≤1.2 pixels through pixel-level boundary calibration, which can completely identify large main buildings and small ancillary facilities, significantly improving the completeness and accuracy of the recognition results. Analysis of the limitations of existing technologies
[0038] Comparative experiments show that the Yolov8n model has significant limitations in recognizing complex urban village scenes, specifically: The problem of missed detection of small-scale buildings: Because it relies on the Feature Pyramid Network (FPN) for feature transfer, there is a "semantic gap" in the identification of small-scale buildings (roof area <30m²) - the low-level features lack the semantic guidance of the high-level features, resulting in a large number of missed detections of temporary sheds, arcade ground floors and other auxiliary buildings.
[0039] Large-scale scene integration defects: The model adopts a local feature aggregation method based on convolution operation, which makes it difficult to capture long-distance spatial relationships between building clusters. In dense building clusters, it is prone to excessive merging of boundaries (blurred boundaries between adjacent buildings) or breakage (incomplete building outlines).
[0040] Low recognition completeness: The above-mentioned dual defects directly result in a significant reduction in the building area recognition completeness of the Yolov8n model compared to the Swin Transformer model, which cannot meet the actual needs of accurate recognition of complex building areas.
[0041] Through model performance comparison experiments, it is clear that the Swin Transformer model adopted in this application, through its hierarchical multi-scale feature fusion architecture and shift window self-attention mechanism, effectively overcomes the shortcomings of existing single-stage detection models that rely on anchor frames, such as insufficient scale adaptability and weak spatial correlation capture ability in complex building scenarios. It shows significant advantages in the accuracy, completeness and adaptability to complex scenarios of urban village building space recognition, and provides reliable technical support for subsequent building space scope definition, area quantification and dynamic monitoring.
[0042] Based on the same inventive concept, this application provides an intelligent identification device for complex building areas and spaces based on multi-time-point high-resolution remote sensing image data, such as... Figure 13 As shown, modular design enables precise identification, temporal tracking, and result output of architectural spaces, specifically including the following functional modules: Image preprocessing and initial recognition module
[0043] This module is the core unit for basic data processing and initial identification of the device. Its main functions include: Image Acquisition and Preprocessing: High-resolution remote sensing images of the area to be identified are acquired. Geometric fine correction is performed on the original images based on ground control points (GCPs). Spatial distortion of the images is corrected through a polynomial transformation model, and image pixel resampling is achieved using a bilinear interpolation algorithm. According to the actual recognition accuracy requirements, spatial resolution parameters are adaptively configured (supporting accuracy adjustment from 10 meters to 1000 meters) to complete the standardization processing of multi-scale images, ensuring the consistency and comparability of building spatial information in geometric dimensions.
[0044] Initial recognition operation: The preprocessed standardized image is input into a complex architectural space deep learning recognition model based on Swin Transformer (this model is pre-trained on the Microsoft COCO dataset and integrates diverse architectural samples of the area to be identified - including feature samples such as dense building clusters and irregularly shaped buildings - optimized by transfer learning and iterative parameter training). Through feature extraction and inference operations of the model, the initial architectural space range, boundary contour and area quantification results of the area to be identified are output. Multi-time point data recognition module
[0045] This module is a time-series data processing and dynamic recognition unit, used to achieve continuous recognition of building spaces at different time points. Its specific functions are as follows: for multi-time point remote sensing images of the area to be identified (such as image data from different years and quarters), it reuses the standardized preprocessing process (including geometric fine correction, resolution unification, etc.) in the image preprocessing and initial recognition module to ensure the consistency of multi-time point images on the spatial reference; the processed images of each time point are sequentially input into the trained deep learning model to generate the building area boundary, spatial distribution characteristics and area quantification results for the corresponding time point, so as to achieve accurate capture of the temporal dynamic changes of building space. Spatial matching and accuracy calibration module
[0046] This module is a unit for optimizing recognition results and ensuring accuracy. It is used to improve the spatial consistency and quantitative accuracy of recognition results. Specific functions include: Spatial matching: The latitude and longitude of the building spatial identification results at each time point are accurately aligned using a geographic coordinate system (such as the WGS84 coordinate system) to eliminate spatial deviations caused by image offsets in time series data and ensure the comparability of identification results at different time points in geographic space.
[0047] Accuracy calibration: Based on the pixel-level architectural space probability map and edge features output by the deep learning model, the accuracy of building boundary extraction is optimized by combining morphological operations (such as erosion operation to remove noise and dilation operation to repair boundary fractures); according to the preset spatial resolution parameters, the area of the building space is accurately quantified by the pixel area conversion formula to ensure that the area calculation error is controlled within the preset threshold range. Results Integration and Output Module
[0048] This module is a data integration and application output unit designed to meet diverse scenario requirements. Its specific functions include: structurally integrating the building area range, boundary contour vector data, and area quantification results at various time points; supporting multi-format data output: it can output building spatial maps at a single time point (including vector formats such as Shapefile or raster formats such as TIFF), and can also generate panel datasets containing time dimensions (multi-time point sequences) and spatial dimensions (latitude and longitude coordinates); the output results are compatible with mainstream Geographic Information System (GIS) platforms and can be directly used for secondary development and visualization of application scenarios such as dynamic monitoring, spatiotemporal evolution analysis, and urban renewal planning.
[0049] Through the coordinated operation of the above modules, this device achieves fully automated processing from multi-time point image preprocessing and intelligent building space recognition to result optimization and output, significantly improving the accuracy, efficiency and dynamic adaptability of spatial recognition in complex building areas.
Claims
1. A method and device for intelligent identification of complex building areas and spaces based on multi-temporal remote sensing image data. Its core technological innovations are reflected in three aspects: First, hierarchical multi-scale feature modeling, which gradually generates multi-resolution feature maps from H / 4×W / 4 to H / 32×W / 32 through the Patch Merging layer, realizing cross-scale information integration from individual building textures (such as roof material and outline details) to the overall layout of the block (such as building density and spatial relationships), breaking through the limitations of the single segmentation scale of traditional methods; Second, a shift window self-attention mechanism, which reduces computational complexity while efficiently capturing the correlation between local features and global space through window cyclic offset and cross-window information interaction, accurately depicting the spatial topological relationship of high-density and heterogeneous buildings in the city; Third, a multi-temporal dynamic identification closed loop, which relies on time-series remote sensing data, and after geometric correction and resolution standardization, combines adaptive learning rate strategy and data augmentation technology to optimize model training, forming a full-process technical framework of "data acquisition-feature extraction-probability fusion-threshold optimization", supporting dynamic monitoring and quantitative evaluation of building spatial evolution.
2. A method for intelligent identification of complex building areas and spaces based on multi-time-point remote sensing image data, specifically including: First, model training and optimization. A deep learning intelligent model was used to construct a building space recognition framework. The Microsoft COCO dataset was used as the pre-training basis, and typical building plots in the area to be identified were selected as the training sample library. The sample annotation strictly followed the "Classification of Urban Land Use and Standard for Planning and Construction Land" (GB50137-2011), covering a variety of building types such as high-rise buildings, low-rise buildings, irregularly shaped buildings, and densely spaced buildings, to ensure the representativeness and randomness of the samples. An adaptive learning rate strategy is adopted during training. In the early stage, a larger learning rate is used to achieve rapid convergence. In the later stage, the model parameters are optimized by fine-tuning the learning rate, and cross-validation is combined to improve the model's learning ability and generalization performance for complex building features. Second, the building space recognition process. High-resolution remote sensing images of the area to be identified are acquired, preprocessed, and then input into a trained intelligent building space recognition model to obtain preliminary building space recognition results for the area to be identified. Based on the recognition results, spatial features are integrated to form a complete building area recognition map. Through probability fusion calculation and boundary optimization, the range, boundary contours, and spatial area information of the building area are accurately extracted.
3. A method for intelligent identification of complex building areas and spaces based on multi-time point remote sensing image data, wherein optional methods for analyzing the temporal changes of building space based on multi-time point data include: Multi-temporal image preprocessing: For remote sensing images at different time points, geometric fine correction is performed based on ground control points (GCPS). A polynomial transformation model and bilinear interpolation are used to unify the spatial resolution of the images, eliminating spatial biases in the temporal data and ensuring accurate alignment and comparability of building spatial information at each time point. Temporal data integration and analysis: The building spatial identification results at each time point are temporally correlated and merged to construct a time-series dataset of building space in the area under test. Based on this dataset, we can track the dynamic evolution of building area boundaries, quantify the changes in building space area at different points in time, and provide data support for the study of regional building space evolution patterns.
4. A method for intelligent identification of complex building areas and spaces based on multi-time-point remote sensing image data. Optionally, the remote sensing images and building sample dataset used in this application specifically include: remote sensing images of buildings in mountainous areas, remote sensing images of buildings in plains, remote sensing images of high-rise buildings, remote sensing images of low-rise buildings, and remote sensing images of buildings in urban areas with different development intensities. At the same time, remote sensing samples of complex building types such as buildings with irregular shapes and buildings with dense spatial layouts are also included. The ability of the model to adapt to complex building scenes is improved through diversified sample coverage.
5. A device for intelligent identification of complex building areas and spaces based on multi-time-point remote sensing image data, the device comprising: The first module is the image preprocessing and initial recognition module. This module is responsible for acquiring and preprocessing high-resolution remote sensing images of the area to be identified: performing geometric fine correction on the original images based on ground control points (GCPs), eliminating spatial distortion through a polynomial transformation model and bilinear interpolation algorithm; adaptively configuring spatial resolution parameters according to recognition accuracy requirements (supporting accuracy adjustment from 10 meters to 1000 meters) to achieve standardized processing of multi-scale images and ensure the geometric consistency of building spatial information. The preprocessed images are input into a complex building space deep learning recognition model based on Swin Transformer (this model is pre-trained on the Microsoft COCO dataset and combined with diverse building samples of the area to be identified—including feature samples such as dense building clusters and irregularly shaped buildings—trained through transfer learning and parameter optimization), outputting the initial building extent and spatial area quantification results of the area to be identified; the second module is the multi-time point data recognition module. For remote sensing images of the area to be identified at different time points, this module reuses the above preprocessing process (to ensure the spatial reference consistency of the time series data), and sequentially inputs the processed multi-temporal images into the trained deep learning model to generate the identification results of the building area boundary, spatial distribution and area corresponding to each time point, so as to realize the temporal dynamic capture of building space. Third is the spatial matching and accuracy calibration module. This module uses a geographic coordinate system (such as WGS84) to achieve accurate latitude and longitude matching of the identification results at each time point, eliminating spatial offset of time series data; Based on the pixel-level probability map and edge features output by the deep learning model, the accuracy of building boundary extraction is optimized by combining morphological operations (such as erosion and dilation), and the building space area is accurately quantified by pixel area conversion (combined with preset resolution parameters) to ensure that the result error is controlled within the threshold range. Fourth is the results integration and output module. This module performs structured integration of building area range, boundary outline, and area data at various time points, and supports multi-format output: it can output building spatial maps of a single time point in vector data (such as Shapefile) or raster data (such as TIFF), and can also generate panel datasets containing time dimensions (multi-temporal) and spatial dimensions (latitude and longitude coordinates), meeting the application needs of dynamic monitoring, spatiotemporal analysis, etc., and is compatible with the secondary development and visualization of mainstream geographic information system (GIS) platforms.