Building storey-adding intelligent identification method and system based on multi-mode remote sensing data
By using multimodal remote sensing data and deep learning models, combined with optical and radar imagery, the low efficiency and accuracy issues of traditional building addition detection have been resolved, enabling efficient and accurate monitoring and management of building additions in large urban areas.
Patent Information
- Application Number
- CN202510861212.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional building floor addition detection relies on manual inspections, which are inefficient and costly, making it difficult to achieve continuous monitoring and timely intervention in large urban areas. Aerial imaging methods are also limited by flight paths and weather factors, affecting detection accuracy and coverage.
An intelligent identification method for building storey additions based on multimodal remote sensing data is adopted. A deep learning prediction model is combined with optical and radar remote sensing images. Through adaptive statistical thresholds and spatial consistency analysis, changes in building heights are identified. An observation dataset is constructed and multi-view, multi-modal remote sensing image training is performed to achieve accurate judgment of building storey additions.
It has achieved full coverage monitoring of large urban areas, improved the accuracy and timeliness of identifying building additions, reduced the false alarm rate and missed alarm rate, provided urban management support from a global perspective, and enabled timely detection of illegal building additions.
Smart Images

Figure SMS_5 
Figure SMS_13 
Figure SMS_22
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of building detection, and in particular relates to a method and system for intelligently identifying building floors based on multimodal remote sensing data. Background Art
[0002] With the acceleration of urbanization and the continuous increase in urban building density, a large number of unauthorized, self-built homes have emerged in some areas. These buildings generally lack professional design and construction, and some residents have added floors without authorization to expand their living space, posing serious structural safety risks. Therefore, regular building addition inspections are necessary to eliminate potential risks.
[0003] Traditional building floor addition detection mainly relies on manual inspections, which have problems such as high labor costs, low efficiency, and long cycles, making it difficult to achieve continuous monitoring and timely intervention in large urban areas.
[0004] In recent years, with the development of image recognition technology, some technologies have adopted aerial images to identify building floors. However, such methods still have the following defects: First, aerial images are limited by factors such as flight paths, weather and operating conditions, and it is difficult to ensure that the image position and angle are completely consistent during multiple shots, which affects the accuracy of subsequent change detection; second, the coverage of aerial photography operations is relatively limited, which cannot meet the needs of efficient and continuous monitoring of large urban areas, restricting the application and promotion of technology in refined urban management.
[0005] Therefore, it is necessary to provide a method and system for intelligent identification of building floors based on multimodal remote sensing data to solve the above problems. Summary of the Invention
[0006] The present invention provides a method and system for intelligently identifying building storey additions based on multimodal remote sensing data. Multi-perspective, multimodal remote sensing images of buildings are used to train a deep learning prediction model to predict building heights. The remote sensing image detection method is well suited for detecting building storey additions over a wide range, and data from different perspectives and modalities can avoid recognition errors caused by single-source data. Furthermore, through adaptive statistical thresholds and spatial consistency analysis methods, the target building and adjacent buildings within a certain range are treated as a whole to determine storey additions, effectively improving recognition accuracy and thereby resolving at least one technical problem described in the background art.
[0007] In order to solve the above-mentioned technical problems, the present invention is achieved as follows: A method for intelligently identifying building floors based on multimodal remote sensing data comprises the following steps: Step S1: constructing an observation dataset, wherein the observation dataset includes remote sensing images of multiple buildings at different viewing angles and in different modalities, with the height of the corresponding building in each remote sensing image being used as a label; Step S2: construct a deep learning prediction model, use the observation data set to train the deep learning prediction model, and learn the mapping relationship between the remote sensing image and the height of the building; Step S3: For any building floor recognition task, a time series containing multiple consecutive time nodes is selected. At each time node, a remote sensing image of the target building is collected and input into the trained deep learning prediction model. The predicted height value of the target building at the current time node is output. The first-order difference between the predicted height value of the target building at the current time node and the predicted height value at the previous time node is calculated as the height change value of the target building. All time nodes are traversed to form a target building height change sequence. Step S4: Select a target area, calculate the height change sequence of all buildings in the target area according to step S3, set a first threshold value at any time point based on the height change value of all buildings in the target area, and if the height change value of the target building exceeds the first threshold value, it is preliminarily determined that the target building has added a floor, and step S5 is executed; otherwise, it is determined that the target building has not added a floor, and the determination ends; Step S5: Use the weighted local spatial anomaly measurement method to calculate the height change consistency index between the target building and other buildings in the target area, set a second threshold, and if the height change consistency index exceeds the second threshold, it is finally determined that the target building has an additional floor, and the determination ends; otherwise, it is finally determined that the target building does not have an additional floor, and the determination ends.
[0008] As a preferred improvement, the building remote sensing images include optical remote sensing images and radar remote sensing images, wherein the optical remote sensing images include high-resolution optical images and wide-area optical images.
[0009] As an optimal improvement, the observation dataset needs to be preprocessed before being input into the deep learning prediction model. The preprocessing process includes geometric correction, orthorectification, radiation correction and image enhancement to ensure data quality. At the same time, the sliding window technology is used to crop the remote sensing image into the required standard size, and a zero-value filling strategy is adopted for the edge areas that cannot be divided evenly to ensure data integrity.
[0010] As a preferred improvement, the deep learning prediction model includes a multimodal feature extraction module, a feature fusion module, a spatial context information enhancement module and an output module. The multimodal feature extraction module uses a long and short attention mechanism to extract multi-scale building features from optical remote sensing images and radar remote sensing images respectively; the feature fusion module uses an adaptive fusion attention mechanism to fuse building features of different modalities and scales to obtain fusion features; the spatial context information enhancement module captures the relationship between the building and the surrounding environment to enhance the fusion features; the output module is used to process the fusion features and output the predicted value of the building height.
[0011] As an optimal improvement, the fusion dice loss is used in the deep learning prediction model training process. , boundary loss and Lovas loss The composite loss function , expressed as: Where, 、 、 Represents dice loss respectively , boundary loss and Lovas loss The weight coefficient of ; where: Dice loss Expressed as: Where, Represents any pixel in a remote sensing image; Represents the total number of pixels in the remote sensing image; Represents the deep learning prediction model for pixels predicted probability; Represents pixels The true pixel value of Indicates the minimum value, used to prevent the denominator from being 0; Boundary loss Expressed as: Where, Represents pixels SDM value; Lovas loss Expressed as: Where, Represents a collection of pixels; Represents the sorted IoU loss gradient sequence; Represents any element in the IoU loss gradient sequence.
[0012] As a preferred improvement, in step S4, the target area is determined by the following process: Step S41: Grid the remote sensing image and calculate the building density in each grid unit. The calculation process is expressed as follows: Where, represents the building density within the grid cell; Represents the total building area within the grid cell, calculated from the building outline data; represents the total area of the grid cells; Step S42: The target area is determined in each grid cell using a sliding window method. The sliding window size is expressed as: Where, Indicates the sliding window size; Indicates the base window size; represents the density adjustment coefficient; Step S43: Select multiple target areas one by one according to the preset window overlap rate and step size, where: Window overlap ratio Expressed as: Where, represents the baseline overlap ratio; represents the overlap adjustment coefficient; step length Expressed as: .
[0013] As a preferred improvement, the process of setting the first threshold specifically includes the following steps: If the height change values of all buildings in the target area obey the normal distribution, the first threshold Set to: Where, Represents the average value of all building height changes in the target area; The coefficient factor that represents the control sensitivity, with a value range of 2-3; Represents the standard deviation of all building height changes within the target area; represents the regional heterogeneity adjustment coefficient, with a value range of 0.1-0.5; represents the coefficient of variation, ; Indicates the total number of samples of buildings in the target area; Indicates the number of benchmark samples, with a value of 50-100; If the height variation of all buildings in the target area is severely skewed, the first threshold Use any of the following methods to set: First: Where, Represents the median of all building height changes within the target area; Represents the absolute median difference of all building height changes in the target area; Second: Where, It represents the 95th percentile of all building height changes in the target area. That is, if the height change of the target building exceeds the height change of 95% of the samples in the area, it is considered abnormal.
[0014] As a preferred improvement, the height change consistency index is represented by the spatial anomaly metric value, and its calculation process is expressed as follows: Where, Indicates the target building The spatial anomaly measure of Indicates the target building The height change value; Indicates buildings within the target area The height change value; Represents the spatial weight matrix, and the target building and adjacent buildings is inversely proportional to the distance; Indicates that the target building is included target area; Indicates the target area The standard deviation of the variation in height of all buildings within the .
[0015] As a preferred improvement, the second threshold value ranges from 2.5 to 3.5, wherein: for densely built areas, the second threshold value is 3.0; for sparsely built areas, the second threshold value is 2.5; for special supervision areas, the second threshold value is 3.5.
[0016] A system for executing the above-mentioned intelligent identification method for building floors based on multimodal remote sensing data, comprising: An observation dataset, comprising remote sensing images of multiple buildings at different viewing angles and in different modalities, with the height of the corresponding building in each remote sensing image used as a label; A deep learning prediction model is trained using an observation dataset to learn the mapping relationship between remote sensing images and building heights. For any building floor recognition task, a time series containing multiple consecutive time nodes is selected. At each time node, a remote sensing image of the target building is collected and input into the trained deep learning prediction model. The predicted height value of the target building at the current time node is output, and the first-order difference between the predicted height value of the target building at the current time node and the predicted height value at the previous time node is calculated as the height change value of the target building. This is then traversed through all time nodes to form a target building height change sequence. A preliminary determination module is configured to select a target area, calculate a height change sequence of all buildings in the target area, and set a first threshold at any time point based on the height change values of all buildings in the target area. If the height change value of the target building exceeds the first threshold, it is preliminarily determined that the target building has an additional floor, and step S5 is executed; otherwise, it is determined that the target building has not an additional floor, and the determination ends. The final judgment module is used to calculate the height change consistency index between the target building and other buildings in the target area using a weighted local spatial anomaly measurement method, set a second threshold, and if the height change consistency index exceeds the second threshold, it is finally determined that the target building has an additional floor and the judgment ends; otherwise, it is finally determined that the target building does not have an additional floor and the judgment ends.
[0017] The beneficial effects of the present invention are: (1) The detection method based on high-resolution remote sensing images breaks through the spatial limitations of traditional monitoring methods and can achieve full coverage monitoring of large urban areas. Whether in the core area or the periphery of the city, the system can provide unified standard monitoring services, solving the problems of uneven coverage and blind spots of traditional methods, and providing a global perspective for urban management. (2) By constructing observation data sets through remote sensing images of different perspectives and modalities, the blind spots and occlusion problems in traditional single-perspective remote sensing monitoring are effectively solved. The system can capture the three-dimensional structural characteristics of buildings from different angles. Through perspective complementarity and integrated learning methods, the accuracy and completeness of building height estimation are significantly improved. The multi-perspective observation strategy is well suited for high-density urban environments and can effectively reduce the monitoring blind spots caused by occlusion of high-rise buildings, thus achieving comprehensive coverage of illegal floor addition monitoring. (3) Deep learning is used to directly predict building height values, rather than the traditional change detection method based on image difference. This method significantly improves detection accuracy and effectively avoids the inherent defects of the image difference method that are affected by lighting conditions, seasonal changes, and shooting angles. It can accurately identify the true vertical changes of buildings and reduce the false alarm rate and missed alarm rate. (4) The design of adaptive statistical thresholds and spatial consistency thresholds can effectively distinguish normal building changes from illegal floor additions, significantly reducing the false alarm rate. The system can intelligently identify and filter changes in non-illegal floor additions such as road vehicles and temporary buildings, thereby improving the accuracy and reliability of identification. (5) By setting a reasonable data collection frequency, a continuous monitoring mechanism for illegal floor additions can be implemented. The system regularly obtains the latest remote sensing data of the target area, which can capture the change information of building height in a timely manner, and realize early detection and intervention of illegal floor additions. This dynamic monitoring mode significantly improves the timeliness of supervision and effectively prevents the spread of illegal construction. (6) According to the real-time detection results, a list of illegal additional-storey buildings and a spatial distribution map can be automatically generated, providing decision-making support for urban management and law enforcement departments. It is of great significance to promote urban building safety supervision, prevent building safety risks, and protect the lives and property of urban residents. DETAILED DESCRIPTION
[0018] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] This embodiment provides a method for intelligently identifying building floors based on multimodal remote sensing data, comprising the following steps: Step S1: construct an observation dataset, which includes remote sensing images of multiple different buildings at different viewing angles and in different modalities, with the height of the corresponding building in each remote sensing image being used as a label.
[0020] Remote sensing imaging technology uses sensors carried on satellites, drones and other carriers to detect and record electromagnetic wave information emitted and reflected by buildings from a distance, and present it in the form of images. In this embodiment, satellites are used as carriers of sensors, and the sensors include optical sensors and radar sensors. Therefore, the obtained remote sensing images of buildings include optical remote sensing images and radar remote sensing images, among which optical remote sensing images include high-resolution optical images and wide-width optical images. The detection method based on high-resolution remote sensing images breaks through the spatial limitations of traditional monitoring methods and can achieve full coverage monitoring of large urban areas; whether in the core area or the peripheral areas of the city, unified standard monitoring services can be provided, solving the problems of uneven coverage and blind spots in supervision of traditional methods, and providing a global perspective for urban management.
[0021] During the remote sensing image acquisition process, image data from multiple satellites should be obtained simultaneously, and optical remote sensing images should be grouped according to the shooting time and coverage area. Each group of optical remote sensing images can be matched with multiple radar remote sensing images to form a multi-perspective observation data set. Each group of observation data sets contains building remote sensing images under different perspectives and different modes (optical mode and radar mode), which can ensure the diversity of data samples, improve the generalization ability of subsequent models, help the model better adapt to different inputs, avoid errors caused by a single data source, and increase the accuracy and robustness of predictions.
[0022] The source of building height is radar point cloud data, DSM or DEM data. It should be noted that in the labeling process, attention should be paid to the temporal and spatial consistency issues, that is, the time and spatial position of the label must be consistent with the time and spatial position corresponding to the target building in the building remote sensing image.
[0023] The observation dataset needs to be preprocessed. The preprocessing process includes geometric correction, orthorectification, radiation correction and image enhancement to ensure data quality. At the same time, the sliding window technology is used to crop the remote sensing image into the required standard size, and the zero-value filling strategy is adopted for the edge areas that cannot be divided evenly to ensure data integrity.
[0024] Step S2: construct a deep learning prediction model, use the observation data set to train the deep learning prediction model, and learn the mapping relationship between the remote sensing image of the building and its height.
[0025] It should be noted that the construction and processing of the deep learning prediction model belong to the existing technology in this field. In this embodiment, the deep learning prediction model improves the loss function based on the traditional deep learning prediction model, and the other architectures remain consistent with the architecture of the traditional deep learning prediction model.
[0026] Deep learning is used to directly predict building height values, rather than the traditional change detection method based on image difference. This method significantly improves detection accuracy and effectively avoids the inherent defects of the image difference method that is affected by lighting conditions, seasonal changes, and shooting angles. It can accurately identify the true vertical changes of buildings and reduce false alarm and omission rates.
[0027] Specifically, the deep learning prediction model includes a multimodal feature extraction module, a feature fusion module, a spatial context information enhancement module and an output module. The multimodal feature extraction module uses a long and short attention mechanism to extract multi-scale building features from optical remote sensing images and radar remote sensing images respectively; the feature fusion module uses an adaptive fusion attention mechanism to fuse building features of different modalities and scales to obtain fused features; the spatial context information enhancement module captures the relationship between the building and the surrounding environment to enhance the fused features; the output module is used to process the fused features and output the predicted value of the building height.
[0028] In this embodiment, a fusion dice loss is designed. , boundary loss and Lovas loss The composite loss function , expressed as: Where, 、 、 Represents dice loss respectively , boundary loss and Lovas loss The weight coefficient of ; where: Dice loss Focus on improving the overall segmentation quality and solving the problem of category imbalance, expressed as: Where, Represents any pixel in a remote sensing image; Represents the total number of pixels in the remote sensing image; Represents the deep learning prediction model for pixels predicted probability; Represents pixels The true pixel value of Indicates the minimum value, used to prevent the denominator from being 0; Boundary loss Focus on improving the building boundary prediction accuracy, expressed as: Where, Represents pixels SDM value; Lovas loss It is used to optimize the IoU performance of the model and improve the spatial consistency of high-level predictions, expressed as: Where, Represents a collection of pixels; Represents the sorted IoU loss gradient sequence; Represents any element in the IoU loss gradient sequence.
[0029] The composite loss function provided in this embodiment Incorporates dice loss , boundary loss and Lovas loss The advantages of the three can greatly improve the prediction accuracy of the model.
[0030] During the deep learning prediction model training process, a preheated learning rate adjustment strategy is adopted to optimize the model convergence process. A phased training strategy is adopted, combined with an early stopping mechanism and model checkpoint preservation to ensure the model generalization ability.
[0031] Step S3: For any building floor recognition task, a time series containing multiple continuous time nodes is selected. At each time node, a remote sensing image of the target building is collected, and the trained deep learning prediction model is input. The height prediction value of the target building at the current time node is output, and the first-order difference between the target building height prediction value at the current time node and the height prediction value at the previous time node is calculated as the height change value of the target building. All time nodes are traversed to form a target building height change sequence.
[0032] The height difference between two adjacent time points reflects the height change of the target building over a short period of time, including whether a height change has occurred, the rate of change, and the magnitude of the change. Preferably, the time series includes at least three time points to eliminate interference from seasonal factors and temporary structural changes.
[0033] Step S4, select the target area, calculate the height change sequence of all buildings in the target area according to step S3, set a first threshold value at any time node according to the height change value of all buildings in the target area, if the height change value of the target building exceeds the first threshold value, it is preliminarily determined that the target building has an additional floor, and step S5 is executed; otherwise, it is determined that the target building does not have an additional floor, and the determination ends.
[0034] The size of the target area is determined by the following process: Step S41: Grid the remote sensing image and calculate the building density in each grid unit. The calculation process is expressed as follows: Where, represents the building density within the grid cell; Represents the total building area within the grid cell, calculated from the building outline data; Represents the total area of the grid cells.
[0035] In this embodiment, the size of the grid unit is 10km×10km, and the total area of the grid unit is 100km 2 .
[0036] Step S42: The target area is determined in each grid cell using a sliding window method. The sliding window size is expressed as: Where, Indicates the sliding window size; The reference window size can be defined according to different regions. In this implementation, it is selected as 0.5 km × 0.5 km. It represents the density adjustment coefficient, with a value range of 0.5-1.5, and is selected as 1.0 in this embodiment.
[0037] The sliding window size is dynamically adjusted according to the building density, ensuring that areas with low building density use larger windows to include enough building samples, and areas with high building density use smaller windows to reduce mutual interference between buildings.
[0038] Step S43: Select multiple target areas one by one according to the preset window overlap rate and step size, where: Window overlap ratio Expressed as: Where, represents the baseline overlap ratio; Represents the overlap adjustment coefficient.
[0039] The window overlap rate is also dynamically adjusted according to the building density, ensuring that areas with higher building density have a higher overlap rate, thereby improving the accuracy and continuity of building floor recognition.
[0040] step length Expressed as: .
[0041] When a building has additional floors, its height will change significantly. The first threshold is set as the allowable height change value of the building. Once the height change value of the target building exceeds the first threshold, it indicates that the target building has additional floors.
[0042] The process of setting the first threshold specifically includes the following steps: If the height change values of all buildings in the target area obey the normal distribution, the first threshold Set to: Where, Represents the average value of all building height changes in the target area; The coefficient factor represents the control sensitivity, with a value range of 2-3. A value of 2 covers approximately 95% of normal changes, which is suitable for scenarios with a small tolerance for false alarms. A value of 3 covers approximately 99.7% of normal changes, which is suitable for strict regulatory scenarios. Represents the standard deviation of all building height changes within the target area; represents the regional heterogeneity adjustment coefficient, with a value range of 0.1-0.5; represents the coefficient of variation, ; Indicates the total number of samples of buildings in the target area; Indicates the number of benchmark samples, ranging from 50 to 100.
[0043] If the height variation of all buildings in the target area is severely skewed, the first threshold Use any of the following methods to set: First: Where, Represents the median of all building height changes within the target area; Represents the absolute median difference of all building height changes in the target area; Second: Where, It represents the 95th percentile of all building height changes in the target area. That is, if the height change of the target building exceeds the height change of 95% of the samples in the area, it is considered abnormal.
[0044] It is understandable that the first threshold is not a fixed value in traditional determination technology, but is adaptively adjusted according to the height changes of all buildings in the target area, which can significantly improve the accuracy of the determination.
[0045] Step S5: Use the weighted local spatial anomaly measurement method to calculate the height change consistency index between the target building and other buildings in the target area, set a second threshold, and if the height change consistency index exceeds the second threshold, it is finally determined that the target building has an additional floor, and the determination ends; otherwise, it is finally determined that the target building does not have an additional floor, and the determination ends.
[0046] Illegal addition of floors to a building is generally manifested as a sudden change in local height, while the height of surrounding buildings remains unchanged. In order to quantify this phenomenon, the present invention studies the adjacent area within a certain range of the target building as a whole, comprehensively considers the height changes of all buildings in the adjacent area of the target building, and characterizes the height change trend of the target building compared with the adjacent buildings by calculating the height change consistency index of the target building and other buildings in the target area. The height change consistency index is represented by a spatial anomaly measurement value. If the spatial anomaly measurement value exceeds the second threshold, it indicates that the height change of the target building is inconsistent with the height change of the adjacent buildings, and it can be determined that the target building has illegally added floors; otherwise, it indicates that the height change of the target building is consistent with the height change of the adjacent buildings, and it can be determined that the target building has not illegally added floors. This judgment method can effectively eliminate the height prediction error caused by the limitation of model accuracy and improve the prediction accuracy of the model.
[0047] The calculation process of spatial anomaly metric is expressed as: Where, Indicates the target building The spatial anomaly measure of Indicates the target building The height change value; Indicates buildings within the target area The height change value; Represents the spatial weight matrix, and the target building and adjacent buildings is inversely proportional to the distance; Indicates that the target building is included target area; Indicates the target area The standard deviation of the variation in height of all buildings within the .
[0048] The value range of the second threshold is 2.5-3.5. For areas with dense buildings, it is recommended to use a second threshold value of 3.0; for areas with sparse buildings, the second threshold value is 2.5; for special supervision areas, the second threshold value is 3.5 to reduce false alarms. The threshold selection should be based on experimental verification, and the optimal threshold should be determined by analyzing the sample area. If necessary, the receiver operating characteristic curve (ROC) analysis can be used to determine the optimal detection threshold for a specific area.
[0049] This application uses two assessments. The first uses the height change of a single target building to preliminarily screen for buildings with abnormal heights. However, this still cannot rule out changes in non-illegal additions such as road vehicles and temporary structures. The second assessment ultimately determines whether the target building has experienced illegal additions by comparing its height change trend with other buildings in the target area. If the target building's height change trend is consistent with that of other buildings, illegal additions can be ruled out. The combination of these two assessments greatly improves the accuracy of determining illegal additions and avoids misjudgments.
[0050] After the judgment is completed, key information such as the precise coordinates of the illegally added-floor buildings, the height of the added floors, remote sensing image screenshots, etc. are generated and output to form a list of illegally added-floor buildings and a spatial distribution map, providing scientific basis and technical support for subsequent law enforcement actions.
[0051] Once the model framework is established, it's only necessary to set a reasonable data collection frequency to collect remote sensing images of all buildings within the target area. This allows for continuous monitoring of illegal floor additions, enabling early detection and intervention. This dynamic monitoring model significantly improves the timeliness of supervision and effectively prevents the spread of illegal construction. Furthermore, as monitoring progresses, the network parameters of the deep learning prediction model are continuously updated and revised, further improving the accuracy of the predictions.
[0052] This embodiment further provides a system for the above-mentioned method for intelligently identifying building floors based on multimodal remote sensing data, comprising: An observation dataset, comprising remote sensing images of multiple buildings at different viewing angles and in different modalities, with the height of the corresponding building in each remote sensing image used as a label; A deep learning prediction model is trained using an observation dataset to learn the mapping relationship between remote sensing images and building heights. For any building floor recognition task, a time series containing multiple consecutive time nodes is selected. At each time node, a remote sensing image of the target building is collected and input into the trained deep learning prediction model. The predicted height value of the target building at the current time node is output, and the first-order difference between the predicted height value of the target building at the current time node and the predicted height value at the previous time node is calculated as the height change value of the target building. This is then traversed through all time nodes to form a target building height change sequence. A preliminary determination module is configured to select a target area, calculate a height change sequence of all buildings in the target area, and set a first threshold at any time point based on the height change values of all buildings in the target area. If the height change value of the target building exceeds the first threshold, it is preliminarily determined that the target building has an additional floor, and step S5 is executed; otherwise, it is determined that the target building has not an additional floor, and the determination ends. The final judgment module is used to calculate the height change consistency index between the target building and other buildings in the target area using a weighted local spatial anomaly measurement method, set a second threshold, and if the height change consistency index exceeds the second threshold, it is finally determined that the target building has an additional floor and the judgment ends; otherwise, it is finally determined that the target building does not have an additional floor and the judgment ends.
[0053] This embodiment further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for intelligent identification of added floors in a building are implemented.
[0054] The above describes the embodiments of the present invention, but the present invention is not limited to the above specific implementation methods. The above specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A method for intelligent identification of building floors based on multimodal remote sensing data, characterized in that: The steps include: Step S1: constructing an observation dataset, wherein the observation dataset includes remote sensing images of multiple buildings at different viewing angles and in different modalities, with the height of the corresponding building in each remote sensing image being used as a label; Step S2: construct a deep learning prediction model, use the observation data set to train the deep learning prediction model, and learn the mapping relationship between the remote sensing image and the height of the building; Step S3: For any building floor recognition task, a time series containing multiple consecutive time nodes is selected. At each time node, a remote sensing image of the target building is collected and input into the trained deep learning prediction model. The predicted height value of the target building at the current time node is output. The first-order difference between the predicted height value of the target building at the current time node and the predicted height value at the previous time node is calculated as the height change value of the target building. All time nodes are traversed to form a target building height change sequence. Step S4: Select a target area, calculate the height change sequence of all buildings in the target area according to step S3, set a first threshold value based on the height change value of all buildings in the target area at any time point, and if the height change value of the target building exceeds the first threshold value, it is preliminarily determined that the target building has added floors, and step S5 is executed; Otherwise, it is determined that the target building does not have an additional floor, and the determination ends; Step S5: Use the weighted local spatial anomaly measurement method to calculate the height change consistency index between the target building and other buildings in the target area, set a second threshold, and if the height change consistency index exceeds the second threshold, it is finally determined that the target building has an additional floor, and the determination ends; otherwise, it is finally determined that the target building does not have an additional floor, and the determination ends.
2. The intelligent identification method for building floors based on multimodal remote sensing data according to claim 1 is characterized in that: Building remote sensing images include optical remote sensing images and radar remote sensing images, among which optical remote sensing images include high-resolution optical images and wide-area optical images.
3. The intelligent identification method for building floors based on multimodal remote sensing data according to claim 1 is characterized in that: The observation dataset needs to be preprocessed before being input into the deep learning prediction model. The preprocessing process includes geometric correction, orthorectification, radiation correction and image enhancement to ensure data quality. At the same time, the sliding window technology is used to crop the remote sensing image into the required standard size, and a zero-value filling strategy is adopted for the edge areas that cannot be divided evenly to ensure data integrity.
4. The intelligent identification method for building floors based on multimodal remote sensing data according to claim 2 is characterized in that: The deep learning prediction model includes a multimodal feature extraction module, a feature fusion module, a spatial context information enhancement module, and an output module. The multimodal feature extraction module uses a long- and short-attention mechanism to extract multi-scale building features from optical remote sensing images and radar remote sensing images respectively. feature The fusion module adopts an adaptive fusion attention mechanism to fuse building features of different modalities and scales to obtain fusion features; The spatial context information enhancement module captures the relationship between the building and the surrounding environment to enhance the fusion features; the output module is used to process the fusion features and output the predicted value of the building height.
5. The intelligent identification method for building floors based on multimodal remote sensing data according to claim 4 is characterized in that: Using fused dice loss in deep learning prediction model training , boundary loss and Lovas loss The composite loss function , expressed as: Where, 、 、 Represents dice loss respectively , boundary loss and Lovas loss The weight coefficient of ; where: Dice loss Expressed as: Where, Represents any pixel in a remote sensing image; Represents the total number of pixels in the remote sensing image; Represents the deep learning prediction model for pixels predicted probability; Represents pixels The true pixel value of Indicates the minimum value, used to prevent the denominator from being 0; Boundary loss Expressed as: Where, Represents pixels SDM value; Lovas loss Expressed as: Where, Represents a collection of pixels; Represents the sorted IoU loss gradient sequence; Represents any element in the IoU loss gradient sequence.
6. The intelligent identification method for building floors based on multimodal remote sensing data according to claim 1 is characterized in that: In step S4, the target area is determined by the following process: Step S41: Grid the remote sensing image and calculate the building density in each grid unit. The calculation process is expressed as follows: Where, represents the building density within the grid cell; Represents the total building area within the grid cell, calculated from the building outline data; represents the total area of the grid cells; Step S42: The target area is determined in each grid cell using a sliding window method. The sliding window size is expressed as: Where, Indicates the sliding window size; Indicates the base window size; represents the density adjustment coefficient; Step S43: Select multiple target areas one by one according to the preset window overlap rate and step size, where: Window overlap ratio Expressed as: Where, represents the baseline overlap ratio; represents the overlap adjustment coefficient; step length Expressed as: 。 7. The intelligent identification method for building floors based on multimodal remote sensing data according to claim 1 is characterized in that: The process of setting the first threshold specifically includes the following steps: If the height change values of all buildings in the target area obey the normal distribution, the first threshold Set to: Where, Represents the average value of all building height changes in the target area; The coefficient factor that represents the control sensitivity, with a value range of 2-3; Represents the standard deviation of all building height changes within the target area; represents the regional heterogeneity adjustment coefficient, with a value range of 0.1-0.5; represents the coefficient of variation, ; Indicates the total number of samples of buildings in the target area; Indicates the number of benchmark samples, with a value of 50-100; If the height variation of all buildings in the target area is severely skewed, the first threshold Use any of the following methods to set: First: Where, Represents the median of all building height changes within the target area; Represents the absolute median difference of all building height changes in the target area; Second: Where, It represents the 95th percentile of all building height changes in the target area. That is, if the height change of the target building exceeds the height change of 95% of the samples in the area, it is considered abnormal.
8. The intelligent identification method for building floors based on multimodal remote sensing data according to claim 1 is characterized in that: The height change consistency index is expressed by the spatial anomaly metric value, and its calculation process is expressed as follows: Where, Indicates the target building The spatial anomaly measure of Indicates the target building The height change value; Indicates buildings within the target area The height change value; Represents the spatial weight matrix, and the target building and adjacent buildings is inversely proportional to the distance; Indicates that the target building is included target area; Indicates the target area The standard deviation of the variation in height of all buildings within the .
9. The intelligent identification method for building floors based on multimodal remote sensing data according to claim 1, characterized in that: The value range of the second threshold is 2.5-3.5, where: for densely built areas, the second threshold is 3.0; for sparsely built areas, the second threshold is 2.5; for special supervision areas, the second threshold is 3.
5.
10. A system for executing the intelligent identification method for building floors based on multimodal remote sensing data according to any one of claims 1 to 9, characterized in that: include: An observation dataset, comprising remote sensing images of multiple buildings at different viewing angles and in different modalities, with the height of the corresponding building in each remote sensing image used as a label; A deep learning prediction model is trained using an observation dataset to learn the mapping relationship between remote sensing images and building heights. For any building floor recognition task, a time series containing multiple consecutive time nodes is selected. At each time node, a remote sensing image of the target building is collected and input into the trained deep learning prediction model. The predicted height value of the target building at the current time node is output, and the first-order difference between the predicted height value of the target building at the current time node and the predicted height value at the previous time node is calculated as the height change value of the target building. This is then traversed through all time nodes to form a target building height change sequence. A preliminary determination module is configured to select a target area, calculate a height change sequence of all buildings in the target area, and set a first threshold value at any time point based on the height change values of all buildings in the target area. If the height change value of the target building exceeds the first threshold value, it is preliminarily determined that the target building has an additional floor, and step S5 is executed; Otherwise, it is determined that the target building does not have an additional floor, and the determination ends; The final judgment module is used to calculate the height change consistency index between the target building and other buildings in the target area using a weighted local spatial anomaly measurement method, set a second threshold, and if the height change consistency index exceeds the second threshold, it is finally determined that the target building has an additional floor and the judgment ends; otherwise, it is finally determined that the target building does not have an additional floor and the judgment ends.
Citation Information
Patent Citations
Dynamic monitoring method and system for urban illegal buildings
CN110243354A
Illegal building detection method and device applied to city management supervision
CN114463624A
Method and device for training preset remote sensing image overlapping shadow segmentation model
CN114529832A
Urban building change remote sensing detection method based on twin multitask network
CN114821354A
Multi-mode building height change detection method for process decoupling
CN119169455A