Building change detection method based on unmanned aerial vehicle image and medium

By constructing a unified aerial triangulation solution system for UAV imagery and deep learning semantic segmentation, the problems of geometric registration deviation and semantic segmentation error in building change detection in UAV low-altitude remote sensing were solved, achieving high-precision and stable building change detection.

CN122329258BActive Publication Date: 2026-08-25Ningbo Institute of Surveying, Mapping and Remote Sensing Technology (Ningbo Natural Resources and Planning Survey and Monitoring Center)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610787244.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-25
Estimated Expiration
2046-06-03

AI Technical Summary

Technical Problem

Existing technologies for detecting building changes in low-altitude UAV remote sensing scenarios are affected by airflow fluctuations, viewing angle differences, and high-resolution parallax, leading to the superposition of geometric registration deviations and semantic segmentation errors, which affect detection accuracy and stability.

Method used

By acquiring pre- and post-terminal images from UAVs, and combining them with ground control points and aerial POS information, a unified aerial triangulation solution system is constructed. Orthophotos are generated and semantic segmentation is performed. Semantic probabilities are extracted using deep learning, cross-temporal consistent object units are constructed, and pixel-level semantic confidence mean judgment and vector reconstruction are performed to form building change detection results.

Benefits of technology

It achieves stable and accurate detection of building changes in UAV imagery, overcoming geometric registration bias and semantic segmentation errors in traditional methods, and improving detection accuracy and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122329258B_ABST
    Figure CN122329258B_ABST
Patent Text Reader

Abstract

The application relates to a building change detection method and medium based on unmanned aerial vehicle images, which comprises the following steps: based on the front and rear phase unmanned aerial vehicle images of the same target area, the unmanned aerial vehicle aerial survey ground control point and the aerial survey POS information are introduced into the aerial triangulation calculation system to jointly calculate the camera exterior orientation elements and the imaging geometric relationship, the front and rear phase orthographic images are generated under the unified coordinate reference, and the front and rear phase building probability graphs are obtained by processing, the merged rear orthographic image obtained by merging the front and rear phase unmanned aerial vehicle images is subjected to cross-phase consistency object unit construction to obtain a plurality of object units, first and second semantic confidence averages corresponding to the front and rear phase orthographic images are further obtained, and whether the two confidence averages meet three preset judgment conditions is judged to mark each object unit as a building change area. In this way, the stable and accurate detection of the building change condition based on the unmanned aerial vehicle images is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of surveying and mapping, and in particular to a method and medium for detecting building changes based on UAV imagery. Background Technology

[0002] Building change detection is a core foundational task in surveying and mapping operations such as monitoring illegal construction, planning verification, and urban renewal verification. Currently, the mainstream approach for building change detection in surveying and mapping typically employs a "post-classification comparison method." This involves first acquiring both pre- and post-contemporary images of the building, then performing semantic segmentation on these images, and finally identifying areas of new addition, demolition, or boundary changes by comparing the segmentation results. This building change detection method based on pre- and post-contemporary images benefits from the rigorous geometric correction, unified benchmarks, and system-level processing capabilities of satellite imagery, making it highly applicable in satellite imagery scenarios. It can achieve near-pixel-level consistent alignment over a wide area, providing a reliable geometric basis for pixel-level differential analysis.

[0003] However, the aforementioned building change detection method based on previous and subsequent temporal images has shortcomings: When applied to low-altitude UAV remote sensing scenarios, this method faces significant technical limitations. Specifically, UAVs are susceptible to fluctuations in flight attitude caused by airflow, slight differences in shooting angle, and changes in flight altitude. Furthermore, the significant parallax and local deformation brought about by ultra-high resolution make registration of previous and subsequent temporal images based solely on image texture features difficult, leading to geometric registration deviations and parallax interference. These geometric registration deviations are amplified in pixel-level differentiation, resulting in frequent band-like or fragmented "pseudo-changes" along building outlines, severely impacting detection accuracy. Moreover, even if the geometric registration of the previous and subsequent temporal images is relatively accurate, single-temporal semantic segmentation is still limited by factors such as blurred edges, similar textures, and complex roof structures, resulting in boundary offsets, missed detections, or false detections. This causes these independently occurring segmentation errors to be superimposed and amplified during the comparison of the previous and subsequent temporal images, leading to a decrease in the stability and verifiability of the building change detection results.

[0004] Therefore, how to accurately detect changes in buildings based on UAV imagery has become a pressing technical problem in the current surveying and mapping field. Summary of the Invention

[0005] The first technical problem to be solved by the present invention is to provide a method for detecting building changes based on UAV imagery, which is in contrast to the above-mentioned prior art.

[0006] The second technical problem to be solved by the present invention is to provide a readable storage medium for implementing the above-mentioned method for detecting building changes based on UAV imagery.

[0007] The technical solution adopted by this invention to solve the first technical problem is: a method for detecting building changes based on UAV imagery, characterized by comprising the following steps 1 to 8: Step 1: Acquire both the preceding and subsequent time-phase UAV images of the same target area; wherein, the preceding and subsequent time-phase UAV images are sets of multiple overlapping single images acquired by aerial photography of the same target area at different time phases. Step 2: Based on the acquired previous and subsequent UAV images, the ground control points and aerial POS information of the UAV aerial survey are incorporated into the aerial triangulation solution system to jointly calculate the camera exterior orientation elements and imaging geometric relationships of the camera carried by the UAV. Step 3: Generate the preceding orthorectified image of the preceding UAV image and the following orthorectified image of the following UAV image under a unified coordinate reference. Then, use semantic segmentation grid to process the preceding and following orthorectified images to output the preceding and following building probability maps. The preceding and following orthorectified images have a consistent spatial reference and pixel correspondence at the pixel level. Step 4: Merge the previous and subsequent UAV images to obtain a merged orthophoto. Then, construct cross-temporal consistent object units in a unified alignment space to obtain multiple cross-temporal consistent object units. Step 5: Perform spatial aggregation operation of pixel-level semantic probability and semantic confidence calculation on each obtained object unit to obtain the mean of the first semantic confidence of the corresponding previous temporal orthophoto and the mean of the second semantic confidence of the corresponding subsequent temporal orthophoto. Step 6: Make a judgment based on the obtained mean of the first semantic confidence score and the mean of the second semantic confidence score: When the mean of the first semantic confidence score and the mean of the second semantic confidence score meet the first preset judgment condition, the current object unit is determined to be a newly built area of ​​the building, and the newly built area is recorded as the building change area. When the mean of the first semantic confidence score and the mean of the second semantic confidence score meet the second preset judgment condition, the current object unit is determined to be the demolished area of ​​the building, and the demolished area is recorded as the building change area. When the mean of the first semantic confidence score and the mean of the second semantic confidence score meet the third preset judgment condition, the current object unit is determined to be an area of ​​uncertain building change, and the process proceeds to step 7; wherein, the first preset judgment condition is: the mean of the first semantic confidence score is less than or equal to a preset low confidence threshold, and the mean of the second semantic confidence score is greater than or equal to a preset high confidence threshold; the second preset judgment condition is: the mean of the first semantic confidence score is greater than or equal to a preset high confidence threshold, and the mean of the second semantic confidence score is less than or equal to a preset low confidence threshold; the third preset judgment condition is: the mean of the first semantic confidence score and the mean of the second semantic confidence score do not meet the first preset judgment condition and do not meet the second preset judgment condition; Step 7: Perform object aggregation and vectorization reconstruction processing on all object units in the uncertain area of ​​building change as determined by the judgment result to obtain vector polygons representing the building change. Step 8: Perform a detailed determination of expansion and shape changes on the obtained vectorized polygons, and finally form a change detection result for uncertain areas of building changes; wherein, the change detection result includes at least one of the following five types of building change detection results: newly built area, demolished area, expanded area, shape change area and unchanged area.

[0008] Improved, in the aforementioned building change detection method based on UAV imagery, in step 2: the ground control points and aerial POS information from UAV aerial surveys are incorporated into the aerial triangulation solution system, and the method for jointly calculating the camera exterior orientation elements and imaging geometric relationship of the camera carried by the UAV is as follows: Step a1: Set the objective function for minimizing the reprojection error with prior constraints; Step a2: Calculate the minimum function value of the objective function that minimizes the reprojection error; Step a3: The camera exterior orientation element corresponding to the minimum function value of the objective function for minimizing the reprojection error is taken as the corrected camera exterior orientation element, and the geometric connection relationship between the camera, image point and ground point determined by the corrected camera exterior orientation element is taken as the imaging geometric relationship.

[0009] Further improvements are made to the building change detection method based on UAV imagery. In step 3, the process of generating the preceding orthophoto image corresponding to the preceding UAV imagery and the following orthophoto image corresponding to the following UAV imagery under a unified coordinate reference includes the following steps: Step b1, Orthorectification and resampling: Based on the calculated exterior orientation elements of the camera and the imaging geometry, orthorectification is performed on the front-time UAV image and the back-time UAV image respectively under a unified coordinate reference. Step b2, generate aligned image pairs: use the orthorectified image of the preceding time-phase UAV image as the generated preceding time-phase orthorectified image, and use the orthorectified image of the following time-phase UAV image as the generated following time-phase orthorectified image.

[0010] Furthermore, in the building change detection method based on UAV imagery, step 3, which processes the preceding and following temporal orthophotos using semantic segmentation grids to output the preceding and following temporal building probability maps, includes the following steps c1 to c4: Step c1, construct a building sample library: collect UAV remote sensing images containing various building types, textures and lighting conditions, and perform pixel-level building mask annotation on the collected UAV remote sensing images to construct a sample dataset for training. Step c2, train the semantic segmentation network: select a deep convolutional neural network, and use the constructed sample dataset to supervise the training of the deep convolutional neural network to train a semantic segmentation model that can recognize the semantic features of buildings in UAV remote sensing images. Step c3, perform probabilistic inference: input the previous phase orthophoto and the subsequent phase orthophoto into the trained semantic segmentation model respectively, and perform forward inference calculation to obtain the original confidence score of each pixel in each orthophoto belonging to the building category. Step c4, output semantic probability map: normalize all the original confidence scores obtained to the score interval [0,1], generate pixel-level pre-temporal building probability map and post-temporal building probability map, and retain the continuous probability distribution features in the output process without performing binarization processing.

[0011] Improved, in the aforementioned building change detection method based on UAV imagery, the method is characterized in that, in step 4, the process of constructing cross-temporal consistent object units from the merged orthophoto image within a unified alignment space to obtain multiple cross-temporal consistent object units includes the following steps 41-42: Step 41, Construct dual-temporal joint features: Use the previous temporal orthophoto and the subsequent temporal orthophoto as a three-dimensional tensor. If any orthophoto contains C Each feature channel will perform a sequential stitching-based concatenation operation on the preceding and following temporal orthophotos while maintaining the image height and width. Step 42: Perform superpixel segmentation based on joint features to generate object units, so as to divide the previous temporal orthophoto and the subsequent temporal orthophoto into multiple initial seed points respectively, and determine the object unit boundary by iteratively optimizing the cluster center. The process of determining the boundaries of object units by iteratively optimizing cluster centers includes an iterative pixel allocation process for assigning pixels to their respective domains; wherein one pixel allocation iteration process includes steps 42a~42e: Step 42a: For any pixel in each orthophoto, initialize a minimum distance label and a category label for that pixel; Step 42b: Traverse all cluster centers and define a local search window around each cluster center; Step 42c: Make a judgment based on the combined distance between any pixel in the obtained orthophoto and the cluster center and the minimum distance label of that pixel: When the overall distance corresponding to any pixel is less than its minimum distance label, update the overall distance to the minimum distance label of that pixel, and update the category label L of that pixel to... k Otherwise, the minimum distance label and its category label for any pixel will not be updated, and the original distance record and category status of any pixel will be retained. Step 42d: After traversing all cluster centers and completing the local search and comprehensive distance update for each cluster center, all pixels are assigned to their respective cluster centers, forming a preliminary object unit partition. After multiple iterations until the cluster center positions no longer change, the multiple object units with cross-temporal consistency are finally output. Step 42e: Calculate the arithmetic mean of the joint feature vector of all pixels in the object unit partition and the arithmetic mean of the spatial coordinates of all pixels, and use the calculated arithmetic mean of the joint feature vector and the arithmetic mean of the pixel spatial coordinates to update the joint feature components and spatial location components of the cluster center respectively.

[0012] Furthermore, in the building change detection method based on UAV imagery, in step 5, the first semantic confidence mean and the second semantic confidence mean are calculated as follows: ; ; in, The mean of the first semantic confidence score. The pixel-level semantic probability is the number of objects within the object unit of the previous temporal orthophoto. This is the pre-temporal building probability map output in step 3; The mean of the second semantic confidence score. The pixel-level semantic probability is the number of objects within the later-phase orthophoto. This is the post-temporal building probability map output in step 3.

[0013] Further improvements are made to the building change detection method based on UAV imagery. In step 7, the process of performing object aggregation and vectorization reconstruction on all object units in areas where the building change is uncertain includes steps e1 to e4: Step e1 involves performing a connected component merging operation on all spatially adjacent object units with consistent building change region determination results to form a complete building change patch with an independent topological structure; wherein, the complete building change patch corresponds to the building change region determination result; Step e2: Perform boundary tracking and vectorization reconstruction on each complete building change patch to generate the corresponding closed vector polygon and the geometric outline of the closed vector polygon respectively; Step e3: Calculate the geometric area of ​​each object unit based on the obtained closed vector polygons and their corresponding geometric contours. Step e4: Perform patch filtering based on the geometric area of ​​each object unit and the preset minimum patch area threshold. If the geometric area of ​​any object unit is less than the minimum area threshold of the top map patch, the complete building change patch corresponding to that object unit is removed; otherwise, the complete building change patch corresponding to that object unit is retained.

[0014] In a further improvement, in the aforementioned building change detection method based on UAV imagery, the refined determination method for expanding and changing the shape of the obtained vectorized polygon in step 8 is as follows: Step f1: Calculate the absolute area increment of the later phase orthophoto relative to the earlier phase orthophoto. Step f2: Make a judgment based on the obtained absolute area increment value and the preset absolute area increment threshold. When the absolute area increment value is greater than the preset absolute area increment threshold, the current object unit is determined to be a building expansion area; otherwise, proceed to step f3. Step f3: Calculate the Euclidean distance from each point on the geometric contour line corresponding to the previous temporal orthophoto to the geometric contour line corresponding to the subsequent temporal orthophoto. Step f4: Calculate the mean distance of all obtained Euclidean distance values, and use the mean distance as a measure of the geometric difference between the previous phase orthophoto and the subsequent phase orthophoto. Step f5: Make a judgment based on the obtained geometric difference measure value: When the geometric difference metric value is greater than the preset geometric difference metric threshold, the current object unit is determined to be a building shape change area; otherwise, the current object unit is determined to be a building unchanged area.

[0015] The technical solution adopted by the present invention to solve the second technical problem is: a readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements any of the building change detection methods based on UAV imagery described in the present invention.

[0016] Compared with the prior art, the advantages of the present invention are as follows: First, the building change detection method based on UAV imagery of this invention incorporates ground control points and aerial POS information from UAV aerial surveys into an aerial triangulation system based on previously acquired and subsequently acquired UAV images of the same target area. This allows for the joint calculation of the camera exterior orientation elements and imaging geometry of the UAV-carried camera. Corresponding previously acquired and subsequently acquired orthorectified images are then generated under a unified coordinate reference. Finally, semantic segmentation grids are used to process the previously acquired and subsequently acquired orthorectified images to obtain the building probability of the previous time period. The method involves combining the preceding and following time-phase building probability maps. The merged orthophoto obtained from the combined UAV imagery of the preceding and following time phases is then processed within a unified alignment space to construct multiple cross-temporal consistent object units. After processing each object unit, the mean first semantic confidence score of the corresponding preceding time-phase orthophoto and the mean second semantic confidence score of the corresponding following time-phase orthophoto are obtained. Based on whether the two mean confidence scores meet the first, second, and third preset judgment conditions, the method determines and labels the current object unit as a change region of the building. This achieves stable and accurate detection of building changes based on UAV imagery.

[0017] Secondly, the building change detection method based on UAV imagery of this invention deeply couples the ground control points of UAV aerial surveys and the POS information of aerial photography into the aerial triangulation and orthorectification system, actively constructing a unified geometric benchmark for sharing dual-temporal images. This overcomes the inherent alignment instability of traditional methods that rely solely on image texture features for passive matching from a physical source, laying a rigorous geometric foundation for high-precision change detection. Furthermore, based on the aforementioned unified geometric benchmark, this invention innovatively constructs cross-temporal consistent object units, strictly limiting the change reasoning process to a unified spatial partition. This mandatory spatial consistency constraint effectively blocks the error propagation and amplification link of single-phase semantic segmentation boundary differences and local misclassifications in the pixel-level difference stage, significantly improving the robustness of the algorithm to image registration and segmentation noise.

[0018] Finally, the building change detection method based on UAV imagery of this invention, under the premise of establishing a unified geometric benchmark to establish high-precision alignment of dual-temporal images, utilizes deep learning to extract semantic probabilities and introduces cross-temporal consistent object units as core constraints, thereby improving the detection dimension from pixel level to object level. Through object-level evidence chains, the superposition interference of registration and segmentation errors is effectively suppressed, significantly enhancing the stability of building change detection results. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the building change detection method based on UAV imagery in an embodiment of the present invention. Detailed Implementation

[0020] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0021] This embodiment provides a method for detecting building changes based on UAV imagery. Specifically, see [link to documentation]. Figure 1 As shown, the building change detection method based on UAV imagery in this embodiment includes the following steps 1-8: Step 1: Acquire both the preceding and subsequent time-phase UAV images of the same target area; wherein, the preceding and subsequent time-phase UAV images are sets of multiple overlapping single images acquired by aerial photography of the same target area at different time phases. Step 2: Based on the acquired previous and subsequent UAV images, the ground control points and aerial POS information of the UAV aerial survey are incorporated into the aerial triangulation solution system to jointly calculate the camera exterior orientation elements and imaging geometric relationships of the camera carried by the UAV. Step 3: Generate the preceding orthorectified image of the preceding UAV image and the following orthorectified image of the following UAV image under a unified coordinate reference. Then, use semantic segmentation grid to process the preceding and following orthorectified images to output the preceding and following building probability maps. The preceding and following orthorectified images have a consistent spatial reference and pixel correspondence at the pixel level. Steps 2 and 3 mentioned above are mainly used to actively construct high-precision dual-temporal geometric consistency and avoid alignment instability caused by relying solely on image texture matching; while this embodiment of the invention aims to provide uncertainty characterization basis for subsequent inference and eliminate boundary jitter interference caused by hard threshold segmentation by retaining continuous probability for object-level statistics. Step 4: Merge the previous and subsequent UAV images to obtain a merged orthophoto. Then, construct cross-temporal consistent object units from the merged orthophoto within a unified alignment space to obtain multiple cross-temporal consistent object units. Secondly, by constructing cross-temporal consistent object units from the merged orthophoto within a unified alignment space, the error amplification effect of minor differences in the two-temporal segmentation boundary and local geometric inconsistencies in the difference operation can be effectively suppressed. Step 5: Perform spatial aggregation operation of pixel-level semantic probability and semantic confidence calculation on each obtained object unit to obtain the mean of the first semantic confidence of the corresponding previous temporal orthophoto and the mean of the second semantic confidence of the corresponding subsequent temporal orthophoto. Step 6: Make a judgment based on the obtained mean of the first semantic confidence score and the mean of the second semantic confidence score: When the mean of the first semantic confidence score and the mean of the second semantic confidence score meet the first preset judgment condition, the current object unit is determined to be a newly built area of ​​the building, and the newly built area is recorded as the building change area. When the mean of the first semantic confidence score and the mean of the second semantic confidence score meet the second preset judgment condition, the current object unit is determined to be the demolished area of ​​the building, and the demolished area is recorded as the building change area. When the mean of the first semantic confidence score and the mean of the second semantic confidence score meet the third preset judgment condition, the current object unit is determined to be an area of ​​uncertain building change, and the process proceeds to step 7; wherein, the first preset judgment condition is: the mean of the first semantic confidence score is less than or equal to a preset low confidence threshold τ1, and the mean of the second semantic confidence score is greater than or equal to a preset high confidence threshold τ2; the second preset judgment condition is: the mean of the first semantic confidence score is greater than or equal to a preset high confidence threshold τ2, and the mean of the second semantic confidence score is less than or equal to a preset low confidence threshold τ1; the third preset judgment condition is: the mean of the first semantic confidence score and the mean of the second semantic confidence score do not meet the first preset judgment condition and do not meet the second preset judgment condition; τ2>τ1; Step 7: Perform object aggregation and vectorization reconstruction processing on all object units in the uncertain area of ​​building change as determined by the judgment result to obtain vector polygons representing the building change. Step 8: Perform a detailed determination of expansion and shape changes on the obtained vectorized polygons, and finally form a change detection result for uncertain areas of building changes; the change detection result includes five types of building change detection results: newly built areas, demolished areas, expanded areas, areas with shape changes, and areas without changes.

[0022] Specifically, in step 2 of this embodiment, the method of incorporating the ground control points and aerial POS information of the UAV aerial survey into the aerial triangulation solution system, and jointly calculating the camera exterior orientation elements and imaging geometric relationship of the camera carried by the UAV, includes the following steps a1~a3: Step a1: Set the objective function for minimizing the reprojection error with prior constraints; wherein the objective function for minimizing the reprojection error is set as follows: ; l g >0, l p >0; in, P i This represents the first [unclear] in any acquired temporal UAV image. i The camera exterior orientation elements corresponding to each image, 1≤ i ≤ M , M This represents the total number of individual images contained in any given time-phase UAV imagery. X j Indicates the first j One ground control point, u ij Indicates the first j The ground control point at the i Image points on the image; π(·) represents ground control points. X j Based on camera exterior elements P i Collinearity equations and projection functions projected onto a two-dimensional image; Indicates ground control points X j The prior values ​​of the coordinates; This represents the external orientation elements of the camera recorded in the aerial photography POS information. P i The prior value; l g This represents the first weighting coefficient for adjusting the ground control point constraint term. l p This represents the second weighting coefficient for adjusting the aerial photography POS information constraint term; Represents the square of the Euclidean distance; for example, This indicates the solution for ground control points. X j and The Euclidean distance; Representing image points u ij Compared to the actual points seen on a single drone image (P) i ,X j The geometric distance between them; Indicates ground control point constraints; Indicates constraints on aerial photography POS information; Step a2: Calculate the minimum function value of the objective function that minimizes the reprojection error; Step a3: The camera exterior orientation element corresponding to the minimum function value of the objective function for minimizing the reprojection error is taken as the corrected camera exterior orientation element, and the geometric connection relationship between the camera, image point and ground point determined by the corrected camera exterior orientation element is taken as the imaging geometric relationship.

[0023] Specifically, in step 3 of this embodiment, the process of generating the preceding orthophoto of the preceding time-phase UAV image and the following orthophoto of the following time-phase UAV image under a unified coordinate reference includes the following steps b1~b2: Step b1, Orthorectification and resampling: Based on the calculated camera exterior orientation elements and the aforementioned imaging geometry, orthorectification is performed on the preceding and following time-phase UAV images under a unified coordinate reference. Step b2, generate aligned image pairs: the orthorectified image of the preceding time-phase UAV image is used as the generated preceding time-phase orthorectified image, and the orthorectified image of the following time-phase UAV image is used as the generated following time-phase orthorectified image.

[0024] Additionally, it should be noted that, specifically in step 3 of this embodiment, the process of using semantic segmentation grids to process the preceding and following temporal orthophotos to output the preceding and following temporal building probability maps includes the following steps c1 to c4: Step c1, construct a building sample library: collect UAV remote sensing images containing various building types, textures and lighting conditions, and perform pixel-level building mask annotation on the collected UAV remote sensing images to construct a sample dataset for training. Step c2, train the semantic segmentation network: select a deep convolutional neural network, and use the constructed sample dataset to supervise the training of the deep convolutional neural network to train a semantic segmentation model that can recognize the semantic features of buildings in UAV remote sensing images. Step c3, perform probabilistic inference: Input the previous phase orthophoto and the subsequent phase orthophoto into the trained semantic segmentation model respectively, and perform forward inference calculation to obtain the original confidence score of each pixel in each orthophoto belonging to the building category; the forward inference calculation here is a conventional technique and will not be described in detail here. Step c4, output semantic probability map: normalize all the original confidence scores obtained to the score interval [0,1], generate pixel-level pre-temporal building probability map and post-temporal building probability map, and retain the continuous probability distribution features in the output process without performing binarization processing.

[0025] Of course, specifically in step 4 of the building change detection method in this embodiment, the aforementioned process of constructing cross-temporal consistent object units from the merged orthophoto image within a unified alignment space to obtain multiple cross-temporal consistent object units includes the following steps 41-42: Step 41, construct the dual-temporal joint feature; wherein, the dual-temporal joint feature is constructed by using the previous temporal orthophoto and the subsequent temporal orthophoto as a three-dimensional tensor, if any orthophoto contains C Each feature channel will perform a sequential stitching concatenation operation on the preceding and following orthophotos while maintaining the image height and width; the input of this concatenation operation is the first tensor T1 corresponding to the preceding orthophoto and the second tensor T2 corresponding to the following orthophoto, and the output is the joint feature tensor F. T 1∈ R H×W×C , T 2∈ R H×W×C For any pixel position in the joint feature tensor F x Its eigenvector f(x) is derived from the previous temporal orthophoto at that location. x All channel components and the post-temporal orthophoto at this location x A 2C-dimensional column vector consisting of all channel components connected end-to-end in sequence: ;in, This represents the value of the first channel of the previous phase orthophoto image. The first phase of the orthophoto image represents the previous phase. C One channel value; This represents the value of the first channel of the post-temporal orthophoto. This represents the C-th channel value of the later-phase orthophoto image. For example, the earlier-phase orthophoto image is an RGB three-band image, i.e. C =3; For any geographic coordinate point, there is exactly one corresponding pixel coordinate point in both the previous and subsequent temporal orthophotos. x , x =( u , v ); u Represents the x-axis, v This represents the ordinate; at this point, the joint feature vector of each pixel not only contains the texture information of the current time, but also the texture information corresponding to another time. If the number of feature channels in the preceding orthophoto and the following orthophoto are different, then the feature channels common to both the preceding and following orthophotos are extracted and concatenated; or, the number of channels in the preceding and following orthophotos are respectively denoted as... C 1 and C 2. The total number of channels in the constructed joint feature tensor F is C 1+ C 2; Step 42: Perform superpixel segmentation based on joint features to generate object units, dividing the previous and subsequent temporal orthophotos into multiple initial seed points, and determining the object unit boundaries by iteratively optimizing the cluster centers; wherein, the generated object units are labeled as o. k The number of initial seed points obtained from the partitioning is K During the iterative optimization of cluster centers, any pixel in each orthophoto... r With cluster center k The combined distance between them is denoted as D ( r , k ): ; ; ; ; in, d feat ( r , k () represents any pixel in the orthophoto. r With cluster center k The joint temporal feature distance between them m d is the compactness parameter used to adjust the weights of feature similarity and spatial compactness; spatial ( r , k () represents any pixel in the orthophoto. r With cluster center k The geometric spatial distance between them is used to constrain the generated object units to maintain a regular and compact shape; S The preset superpixel grid step size, N This represents the total number of pixels in the orthophoto. K f is the total number of superpixels in the joint feature tensor F. r Represents pixels r The joint eigenvector of the two temporal phases, f k Representing cluster centers k The joint eigenvectors of the two temporal phases, This indicates the calculation of Euclidean distance; Represents pixels r The first of the two-phase joint eigenvectors c Each channel value Representing cluster centers k The first of the two-phase joint eigenvectors c Each channel value x r Represents pixels r spatial coordinate vector, x k Representing cluster centers k The spatial coordinate vector; u r Represents pixels r The horizontal coordinate on the orthophoto image v r Represents pixels r The vertical coordinate on the orthophoto u k Representing cluster centers k The horizontal coordinate on the orthophoto image v k Representing cluster centers k The vertical coordinate on the orthophoto.

[0026] By calculating the aforementioned dual-temporal joint feature distance d feat ( r , k This forces the segmentation boundary to form at the gradient edge shared by the two phases of the image, ensuring the spatial consistency of atomic units across both temporal phases; and by calculating the aforementioned geometric spatial distance d spatial ( r , k ), which is used to constrain the generated object units to maintain a compact and regular shape.

[0027] Furthermore, in step 42 mentioned above, the iterative process of determining the object unit boundary through iterative optimization of cluster centers includes the following steps 42a~42e: Step 42a: For any pixel in each orthophoto, initialize a minimum distance label and a category label for that pixel; wherein, the minimum distance label for that pixel is denoted as d. min The category label of any given pixel is marked as L; Step 42b: Traverse all cluster centers, defining a local search window around each cluster center; wherein the size of the local search window is 2. S ×2 S ; Step 42c: Make a judgment based on the combined distance between any pixel in the obtained orthophoto and the cluster center and the minimum distance label of that pixel: When the overall distance corresponding to any pixel is less than its minimum distance label, update the overall distance to the minimum distance label of that pixel, and update the category label L of that pixel to... k Otherwise, the minimum distance label and its category label for any pixel will not be updated, and the original distance record and category status of any pixel will be retained. Step 42d: After traversing all cluster centers and completing the local search and overall distance update for each cluster center's local search window, all pixels are assigned to their respective cluster centers, forming preliminary object units; whereby the object units are labeled as... o k ; Step 42e, calculate the units belonging to the object. o k The arithmetic mean of the joint feature vectors of all pixels within the range And will use the arithmetic mean As the updated cluster center k The joint feature components; and, calculating the elements belonging to the object unit. o k Arithmetic mean of the spatial coordinate vectors of all pixels within the area and the arithmetic mean As the updated cluster center k The spatial location components. The update formula is shown below: ; ; Among them, | o k | Represents an object unit o k The total number of pixels contained within, f r Represents pixels r The joint eigenvectors of the two temporal phases, x r Represents pixels r The spatial coordinate vector. After superpixel segmentation, atomic units with strict consistency in spatial coverage across both temporal phases are obtained, thereby ensuring that subsequent feature statistics and change discrimination are constrained to be completed within the same spatial partitioning reference, eliminating interpretation interference caused by spatial misalignment.

[0028] Specifically, in step 5 of this embodiment, the first semantic confidence mean and the second semantic confidence mean are calculated as follows: ; ; in, The mean of the first semantic confidence score. The pixel-level semantic probability is the number of objects within the object unit of the previous temporal orthophoto. This is the pre-temporal building probability map output in step 3; The mean of the second semantic confidence score. The pixel-level semantic probability is the number of objects within the later-phase orthophoto. This is the post-temporal building probability map output in step 3.

[0029] Furthermore, specifically regarding step 7 of this embodiment, the process of performing object aggregation and vectorization reconstruction on all object units in the area where the determination result is an uncertain change in the building includes steps e1 to e4: Step e1 involves performing a connected component merging operation on all spatially adjacent object units with consistent building change region determination results to form a complete building change patch with an independent topological structure; wherein, the complete building change patch corresponds to the building change region determination result; Step e2 involves performing boundary tracing and vectorization reconstruction on each formed complete building change patch, generating corresponding closed vector polygons and their corresponding geometric contour lines; wherein, the closed vector polygons are marked as... V Closed vector polygon V The corresponding geometric contour line is marked as ; Step e3: Based on the obtained closed vector polygons and their corresponding geometric contours, calculate the geometric area of ​​each object unit; wherein, the geometric area of ​​any object unit is denoted as... A , A = Area ( V ); Step e4: Perform patch filtering based on the geometric area of ​​each object unit and the preset minimum patch area threshold. When the geometric area of ​​any object unit is less than the minimum area threshold of the top image patch, the complete building change patch corresponding to that object unit is removed to suppress pseudo-changes caused by high-frequency noise or non-substantial texture fluctuations in the image, ensuring that the final output vector result has rigorous geometric quality and engineering usability; otherwise, the complete building change patch corresponding to that object unit is retained.

[0030] Specifically, in step 8 of this embodiment, the refined determination method for expanding and transforming the obtained vectorized polygon includes steps f1 to f5: Step f1: Calculate the absolute area increment of the later-phase orthophoto relative to the earlier-phase orthophoto; where the area of ​​the earlier-phase orthophoto is denoted as... A T1 The area of ​​the later phase orthophoto is marked as A T2 The absolute area increment value is marked as A abs ; A abs = A T2 - A T1 ; Step f2: Make a judgment based on the obtained absolute area increment value and the preset absolute area increment threshold. When the absolute area increment value is greater than the preset absolute area increment threshold, that is, the absolute area increment value A abs Preset absolute area increment threshold A min If the current object is a building expansion area, determine that the current object unit is a building expansion area; otherwise, proceed to step f3. Step f3: Calculate the Euclidean distance from each point on the geometric contour line corresponding to the preceding orthophoto to the geometric contour line corresponding to the following orthophoto; wherein, the geometric contour line corresponding to the preceding orthophoto is marked as... Geometric outline Any point on is marked as p The geometric contour lines corresponding to the later phase orthophoto are marked as This point p The geometric contour line corresponding to the later phase orthophoto The Euclidean distance value is marked as ; Step f4: Calculate the mean distance of all obtained Euclidean distance values, and use this mean distance as the geometric difference measure between the previous and subsequent temporal orthophotos; wherein, the geometric difference measure is denoted as... ; ; Step f5: Make a judgment based on the obtained geometric difference metric value: when the geometric difference metric value is greater than the preset geometric difference metric threshold, i.e., the geometric difference metric value is... Preset geometric difference measurement threshold τ TH If the current object unit is determined to be a building shape change area, then the current object unit is determined to be a building unchanged area.

[0031] Of course, this embodiment also provides a readable storage medium. Specifically, the readable storage medium stores a computer program, which, when executed by a processor, implements the above-described method for detecting building changes based on UAV imagery.

[0032] Although preferred embodiments of the present invention have been described in detail above, it should be clearly understood that various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for detecting building changes based on UAV imagery, characterized in that, Includes the following steps 1-8: Step 1: Acquire both front-time and back-time UAV images of the same target area. Step 2: Based on the acquired previous and subsequent UAV images, the ground control points and aerial POS information of the UAV aerial survey are incorporated into the aerial triangulation solution system to jointly calculate the camera exterior orientation elements and imaging geometric relationships of the camera carried by the UAV. Step 3: Generate the orthorectified image of the preceding time phase UAV image and the orthorectified image of the following time phase UAV image under a unified coordinate reference. Then, use semantic segmentation grid to process the preceding time phase orthorectified image and the following time phase orthorectified image respectively, and output the preceding time phase building probability map and the following time phase building probability map. Step 4: Merge the previous and subsequent UAV images to obtain a merged orthophoto. Then, construct cross-temporal consistent object units in a unified alignment space to obtain multiple cross-temporal consistent object units. Step 5: Perform spatial aggregation operation of pixel-level semantic probability and semantic confidence calculation on each obtained object unit to obtain the mean of the first semantic confidence of the corresponding previous temporal orthophoto and the mean of the second semantic confidence of the corresponding subsequent temporal orthophoto. Step 6: Make a judgment based on the obtained mean of the first semantic confidence score and the mean of the second semantic confidence score: When the mean of the first semantic confidence score and the mean of the second semantic confidence score meet the first preset judgment condition, the current object unit is determined to be a newly built area of ​​the building, and the newly built area is recorded as the building change area. When the mean of the first semantic confidence score and the mean of the second semantic confidence score meet the second preset judgment condition, the current object unit is determined to be the demolished area of ​​the building, and the demolished area is recorded as the building change area. When the mean of the first semantic confidence score and the mean of the second semantic confidence score meet the third preset judgment condition, the current object unit is determined to be an area of ​​uncertain building change, and the process proceeds to step 7; wherein, the first preset judgment condition is: the mean of the first semantic confidence score is less than or equal to a preset low confidence threshold, and the mean of the second semantic confidence score is greater than or equal to a preset high confidence threshold; the second preset judgment condition is: the mean of the first semantic confidence score is greater than or equal to a preset high confidence threshold, and the mean of the second semantic confidence score is less than or equal to a preset low confidence threshold; the third preset judgment condition is: the mean of the first semantic confidence score and the mean of the second semantic confidence score do not meet the first preset judgment condition and do not meet the second preset judgment condition; Step 7: Perform object aggregation and vectorization reconstruction processing on all object units in the uncertain area of ​​building change as determined by the judgment result to obtain vector polygons representing the building change. Step 8: Expand and finely determine the shape changes of the obtained vectorized polygons to finally form the change detection results for uncertain areas of building changes.

2. The method for detecting building changes based on UAV imagery according to claim 1, characterized in that, In step 1, the preceding time-phase UAV image and the following time-phase UAV image are sets of multiple overlapping single images acquired by aerial photography of the same target area at different time phases. In step 2: The ground control points and aerial POS information from UAV aerial surveys are incorporated into the aerial triangulation solution system. The method for jointly calculating the exterior orientation elements and imaging geometric relationship of the camera carried by the UAV is as follows: Step a1: Set the objective function for minimizing the reprojection error with prior constraints; Step a2: Calculate the minimum function value of the objective function that minimizes the reprojection error; Step a3: The camera exterior orientation element corresponding to the minimum function value of the objective function for minimizing the reprojection error is used as the corrected camera exterior orientation element, and the geometric connection relationship between the camera, image point and ground point determined by the corrected camera exterior orientation element is used as the imaging geometric relationship. In step 3, the preceding orthophoto and the following orthophoto have a consistent spatial reference and pixel correspondence at the pixel level.

3. The method for detecting building changes based on UAV imagery according to claim 2, characterized in that, In step 3, the process of generating the preceding orthophoto of the preceding time-phase UAV image and the following orthophoto of the following time-phase UAV image under a unified coordinate reference includes the following steps: Step b1, Orthorectification and resampling: Based on the calculated exterior orientation elements of the camera and the imaging geometry, orthorectification is performed on the front-time UAV image and the back-time UAV image respectively under a unified coordinate reference. Step b2, generating aligned image pairs: using a single global image obtained by orthorectifying and stitching the preceding time-phase UAV image as the generated preceding time-phase orthorectified image, and using a single global image obtained by orthorectifying and stitching the following time-phase UAV image as the generated following time-phase orthorectified image.

4. The method for detecting building changes based on UAV imagery according to claim 3, characterized in that, In step 3, the semantic segmentation grid is used to process the preceding and following temporal orthophotos respectively, and the process of outputting the preceding and following temporal building probability maps includes the following steps c1 to c4: Step c1, construct a building sample library: collect UAV remote sensing images containing various building types, textures and lighting conditions, and perform pixel-level building mask annotation on the collected UAV remote sensing images to construct a sample dataset for training. Step c2, train the semantic segmentation network: select a deep convolutional neural network, and use the constructed sample dataset to supervise the training of the deep convolutional neural network to train a semantic segmentation model that can recognize the semantic features of buildings in UAV remote sensing images. Step c3, perform probabilistic inference: input the previous phase orthophoto and the subsequent phase orthophoto into the trained semantic segmentation model respectively, and perform forward inference calculation to obtain the original confidence score of each pixel in each orthophoto belonging to the building category. Step c4, output semantic probability map: normalize all the original confidence scores obtained to the score interval [0,1], generate pixel-level pre-temporal building probability map and post-temporal building probability map, and retain the continuous probability distribution features in the output process without performing binarization processing.

5. The building change detection method based on UAV imagery according to claim 4, characterized in that, In step 4, the merged orthophoto image is processed to construct cross-temporal consistent object units within a unified alignment space, resulting in multiple cross-temporal consistent object units. This process includes the following steps 41-42: Step 41, construct the dual-temporal joint features; Step 42: Perform superpixel segmentation based on joint features to generate object units, so as to divide the previous and subsequent temporal orthophotos into multiple initial seed points, and determine the object unit boundary by iteratively optimizing the cluster center.

6. The method for detecting building changes based on UAV imagery according to claim 5, characterized in that, In step 8, the change detection results include five categories of building change detection results: newly built areas, demolished areas, expanded areas, areas with morphological changes, and areas without changes. Furthermore, in step 8, the precise method for determining the expansion and shape transformation of the obtained vectorized polygon is as follows: Step f1: Calculate the absolute area increment of the later phase orthophoto relative to the earlier phase orthophoto. Step f2: Make a judgment based on the obtained absolute area increment value and the preset absolute area increment threshold. When the absolute area increment value is greater than the preset absolute area increment threshold, the current object unit is determined to be a building expansion area; Otherwise, proceed to step f3; Step f3: Calculate the Euclidean distance from each point on the geometric contour line corresponding to the previous temporal orthophoto to the geometric contour line corresponding to the subsequent temporal orthophoto. Step f4: Calculate the mean distance of all obtained Euclidean distance values, and use the mean distance as a measure of the geometric difference between the previous phase orthophoto and the subsequent phase orthophoto. Step f5: Make a judgment based on the obtained geometric difference measure value: When the geometric difference metric value is greater than the preset geometric difference metric threshold, the current object unit is determined to be a building shape change area; otherwise, the current object unit is determined to be a building unchanged area.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the building change detection method based on UAV imagery as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Dam slope deformation monitoring system and method

    CN110453731A

  • Unmanned aerial vehicle photogrammetry registration method and system and computer equipment

    CN117670957A