A building update type identification method and device based on time-series street view images

By acquiring street view images from different periods, and utilizing a building update type recognition model based on YOLO and an improved Mask2Former architecture, the problems of insufficient temporal and multi-view fusion in traditional methods are solved, achieving efficient and accurate recognition and management of building update types.

CN120913084BActive Publication Date: 2025-12-30UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511443391.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-12-30
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Traditional building renovation type identification methods lack temporal and multi-view fusion capabilities, making it impossible to reliably identify changes in the appearance of buildings at different times and from different shooting angles. Furthermore, manual surveys are inefficient and costly, and remote sensing images have limited resolution, making it difficult to identify details of building facades.

Method used

By acquiring street view images from different periods and perspectives, building targets are detected using the YOLO object detection model. A geographic information database is constructed by combining geographic coordinates and object re-identification algorithms. An improved Mask2Former architecture building update type recognition model is also built to achieve accurate identification of building update types.

Benefits of technology

It achieves stable identification of buildings at different times and from different perspectives, can accurately distinguish micro-update types such as structural reinforcement and exterior decoration, meets the needs of large-scale and high-frequency updates, improves identification efficiency and accuracy, and reduces human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913084B_ABST
    Figure CN120913084B_ABST
Patent Text Reader

Abstract

The application discloses a building update type identification method and device based on time sequence street view images, and relates to the technical field of civil engineering and computer vision. The method comprises the following steps: acquiring multiple street view images; outputting a target detection result of a building through a YOLO target detection model; determining geographical coordinates of the building according to metadata information of the street view images and the target detection result; constructing an initial building group geographical information database according to the geographical coordinates and in combination with a target re-identification algorithm; constructing a building update type identification data set according to the initial building group geographical information database; constructing a building update type identification model; training the building update type identification model by using the building update type identification data set; outputting an update type identification result of the building through the trained building update type identification model; and storing the update type identification result into the initial building group geographical information database to obtain a target building group geographical information database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of civil engineering and computer vision technology, and in particular to a method and device for identifying building update types based on time-series street view images. Background Technology

[0002] With the accelerating pace of urbanization, the number and density of urban building complexes continue to rise, leading to increasingly frequent demolition, renovation, reinforcement, and reconstruction of buildings. Accurately understanding the status of building renovations is of significant practical importance for urban planning, building safety supervision, disaster risk assessment, and the preservation of historic districts.

[0003] In existing technologies, traditional building renewal type identification methods are based on single-time point images or single-view information, lacking temporal and multi-view fusion capabilities, resulting in the inability to reliably identify changes in the appearance of buildings at different times and from different shooting angles.

[0004] In addition, traditional methods for identifying building renewal types mainly rely on manual surveys or remote sensing image change detection.

[0005] However, manual surveys are inefficient, costly, and slow to update, making it difficult to meet the needs of large-scale, high-frequency updates. While change detection based on remote sensing images can achieve large-scale coverage, its limited image resolution makes it difficult to identify details of building facades. It can usually only determine whether a building has disappeared or been added, and cannot accurately distinguish micro-update types such as structural reinforcement and exterior decoration. Summary of the Invention

[0006] To address the shortcomings of existing technologies, traditional building renewal type identification methods rely on single-time-point images or single-view information, lacking temporal and multi-view fusion capabilities. This results in an inability to reliably identify changes in building appearance at different times and from different shooting angles. Furthermore, manual surveys are inefficient, costly, and prone to delays, failing to meet the demands of large-scale, high-frequency updates. Change detection based on remote sensing images has limited image resolution, making it difficult to identify building facade details; it typically only determines whether a building has disappeared or been added, failing to accurately distinguish micro-level renewal types such as structural reinforcement and exterior decoration. Therefore, this invention provides a building renewal type identification method and device based on temporal street view images. The technical solution is as follows:

[0007] On the one hand, a method for recognizing building update types based on time-series street view images is provided. This method is implemented by a building update type recognition device based on time-series street view images, and includes:

[0008] S1: Acquire multiple street view images containing a building complex taken from different time periods and perspectives. The building complex consists of multiple buildings.

[0009] S2: Input the street view image into the YOLO object detection model and output the object detection results of buildings;

[0010] S3: Determine the geographic coordinates of buildings in geospatial space based on the metadata information of the street view image and the target detection results;

[0011] S4: Based on geographic coordinates and combined with target re-identification algorithms, construct an initial geographic information database of building clusters;

[0012] S5: Construct a building update type identification dataset based on the initial building cluster geographic information database;

[0013] S6: Construct a building update type identification model based on the improved Mask2Former architecture;

[0014] S7: Use the building renovation type recognition dataset as training data to train the building renovation type recognition model;

[0015] S8: Input street view images of the same building at different times from the initial building cluster geographic information database into the trained building update type recognition model, and output the building update type recognition result;

[0016] S9: Store the updated type identification results into the initial building cluster geographic information database to obtain the target building cluster geographic information database.

[0017] On the other hand, a building renewal type recognition device based on time-series street view images is provided. The building renewal type recognition device based on time-series street view images includes: a processor; a memory, wherein computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, any one of the methods described above for building renewal type recognition based on time-series street view images is implemented.

[0018] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement any of the above-described methods for identifying building update types based on time-series street view images.

[0019] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0020] (1) By acquiring multiple street view images containing building clusters taken from different periods and perspectives, the street view images are input into the YOLO target detection model, and the target detection results of the buildings are output. Based on the metadata information of the street view images and the target detection results, the geographic coordinates of the buildings in the geographic space are determined. Based on the geographic coordinates and the target re-identification algorithm, an initial building cluster geographic information database is constructed. Based on the initial building cluster geographic information database, a building update type identification dataset is constructed. It is no longer based on single-time point images or single-view information, but has the ability to integrate temporal and multi-view information, and can stably identify the appearance changes of buildings at different times and different shooting perspectives.

[0021] (2) By constructing a building update type identification model based on the improved Mask2Former architecture, and inputting the street view images of the same building at different times in the initial building cluster geographic information database into the trained building update type identification model, the building update type identification results are output. This avoids the problems of low efficiency, high cost and delayed updates that exist in manual surveys. It can meet the needs of large-scale and high-frequency updates, and avoids the problem of limited image resolution in remote sensing image change detection. It can identify building facade details, and can not only determine whether a building has disappeared or been added, but also accurately distinguish micro-update types such as structural reinforcement and exterior decoration. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a building update type recognition method based on temporal street view images provided by an embodiment of the present invention;

[0024] Figure 2 This is a schematic diagram of an improved Mask2Former architecture provided by an embodiment of the present invention;

[0025] Figure 3 This is a schematic diagram of the structure of a building update type recognition device based on time-series street view images provided in an embodiment of the present invention. Detailed Implementation

[0026] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0027] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0028] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0029] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0030] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0031] This invention provides a method for identifying building update types based on time-series street view images. This method can be implemented by a device for identifying building update types based on time-series street view images, which can be a terminal or a server. Figure 1 The flowchart shown is for a building update type recognition method based on time-series street view images. The processing flow of this method may include the following steps:

[0032] S1: Acquire multiple street view images of a building complex taken at different times and from different perspectives. The building complex consists of multiple buildings.

[0033] Optionally, multiple street view images containing building clusters can be acquired from different perspectives at different times within a city or a predetermined area.

[0034] Specifically, street view images are collected at different times by sampling points selected at preset intervals (e.g., every 7.5 meters) along the urban road network, covering building clusters within a city or predetermined area. Each sampling point corresponds to a panoramic image covering 360 degrees horizontally and -90 to +90 degrees vertically, with an original image size of 4096×2048 pixels and containing RGB channels. The collected images must cover the main streets and alleys within the target area to ensure multi-angle and multi-time-point observation coverage of buildings.

[0035] In this embodiment of the invention, by selecting sampling points at preset distances and acquiring street view images from multiple perspectives and different time points, this method ensures comprehensive observation of buildings from multiple angles and at multiple times. This helps improve the accuracy and comprehensiveness of building status identification and adapts to changes in complex urban environments. Ultimately, it can provide high-quality and accurate dynamic update monitoring of buildings from multiple angles and at multiple time dimensions, providing important data support for subsequent decision-making and analysis.

[0036] S2: Input the street view image into the YOLO object detection model and output the object detection results of the buildings.

[0037] It should be noted that YOLO (You Only Look Once) is a real-time object detection model based on deep learning. Its core feature is that it simultaneously locates and classifies objects in an image during a single forward propagation. Compared with traditional detection methods, YOLO uses the entire image as input and transforms the detection task into a unified regression problem, thereby greatly improving detection speed and processing efficiency. YOLO divides the image into multiple grids and predicts the position, confidence score, and class probability of the bounding box in each grid, achieving fast and accurate identification of multiple objects in the image. Its compact structure and fast inference speed make it particularly suitable for object extraction in large-scale image scenes, such as the automatic detection of buildings in street view images in this invention.

[0038] Furthermore, the YOLO object detection model is an existing model, and those skilled in the art can choose the type of YOLO object detection model according to actual needs, which will not be elaborated here.

[0039] The target detection results include: the bounding box coordinates of the building, the category label, and the confidence score.

[0040] In this embodiment of the invention, street view images are input into the YOLO object detection model for building recognition. This method efficiently and accurately extracts building target areas and outputs detection results including bounding box coordinates, category labels, and confidence scores. This approach automates the processing of large-scale street view images, avoiding subjective errors caused by manual annotation and improving system recognition efficiency and data consistency. Furthermore, the confidence score and pixel area filtering mechanism further enhances detection quality, providing a reliable data foundation for subsequent geolocation, re-identification matching, and updated type recognition.

[0041] In one possible implementation, after S2 and before S3, the process further includes:

[0042] S2A: Determine if the confidence level of the building is greater than or equal to the preset confidence level. If yes, retain the street view image containing the building. Otherwise, discard the street view image containing the building.

[0043] S2B: Determine if the pixel area of ​​a building is greater than or equal to a preset pixel area. If yes, retain the street view image containing the building. Otherwise, discard the street view image containing the building.

[0044] It should be noted that those skilled in the art can set the preset confidence level and the size of the preset pixel area according to actual needs, and this invention does not limit these settings.

[0045] In this embodiment of the invention, after the YOLO model completes building target detection, a dual screening mechanism based on confidence level and pixel area is introduced to effectively improve the reliability of the detection results and image quality. By setting a confidence level threshold, false detections or uncertain targets can be eliminated, ensuring the accuracy of building identification. Simultaneously, by setting a minimum pixel area threshold, insufficient image information due to excessively small target areas is avoided, enhancing the stability of subsequent georeferencing and feature extraction. This strategy not only improves the overall accuracy of building localization and identification but also significantly reduces invalid data interference, improving system processing efficiency and model training quality.

[0046] S3: Determine the geographic coordinates of buildings in geospatial space based on the metadata information of the street view image and the target detection results.

[0047] The metadata information includes: the time the street view image was captured, the shooting angle, the latitude and longitude coordinates of the shooting location, and a unique identifier.

[0048] It should be noted that geospatial refers to a collection of spatial information with geographical location attributes, typically including the location, shape, extent, and relationships of objects, scenes, or phenomena on the Earth's surface in two-dimensional or three-dimensional space. It involves not only coordinates (such as latitude and longitude) and geometric shapes, but also temporal, attribute, and semantic information related to a specific location.

[0049] In one possible implementation, S3 specifically includes sub-steps S301 and S302:

[0050] S301: Determine the azimuth angle of the building in the street view image based on the shooting angle and bounding box coordinates.

[0051] S302: Determine the geographic coordinates of the building in geographic space based on the latitude and longitude coordinates of the shooting location and the azimuth angle of the building in the street view image.

[0052] Specifically, the pixel values ​​of street view images are typically 1000. The azimuth angle range is -180° to 180° in the horizontal direction and -90° to 90° in the vertical direction. Street view images are generally taken along the road network, so the horizontal azimuth angle of the roads is 0°. Based on the location of buildings in the street view image, their angles relative to the shooting point can be determined. By using multiple street view images or combining them with OSM building vector data, the geographic coordinates of the buildings can be further determined.

[0053] In this embodiment of the invention, by utilizing the shooting perspective of the street view image, the target detection results, and the latitude and longitude coordinates of the shooting location, combined with the bounding box position of the building in the image, the azimuth angle of the building in the image is calculated. Based on this azimuth angle, the coordinates of the building in geographic space are extrapolated from the shooting point, achieving accurate positioning of the building in the real map. This method requires no additional hardware calibration, relying only on the metadata carried by the street view image itself, and is highly efficient and scalable, providing a high-precision spatial foundation for subsequent building matching, update recognition, and geographic information database construction.

[0054] S4: Based on geographic coordinates and combined with target re-identification algorithms, construct an initial geographic information database of building clusters.

[0055] In one possible implementation, S4 specifically includes sub-steps S401 to S406:

[0056] S401: Obtain the building outlines in geographic coordinates from an open street vector map.

[0057] Optionally, the open block vector map is OSM.

[0058] It should be noted that OSM (OpenStreetMap) is a global open geographic information project that aims to create a set of free, editable, and usable digital map data through crowdsourcing. This platform is maintained collaboratively by volunteers worldwide, and the data includes geographic elements such as roads, buildings, land use, natural features, and public facilities, all stored in vector format, supporting spatial analysis and map visualization. OSM is widely used in navigation, urban planning, environmental monitoring, and other fields, and has the advantages of strong openness, rapid updates, and wide coverage. In this invention, OSM serves as the data source for the open street vector map, used to obtain the outline information of buildings at given geographic coordinates to assist in the spatial matching and identification of building targets.

[0059] S402: Calculate the angle between the azimuth of a building in a street view image and the azimuth of the building's outline center relative to the shooting device.

[0060] S403: Determine if the included angle is less than or equal to a preset included angle. If yes, confirm that the building and building outline match successfully, add a first identifier to the building, and proceed to S406. Otherwise, confirm that the building and building outline match fails, and proceed to S404.

[0061] It should be noted that those skilled in the art can set the size of the preset angle according to actual needs, and this invention does not limit it.

[0062] S404: Using a target re-identification algorithm, cluster the buildings that failed to match, and determine the clustering results as potential buildings.

[0063] It should be noted that Re-Identification (ReID) is a computer vision technique designed to identify the same object across cameras or image viewpoints. This algorithm typically uses deep neural networks to extract discriminative feature vectors of targets (such as pedestrians, vehicles, and buildings) in images and calculates feature similarity across multiple images to achieve automatic matching and identification of the same target. Unlike general object detection, ReID focuses on "whether they are the same object" rather than "which category they belong to." In this invention, the ReID algorithm is used to encode and cluster features of building images that failed in spatial matching, thereby discovering building instances that are essentially the same despite different viewpoints, identifying potential building units, and assigning them a unified identifier.

[0064] Specifically, image features of buildings that failed to match in street view images are extracted, and the image features are encoded using a target re-identification (ReID) algorithm to construct a feature similarity matrix. Based on this, hierarchical clustering or density clustering algorithms (such as DBSCAN) are used to aggregate and analyze the feature vectors, grouping images with high feature similarity, close spatial locations, and matching shooting angles into the same group. Each aggregation result is considered a "potential building".

[0065] S405: Add a second identifier to the potential building.

[0066] S406: Combine the street view images and metadata information corresponding to the buildings after adding identifiers to construct an initial geographic information database of building clusters.

[0067] It should be noted that the purpose of adding the first and second identifiers is to classify, label, and uniquely encode buildings identified from different sources and matching methods, facilitating unified management and accurate identification in the future. The first identifier is used to mark buildings that successfully match open street vector maps (such as OSM) based on angular relationships, indicating a high-confidence map correspondence in their spatial location. The second identifier, on the other hand, is used to mark potential buildings identified through target re-identification and feature clustering, suitable for building identification in situations where map data is missing or viewpoint information is insufficient. By distinguishing between these two types of identifiers, the structure of the building geographic information database can be ensured to be clear, its sources traceable, and its processing logic controllable, providing a reliable indexing basis for subsequent data fusion, update detection, and spatial analysis.

[0068] In this embodiment of the invention, by combining the geographic coordinates of street view images with a target re-identification algorithm, accurate calibration and unified identification of buildings in urban space are achieved. First, buildings that successfully match are obtained by azimuth matching with a vector map and assigned a unique identifier. For targets that fail to match, potential building units are identified through re-identification feature extraction and cluster analysis, and a second identifier is added. Finally, the identification results are integrated with image metadata to construct a structured initial geographic information database of building clusters, providing high-quality data support for subsequent building update detection, temporal analysis, and urban planning.

[0069] S5: Based on the initial building cluster geographic information database, construct a building update type identification dataset.

[0070] In one possible implementation, S5 specifically includes sub-steps S501 to S504:

[0071] S501: In the initial building cluster geographic information database, select street view images corresponding to buildings with the same identifier, and combine two street view images whose latitude and longitude coordinates differ from the preset latitude and longitude coordinates into an image pair.

[0072] It should be noted that those skilled in the art can set the magnitude of the preset latitude and longitude coordinate difference according to actual needs, and this invention does not limit this.

[0073] Specifically, street view image samples with the same building identifiers are selected from the initial building cluster geographic information database. Based on the latitude and longitude coordinates of the images in the image metadata, the spatial distance between image pairs is calculated. When this distance is less than or equal to a preset latitude and longitude difference threshold (e.g., 10 meters), they are considered to have similar shooting locations. At the same time, two images with a large difference in shooting time are preferentially selected to be combined into image pairs.

[0074] S502: Perform enhancement operations on the image pair to obtain an enhanced image pair.

[0075] Optionally, enhancement operations include: adjusting the contrast, saturation, multi-angle projection, noise reduction, or blurring of image pairs.

[0076] S503: Annotate the enhanced image pairs.

[0077] Specifically, each pair of enhanced images is arranged chronologically according to the shooting time. By comparing the differences in architectural appearance features between the images, the changes in their update status are initially determined. Pre-trained change detection models or semantic segmentation models are prioritized for automatic identification and difference extraction of architectural regions within the image pairs. Initial annotation results are generated by combining the location, contour, texture, and color features of the differing regions. A high-confidence screening strategy is employed, retaining only labeled samples with a model confidence score higher than a preset threshold (e.g., 80%). For samples with low confidence or ambiguous boundaries, correction and confirmation are performed with manual assistance or human-computer interaction. Finally, pseudo-labels are generated and the update type is classified, including but not limited to: structural reinforcement (e.g., adding supporting components), exterior decoration (e.g., facade painting or decorative renovation), demolition and reconstruction (significant changes in architectural style and contour), and ongoing construction (e.g., scaffolding or obstruction), providing accurate supervision information for the subsequent construction of an update type recognition dataset.

[0078] S504: Construct a building update type recognition dataset based on the annotated enhanced image pairs.

[0079] In this embodiment of the invention, street view images with the same identifier and similar shooting locations are selected from the initial building cluster geographic information database to construct spatially consistent image pairs with significant temporal differences. Image enhancement operations are then combined to improve sample diversity. Subsequently, a pre-trained model is used for preliminary labeling, supplemented by human-computer interaction correction, to generate pseudo-labels including categories such as structural reinforcement, exterior decoration, demolition and reconstruction, and under construction. Finally, a high-quality, strongly supervised building update type recognition dataset is constructed, providing accurate, stable, and scalable data support for subsequent model training and building appearance change recognition.

[0080] S6: Construct a building update type identification model based on the improved Mask2Former architecture.

[0081] It should be noted that Mask2Former is an advanced unified full-image mask prediction architecture, originally designed for image segmentation tasks, capable of simultaneously supporting semantic segmentation, instance segmentation, and panoptic segmentation. The core idea of ​​this architecture is to generate a set of image-level learnable queries using a Transformer-based mask decoder, with each query corresponding to a target mask. It then extracts associated region information from image features through a cross-attention mechanism. Mask2Former uses a Pixel Decoder + Transformer Decoder decoding structure, effectively combining local pixel features with global contextual relationships, resulting in powerful target representation capabilities and segmentation accuracy. In this invention, Mask2Former is further improved to identify changes in the updated state of buildings in images from different periods, capturing subtle differences in building appearance through its accurate region segmentation capabilities.

[0082] In one possible implementation, S6 specifically includes sub-steps S601 and S602:

[0083] S601: Replace the Swin Transformer backbone network in the original Mask2Former architecture with a hybrid convolutional-Transformer backbone network.

[0084] S602: After the intermediate module of the original Mask2Former architecture and before the maskdecoder module of the original Mask2Former architecture, a temporal image processing network is introduced to build a building update type recognition model based on the improved Mask2Former architecture.

[0085] Optionally, the hybrid convolutional-Transformer backbone network includes: standard convolutional layers (i.e., Figure 2 The traditional convolutional layers, depthwise separable convolutional layers, depthwise convolutional layers, depthwise convolutional layers with residual connections, first Transformer encoder layers, second Transformer encoder layers, third Transformer encoder layers, and Transformer encoder with residual connections are all included.

[0086] It should be noted that the hybrid convolutional-Transformer backbone network adopts a four-stage hierarchical structure (Stage 1 to Stage 4), each responsible for different levels of feature extraction and semantic modeling, taking into account both local texture details and global contextual information. Stage 1 includes two convolutional layers (standard convolutional layer and depthwise separable convolutional layer): First, a standard 3×3 convolutional kernel (stride=2) is used to extract the basic features of the image, with 32 output channels. Subsequently, depthwise separable convolution is introduced to further refine the extraction of local texture information, with the number of output channels still at 32, supplemented by batch normalization and activation functions (such as ReLU) for non-linear mapping and feature normalization. Stage 2 consists of two depthwise convolutional (3×3) layers (depthwise convolutional layer and depthwise convolution + residual connection layer): the first is a standard depthwise convolutional layer, and the second is a depthwise convolutional layer integrating residual structures. The number of output channels increases from 32 to 64. This stage introduces a residual connection mechanism to optimize information flow, effectively alleviate the gradient vanishing problem, and enhance the modeling ability for building edge contours and detailed textures. Stage 3 contains two Transformer encoder layers (first and second Transformer encoder layers), each consisting of a multi-head attention mechanism and a feed-forward network. It uses four attention heads, each with a dimension of 64, to capture complex spatial relationships and contextual dependencies within the building area, improving the model's ability to express cross-scale features. Stage 4 is responsible for deeper feature aggregation and semantic enhancement, also containing a two-layer Transformer encoder structure (third Transformer encoder layer and Transformer encoder + residual connection layer). Compared to Stage 3, Stage 4 adds cross-layer residual connections and layer normalization mechanisms to ensure the stability of deep information transmission and improve the model's discriminative ability and generalization performance in handling complex building appearance changes.

[0087] The intermediate modules include: a positional encoding layer, a multi-scale fusion layer, an upsampling layer, and a 1×1 convolutional layer.

[0088] Temporal image processing networks include: self-attention modules and cross-attention modules.

[0089] Furthermore, the temporal image processing network also includes two Add & norm integration layers.

[0090] In one possible implementation, the temporal image processing network is specifically used for:

[0091] The input image sequence is processed by the backbone network and intermediate modules to generate feature map sequences of buildings from different periods. For example, if a building contains street view images from years t1 and t2, its feature map sequence will include architectural feature images from two different years (t1 and t2). and ,in, and The images represent architectural features from years t1 and t2, respectively. R represents the range of pixel values ​​in the image, H and W represent the height and width of the image, respectively, and C represents the number of channels in the image.

[0092] A self-attention module is used to receive the image feature sequence output from the feature extraction module. By modeling the attention relationship between spatial locations within the image at each time step, the fusion of local and global features is achieved, thereby enhancing the semantic consistency and spatial expressiveness within the image and providing a foundation for cross-temporal comparison by the subsequent cross-attention module.

[0093] A cross-attention module is used to introduce a cross-attention mechanism to align and fuse features of images at different time points. Specifically, a trainable projection matrix W is used. Q and W K right and Perform a linear transformation to align them to a shared low-dimensional space: , ,in, , Let be the projection matrix, Q and K be the projected feature representations (representing the transformed features of the images from years t1 and t2, respectively), and d be the mapped feature dimension.

[0094] Specifically, Q and K are input to the cross-attention module for alignment, the feature weights of year t1 to year t2 are calculated, and a new feature map is obtained. :

[0095]

[0096] in, (That is, the keys and values ​​in the self-attention mechanism are the same, V is the same matrix as K, representing "value"), softmax represents the softmax activation function (used to normalize weights and ensure the weighted synthesis of each feature).

[0097] Quantify Q using the L1 norm and Difference values:

[0098]

[0099] in, This represents the difference value at each spatial location, used to measure local changes in the image.

[0100] The feature difference value D is converted into a similarity score, reflecting the significance of the changed regions:

[0101]

[0102] Where α represents an adjustable hyperparameter that controls the sensitivity of similarity to feature differences. The similarity score represents the significance of changes in an image region, and exp represents the exponential function.

[0103] Finally, the similarity matrix S is multiplied by the feature map of year t1 to obtain a weighted feature map, which is then fed into the mask decoder module for more refined target segmentation, ultimately obtaining the mask of the changed area before and after the building.

[0104] The mask decoder module includes: a pixel decoder (i.e., Figure 2 (Pixel decoder) and Transformer decoder.

[0105] In this embodiment of the invention, the improved Mask2Former architecture significantly enhances the performance of the building update type recognition model by introducing a hybrid convolutional-transformer backbone network and a temporal image processing network. The hybrid convolutional-transformer backbone combines the local feature extraction capability of convolutional networks with the global context modeling capability of Transformers, effectively enhancing the model's sensitivity to changes in buildings at different scales. The temporal image processing network, through self-attention and cross-attention modules, improves the model's accuracy in recognizing subtle changes between images from different periods. The intermediate module ensures accurate segmentation of the building update region through multi-scale feature fusion and upsampling. Simultaneously, the use of depthwise separable convolutions and residual connections optimizes the model's computational efficiency and robustness, making it suitable for processing large-scale street view image data. This invention effectively improves the accurate recognition capability of building update status and is applicable to application scenarios such as urban building update monitoring, urban planning, and building safety management.

[0106] S7: Use the building renewal type recognition dataset as training data to train the building renewal type recognition model.

[0107] In this embodiment of the invention, a high-quality, finely annotated building update type identification dataset is used as training data to train the improved identification model. This helps the model fully learn the feature differences and change patterns between different update types, thereby significantly improving its identification accuracy and classification accuracy for minor changes in building appearance. It also enhances the model's generalization ability and robustness in real complex urban street scenes, providing a solid intelligent foundation for subsequent automated building update detection.

[0108] S8: Input street view images of the same building at different times from the initial building cluster geographic information database into the trained building update type recognition model, and output the building update type recognition result.

[0109] It should be noted that the updated type identification results are structural reinforcement (such as adding supporting components), appearance decoration (such as exterior painting or decoration renovation), demolition and reconstruction (significant changes in architectural style and outline), and construction in progress (such as the presence of scaffolding or obstruction).

[0110] In this embodiment of the invention, street view images of the same building at different times from the initial building cluster geographic information database are input into the trained recognition model to achieve intelligent identification of the target building's renewal type. This not only quickly distinguishes specific renewal states such as structural reinforcement, exterior decoration, demolition and reconstruction, and ongoing construction, but also significantly improves the efficiency and automation level of renewal identification, reduces manual intervention, and helps promote the intelligent advancement of application scenarios such as urban renewal, regulatory decision-making, and digital building archive management.

[0111] S9: Store the updated type identification results into the initial building cluster geographic information database to obtain the target building cluster geographic information database.

[0112] In this embodiment of the invention, the update type identification result is stored in the initial building cluster geographic information database to obtain the target building cluster geographic information database, realizing comprehensive dynamic management of buildings from spatial location, image metadata to update status. This method not only constructs a traceable, time-series building archive but also improves data integration and retrieval efficiency, avoiding information fragmentation and redundant processing. Simultaneously, the database's dynamic update capability provides high-precision data support for urban planning, building safety supervision, and the construction of digital twin cities, significantly enhancing the system's practicality, scalability, and decision-making service capabilities.

[0113] Figure 3 This is a schematic diagram of a building update type recognition device based on time-series street view images provided in an embodiment of the present invention. Optionally, the building update type recognition device 410 based on time-series street view images may include a first processor 2001.

[0114] Optionally, the building update type recognition device 410 based on time-series street view images may also include a memory 2002 and a transceiver 2003.

[0115] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.

[0116] The following is combined with Figure 3 The components of the building update type recognition device 410 based on time-series street view images are described in detail below:

[0117] The first processor 2001 is the control center of the building update type recognition device 410 based on time-series street view images. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).

[0118] Optionally, the first processor 2001 can perform various functions of the building update type recognition device 410 based on time-series street view images by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0119] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 3 CPU0 and CPU1 are shown in the diagram.

[0120] In a specific implementation, as one example, the building update type recognition device 410 based on temporal street view images may also include multiple processors, for example... Figure 3 The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).

[0121] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.

[0122] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently, and may be connected via the interface circuit of the building update type recognition device 410 based on time-series street view images. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0123] The transceiver 2003 is used to communicate with network devices or with terminal devices.

[0124] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 3 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.

[0125] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and can be connected to the interface circuit of the building update type recognition device 410 based on time-series street view images. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0126] It should be noted that, Figure 3 The structure of the building update type identification device 410 based on time-series street view images shown does not constitute a limitation on this router. Actual building update type identification devices based on time-series street view images may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0127] Furthermore, the technical effects of the building renewal type recognition device 410 based on time-series street view images can be referred to the technical effects of the building renewal type recognition method based on time-series street view images described in the above method embodiments, and will not be repeated here.

[0128] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or it may be any conventional processor, etc.

[0129] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0130] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0131] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0132] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0133] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0134] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0135] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0136] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0137] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0138] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0139] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0140] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for identifying building update types based on temporal street view images, characterized in that, The method comprises: S1: acquiring multiple street view images containing a building group taken at different time periods and different perspectives, the building group being composed of multiple buildings; S2: inputting the street view images into a YOLO target detection model to output a target detection result of the buildings; S3: determining geographical coordinates of the buildings in geographical space according to metadata information of the street view images and the target detection result; S4: constructing an initial building group geographical information database according to the geographical coordinates in combination with a target re-identification algorithm; S5: constructing a building update type identification data set according to the initial building group geographical information database; S6: constructing a building update type identification model based on an improved Mask2Former architecture; S7: training the building update type identification model by taking the building update type identification data set as training data; S8: inputting street view images of the same building at different time periods in the initial building group geographical information database into the trained building update type identification model to output an update type identification result of the building; S9: storing the update type identification result into the initial building group geographical information database to obtain a target building group geographical information database; The S6 specifically comprises: S601: replacing a Swin Transformer backbone network in an original Mask2Former architecture with a hybrid convolution-Transformer backbone network; S602: introducing a time series image processing network after a middle module of the original Mask2Former architecture and before a mask decoder module of the original Mask2Former architecture to construct the building update type identification model based on the improved Mask2Former architecture; The hybrid convolution-Transformer backbone network comprises a standard convolution layer, a depth separable convolution layer, a depth convolution layer, a depth convolution+residual connection layer, a first Transformer encoder layer, a second Transformer encoder layer, a third Transformer encoder layer, and a Transformer encoder+residual connection layer; The middle module comprises a position encoding layer, a multi-scale fusion layer, an up-sampling layer, and a 1x1 convolution layer; The time series image processing network comprises a self-attention module and a cross-attention module; The mask decoder module comprises a Pixel decoder and a Transformer decoder. 2.The building update type identification method based on time-sequential street view images according to claim 1, wherein, The target detection result comprises boundary box coordinates, a class label, and a confidence of the building; The metadata information comprises a shooting time, a shooting perspective, a shooting position longitude and latitude coordinates, and a unique identifier of the street view images. 3.The building update type identification method based on time-sequential street view images according to claim 2, characterized in that, After the S2 and before the S3, the method further comprises: S2A: determining whether the confidence of the building is greater than or equal to a preset confidence; if yes, retaining the street view images containing the building; otherwise, discarding the street view images containing the building; S2B: judging whether the pixel area of the building is greater than or equal to a preset pixel area; if yes, retaining the street view image containing the building; otherwise, removing the street view image containing the building. 4.The building update type identification method based on time-sequential street view images according to claim 2, wherein, The S3 specifically comprises: S301: determining an azimuth angle of the building in the street view image according to the shooting angle and the bounding box coordinates; S302: determining a geographical coordinate of the building in geographical space according to the shooting position latitude-longitude coordinates and the azimuth angle of the building in the street view image. 5.The building update type identification method based on time-sequential street view images according to claim 1, wherein, The S4 specifically comprises: S401: obtaining a building contour at the geographical coordinate in the open block vector map; S402: calculating an angle included angle between the azimuth angle of the building in the street view image and the azimuth angle of the center of the building contour relative to the shooting device; S403: judging whether the angle included angle is less than or equal to a preset angle included angle; if yes, determining that the building and the building contour match successfully, adding a first identifier to the building, and entering S406; otherwise, determining that the building and the building contour match unsuccessfully, and entering S404; S404: clustering each building matching unsuccessfully by using the target re-identification algorithm, and determining a clustering result as a potential building; S405: adding a second identifier to the potential building; S406: combining the street view image corresponding to the building after adding the identifier and the metadata information to construct the initial building group geographical information database. 6.The building update type identification method based on time-sequential street view images according to claim 1, wherein, The S5 specifically comprises: S501: selecting street view images corresponding to buildings with the same identifier in the initial building group geographical information database, and combining two street view images with a shooting position latitude-longitude coordinate difference less than or equal to a preset latitude-longitude coordinate difference into an image pair; S502: performing an enhancement operation on the image pair to obtain an enhanced image pair; S503: labeling the enhanced image pair; S504: constructing the building update type identification data set according to the labeled enhanced image pair.

7. A building update type recognition device based on time-series street view images, characterized by, The building update type identification device based on time-series street view images comprises: a processor; a memory, wherein the memory stores computer readable instructions, and the computer readable instructions are executed by the processor to implement the method in any one of claims 1 to 6.

8. A computer readable storage medium, characterized in that, The computer readable storage medium stores program codes, and the program codes can be called and executed by the processor to implement the method in any one of claims 1 to 6.

Citation Information

Patent Citations

  • High-resolution remote sensing image target state discrimination method and device, and storage medium

    CN118691877A

  • Urban non-fortified building reinforcement identification method and device based on historical streetscape pictures

    CN119418207A