Building update type identification method and device based on time sequence streetscape image
By using a building update type recognition method based on temporal street view images, and leveraging YOLO object detection and an improved Mask2Former architecture, the problem of insufficient temporal and multi-view fusion in traditional methods is solved, achieving efficient and accurate recognition and large-scale monitoring of building update types.
Patent Information
- Application Number
- CN202511443391.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Traditional building renovation type identification methods lack temporal and multi-view fusion capabilities, making it impossible to reliably identify changes in the appearance of buildings at different times and from different shooting angles. Furthermore, manual surveys are inefficient and costly, and remote sensing images have limited resolution, making it difficult to identify details of building facades.
A building update type identification method based on time-series street view images is adopted. Building targets are extracted by YOLO target detection model, and a geographic information database is constructed by combining geographic coordinates and target re-identification algorithm. An improved Mask2Former architecture building update type identification model is used to identify the update type of buildings.
It achieves stable identification of building appearance changes at different times and from different perspectives, and can accurately distinguish micro-update types such as structural reinforcement and exterior decoration, meeting the needs of large-scale and high-frequency updates, reducing the cost of manual surveys, and improving identification efficiency and accuracy.
Smart Images

Figure CN120913084A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of civil engineering and computer vision technology, in particular to a building update type identification method and device based on time-series street view images. BACKGROUND
[0002] With the continuous acceleration of urbanization, the number and density of urban building groups continue to rise, and the update behaviors such as demolition, renovation, reinforcement and reconstruction of buildings are increasingly frequent. Accurately grasping the update state of buildings has important practical significance for the fields of urban planning, building safety supervision, disaster risk assessment and historical block protection.
[0003] In the prior art, the traditional building update type identification method is based on single-time-point images or single-view information, lacks time-series and multi-view fusion capabilities, and cannot stably identify the appearance changes of buildings at different times and different shooting angles.
[0004] In addition, the traditional building update type identification method mainly relies on manual investigation or remote sensing image change detection.
[0005] However, manual investigation has low efficiency, high cost and update lag, and cannot meet the needs of large-scale and high-frequency updates; while the change detection based on remote sensing images can achieve large-scale coverage, but due to the limited image resolution, it is difficult to identify the details of building facades, and can only judge whether the building disappears or is newly added, and cannot accurately distinguish the micro update types such as structural reinforcement and appearance decoration. SUMMARY
[0006] In order to solve the technical problems in the prior art that the traditional building update type identification method is based on single-time-point images or single-view information, lacks time-series and multi-view fusion capabilities, and cannot stably identify the appearance changes of buildings at different times and different shooting angles, and manual investigation has low efficiency, high cost and update lag, and cannot meet the needs of large-scale and high-frequency updates, and the change detection based on remote sensing images has limited image resolution, and it is difficult to identify the details of building facades, and can only judge whether the building disappears or is newly added, and cannot accurately distinguish the micro update types such as structural reinforcement and appearance decoration, the embodiments of the present application provide a building update type identification method and device based on time-series street view images. The technical solution is as follows:
[0007] On the one hand, a building update type identification method based on time-series street view images is provided, which is realized by a building update type identification device based on time-series street view images, and the method comprises:
[0008] S1: acquiring multiple street view images containing a building group taken at different periods and different angles, the building group being composed of multiple buildings;
[0009] S2: inputting the street view image into a YOLO target detection model to output a target detection result of the building;
[0010] S3: determining geographical coordinates of the building in a geographical space according to metadata information of the street view image and the target detection result;
[0011] S4: constructing an initial building group geographical information database according to the geographical coordinates and in combination with a target re-identification algorithm;
[0012] S5: constructing a building update type identification dataset according to the initial building group geographical information database;
[0013] S6: constructing a building update type identification model based on an improved Mask2Former architecture;
[0014] S7: training the building update type identification model by taking the building update type identification dataset as training data;
[0015] S8: inputting street view images of the same building in different periods in the initial building group geographical information database into the trained building update type identification model to output an update type identification result of the building;
[0016] S9: storing the update type identification result into the initial building group geographical information database to obtain a target building group geographical information database.
[0017] On the other hand, a building update type identification device based on time-series street view images is provided, which comprises a processor and a memory having computer readable instructions stored thereon, the computer readable instructions being executed by the processor to implement any one of the above building update type identification methods based on time-series street view images.
[0018] On the other hand, a computer readable storage medium is provided, which stores at least one instruction, the at least one instruction being loaded and executed by a processor to implement any one of the above building update type identification methods based on time-series street view images.
[0019] The technical solutions provided by the embodiments of the present application have at least the following beneficial effects:
[0020] (1) By acquiring multiple street view images containing building groups taken at different time periods and different angles, inputting the street view images into a YOLO target detection model, outputting the target detection results of the buildings, determining the geographic coordinates of the buildings in the geographic space according to the metadata information of the street view images and the target detection results, and according to the geographic coordinates, combining a target re-identification algorithm, an initial building group geographic information database is constructed, and according to the initial building group geographic information database, a building update type identification dataset is constructed, which no longer relies on single time point images or single angle information, has the ability of time sequence and multi-angle fusion, and can stably identify the appearance changes of buildings at different times and different shooting angles.
[0021] (2) By constructing a building update type identification model based on the improved Mask2Former architecture, and inputting the street view images of the same building at different time periods in the initial building group geographic information database into the trained building update type identification model, the update type identification result of the building is output, which avoids the low efficiency, high cost and update lag phenomenon of manual investigation, can meet the demand of large-scale and high-frequency update, avoids the phenomenon of limited image resolution of change detection based on remote sensing images, can identify the building facade details, and can not only judge whether the building disappears or is newly added, but also accurately distinguish the microscopic update types such as structure reinforcement and appearance decoration. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0023] Figure 1 is a building update type identification method flowchart provided by the embodiment of the present application based on time sequence street view images;
[0024] Figure 2 is a structure diagram of an improved Mask2Former architecture provided by the embodiment of the present application;
[0025] Figure 3 is a structure diagram of a building update type identification device based on time sequence street view images provided by the embodiment of the present application. DETAILED DESCRIPTION
[0026] The technical solutions in the present application will be described below in combination with the drawings.
[0027] In the embodiments of the present application, the words such as "example", "for example" are used to represent an example, illustration, or description. Any embodiment or design scheme described as "example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific manner. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be one of the two.
[0028] In the embodiments of the present application, "image" and "picture" can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized. "Of", "corresponding" and "corresponding" can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized.
[0029] In the embodiments of the present application, sometimes the subscript such as W1 can be written in the form of non-subscript such as W1, and the meanings expressed are consistent when the distinction is not emphasized.
[0030] In order to make the technical problems, technical schemes and advantages to be solved by the present application more clear, the following will be described in detail in conjunction with the drawings and specific embodiments.
[0031] The embodiments of the present application provide a building update type identification method based on time sequence street view images, which can be realized by a building update type identification device based on time sequence street view images. The building update type identification device based on time sequence street view images can be a terminal or a server. As shown in the flow chart of the building update type identification method based on time sequence street view images, the processing flow of the method can include the following steps: Figure 1
[0032] S1: Obtain multiple street view images containing a building group taken at different time periods and different angles, the building group being composed of multiple buildings.
[0033] Optionally, multiple street view images containing a building group taken at different time periods and different angles in a city or a predetermined area are obtained.
[0034] Specifically, the street view images are collected at different time points by sampling points selected at a preset distance (for example, every 7.5 meters) along the city road network in the city or the predetermined area. The image corresponding to each sampling point is a panoramic image with a horizontal direction of 360 degrees and a vertical direction covering -90 to +90 degrees. The original image size is 4096x2048 pixels, and the image contains three channels of RGB. The collected images need to cover the main streets and alleys in the target area to ensure the observation coverage of the buildings at multiple angles and multiple time points.
[0035] In the embodiment of the present application, by selecting sampling points at a preset distance and collecting street view images from multiple perspectives and different time points, this method can ensure all-around observation of buildings from multiple perspectives and different time points, which helps to improve the accuracy and comprehensiveness of building update state recognition and adapt to changes in complex urban environments. Ultimately, high-quality and accurate dynamic update monitoring of buildings can be provided in multiple perspectives and time dimensions, providing important data support for subsequent decision-making and analysis.
[0036] S2: inputting the street view image into a YOLO target detection model to output a target detection result of the building.
[0037] It should be noted that YOLO (You Only Look Once) is a real-time target detection model based on deep learning, and its core feature is to complete the positioning and classification of targets in an image in a single forward propagation. Compared with traditional detection methods, YOLO takes the whole image as input and converts the detection task into a unified regression problem, thereby greatly improving the detection speed and processing efficiency. YOLO divides the image into multiple grids and predicts the position, confidence and class probability of the bounding box in each grid to achieve fast and accurate identification of multiple targets in the image. Its structure is compact and the inference speed is fast, and it is particularly suitable for target extraction in large-scale image scenarios, such as automatic detection of buildings in street view images in the present application.
[0038] Further, the YOLO target detection model is an existing model, and those skilled in the art can select the type of YOLO target detection model according to actual needs, which will not be described herein.
[0039] The target detection result includes the bounding box coordinates, class label and confidence of the building.
[0040] In the embodiment of the present application, the street view image is input into the YOLO target detection model for building recognition, which can efficiently and accurately extract the building target area and output the detection result including the bounding box coordinates, class label and confidence. This method realizes the automatic processing of large-scale street view images, avoids subjective errors caused by manual annotation, improves the system recognition efficiency and data consistency. Through the confidence and pixel area screening mechanism, the detection quality is further improved, providing a reliable data basis for subsequent geographic positioning, re-identification matching and update type recognition.
[0041] In a possible implementation, after S2 and before S3, the method further includes:
[0042] S2A: determining whether the confidence of the building is greater than or equal to a preset confidence. If yes, the street view image containing the building is retained. Otherwise, the street view image containing the building is removed.
[0043] S2B: determining whether the pixel area of the building is greater than or equal to a preset pixel area. If yes, the street view image containing the building is retained. Otherwise, the street view image containing the building is removed.
[0044] It should be noted that the skilled person in the art can set the size of the preset confidence and the preset pixel area according to actual needs, which is not limited in the present application.
[0045] In the embodiments of the present application, after the YOLO model completes the building target detection, a double screening mechanism based on confidence and pixel area is introduced, which effectively improves the reliability of the detection result and the image quality. By setting the confidence threshold, the model mis-detection or uncertain target can be removed, ensuring the accuracy of building recognition. At the same time, by setting the minimum pixel area threshold, the lack of image information caused by too small target area is avoided, and the stability of subsequent geographic registration and feature extraction is enhanced. This strategy not only improves the overall precision of building positioning and recognition, but also significantly reduces the interference of invalid data, improves the system processing efficiency and model training quality.
[0046] S3: determining the geographic coordinates of the building in the geographic space according to the metadata information of the street view image and the target detection result.
[0047] The metadata information includes the shooting time, shooting angle, shooting location latitude and longitude coordinates, and unique identifier of the street view image.
[0048] It should be noted that the geographic space refers to a collection of spatial information with geographic location attributes, usually including the position, shape, range and relationship of objects, scenes or phenomena on the earth's surface in two or three-dimensional space. It not only involves coordinates (such as latitude and longitude) and geometric shapes, but also includes time, attributes and semantic information related to specific locations.
[0049] In one possible implementation, S3 specifically includes sub-steps S301 and S302:
[0050] S301: determining the azimuth angle of the building in the street view image according to the shooting angle and the bounding box coordinates.
[0051] S302: determining the geographic coordinates of the building in the geographic space according to the shooting location latitude and longitude coordinates and the azimuth angle of the building in the street view image.
[0052] Specifically, the pixel value of the street view image is usually , indicating that the azimuth angle range in the horizontal direction is -180° to 180°, and the azimuth angle range in the vertical direction is -90° to 90°. Street view images are generally taken along the road network, so the horizontal azimuth angle of the road is 0°. According to the position of the building in the street view image, the included angle of the building relative to the shooting point can be determined. Through multiple street view images or in combination with OSM building vector data, the geographical coordinates of the building can be further determined.
[0053] In the embodiment of the present application, by using the shooting angle of the street view image, the target detection result and the latitude and longitude coordinates of the shooting position, and combining the position of the building boundary box in the image, the azimuth angle of the building in the image is calculated, and the coordinates of the building in the geographical space are calculated outward from the shooting point based on the azimuth angle, thereby realizing accurate positioning of the building in the real map. This method does not require additional hardware calibration, but only relies on the metadata carried by the street view image itself, has high efficiency and scalability, and provides a high-precision spatial basis for subsequent building matching, updating identification and geographical information database construction.
[0054] S4: According to the geographical coordinates, an initial building group geographical information database is constructed in combination with a target re-identification algorithm.
[0055] In a possible implementation, S4 specifically includes sub-steps S401 to S406:
[0056] S401: Obtain the building contour at the geographical coordinates in the open street block vector map.
[0057] Optionally, the open street block vector map is OSM.
[0058] It should be noted that OSM (OpenStreetMap) is a global open geographic information project, aiming to create a set of free editable, free to use digital map data through crowdsourcing. The platform is maintained by global volunteers, and the data includes roads, buildings, land use, natural features, public facilities and other geographic elements, which are stored in vector form and support spatial analysis and map visualization. OSM is widely used in navigation, urban planning, environmental monitoring and other fields, and has the advantages of strong openness, fast updating and wide coverage. In the present application, OSM is used as a data source for the open street block vector map to obtain the contour information of the building at a given geographical coordinate, to assist the spatial matching and identification of the building target.
[0059] S402: Calculate the included angle between the azimuth angle of the building in the street view image and the azimuth angle of the building contour center relative to the shooting device.
[0060] S403: Determine whether the angle is less than or equal to the preset angle. If yes, determine that the building and the building contour match successfully, add a first identifier to the building, and enter S406. Otherwise, determine that the building and the building contour match unsuccessfully, and enter S404.
[0061] It should be noted that the size of the preset angle can be set by the person skilled in the art according to actual needs, and the present application does not limit it.
[0062] S404: Cluster each building that fails to match by using a target re-identification algorithm, and determine the clustering result as a potential building.
[0063] It should be noted that the target re-identification algorithm (Re-Identification, ReID for short) is a computer vision technology that aims to identify the same object across cameras or across image perspectives. The algorithm usually extracts a discriminative feature vector of a target (such as a pedestrian, a vehicle, a building, etc.) in an image through a deep neural network, and calculates the feature similarity between multiple images to achieve automatic matching and identification of the same target. Unlike general object detection, re-identification focuses on "whether it is the same object", rather than "which category it belongs to". In the present application, the target re-identification algorithm is used to encode and cluster the building images that fail in spatial matching, so as to discover building instances with different perspectives but the same essence, identify potential building units and assign a unified identifier.
[0064] Specifically, the image features of the buildings that fail to match in the street view image are extracted, the image features are encoded using a target re-identification (ReID) algorithm, and a feature similarity matrix is constructed. On this basis, a hierarchical clustering or density clustering algorithm (such as DBSCAN) is used to aggregate and analyze the feature vectors, and images with high feature similarity, close spatial position and consistent shooting angle are classified into the same group. Each group of aggregation results is considered as a "potential building".
[0065] S405: Add a second identifier to the potential building.
[0066] S406: Combine the street view images and metadata information corresponding to the buildings after adding the identifiers to construct an initial building group geographic information database.
[0067] It should be noted that the purposes of adding the first identifier and the second identifier are to classify and uniquely encode the buildings identified under different sources and matching modes, so as to be uniformly managed and accurately identified subsequently. The first identifier is used to mark those buildings successfully matched by the angle included angle judgment and the open street block vector map (such as OSM), indicating that the spatial position has a high confidence map correspondence. The second identifier is used to mark the potential buildings determined by the target re-identification and feature clustering method, which is suitable for building identification under the condition of missing map data or insufficient view information. By distinguishing the two types of identifiers, the structure of the building geographic information database can be clear, the source can be traced, and the processing logic can be controlled, providing reliable index basis for subsequent data fusion, update detection and spatial analysis.
[0068] In the embodiment of the present application, by combining the geographic coordinates of the street view image and the target re-identification algorithm, accurate calibration and unified identification of buildings in urban space are realized. First, the successfully matched buildings are obtained by using the azimuth angle matching with the vector map, and a unique identifier is given. For the targets that fail to match, potential building units are identified by re-identifying feature extraction and clustering analysis, and a second identifier is added. Finally, the identification results and image metadata are integrated to construct a structured initial building group geographic information database, which provides high-quality data support for subsequent building update detection, time series analysis and urban planning.
[0069] S5: Constructing a building update type identification data set according to the initial building group geographic information database.
[0070] In one possible implementation, S5 specifically includes sub-steps S501 to S504:
[0071] S501: In the initial building group geographic information database, select the street view images corresponding to the buildings with the same identifier, and combine two street view images with a shooting position latitude and longitude coordinate difference less than or equal to a preset latitude and longitude coordinate difference into an image pair.
[0072] It should be noted that the size of the preset latitude and longitude coordinate difference can be set by the person skilled in the art according to actual needs, which is not limited in the present application.
[0073] Specifically, the street view image samples with the same building identifier are selected from the initial building group geographic information database, and the spatial distance between the image pairs is calculated according to the shooting latitude and longitude coordinate information in the image metadata. When the distance is less than or equal to the preset latitude and longitude difference threshold (such as 10 meters), it is considered that the shooting positions are close. At the same time, two images with a larger time difference are preferentially selected to form an image pair.
[0074] S502: Performing an enhancement operation on the image pair to obtain an enhanced image pair.
[0075] Optionally, the enhancement operation comprises adjusting contrast, saturation, multi-angle projection, noise reduction or blur processing of the image pair.
[0076] S503: Label the enhanced image pair.
[0077] Specifically, each pair of enhanced images is arranged in chronological order, and the update state change is preliminarily judged by comparing the differences in building appearance features between the images. Preferably, a pre-trained change detection model or a semantic segmentation model is used to automatically identify and extract differences in the building area in the image pair, and an initial labeling result is generated by combining the features such as the location, contour, texture and color of the difference area. A high-confidence screening strategy is adopted to retain only the labeling samples with a model confidence higher than a preset threshold (e.g. 80%). For samples with low confidence or fuzzy boundaries, manual assistance or human-computer interaction is used for correction and confirmation. Finally, pseudo-labels are generated and classified to label the update types, including but not limited to structural reinforcement (such as adding support members), appearance decoration (such as exterior surface painting or decoration reconstruction), demolition and reconstruction (significant changes in building style and contour), and construction in progress (such as the presence of scaffolding or being blocked), which provides accurate supervision information for subsequent construction of update type recognition dataset.
[0078] S504: Construct a building update type recognition dataset according to the labeled enhanced image pair.
[0079] In the embodiments of the present application, by screening street view images with the same identifier and close shooting positions from the initial building group geographic information database, image pairs with spatial consistency and obvious time difference are constructed, the sample diversity is improved by combining image enhancement operation, then a pre-trained model is used for preliminary labeling and human-computer interaction is used for correction, pseudo-labels including structural reinforcement, appearance decoration, demolition and reconstruction, and construction in progress are generated, and finally a high-quality and strongly supervised building update type recognition dataset is constructed, which provides accurate, stable and expandable data support for subsequent model training and building appearance change recognition.
[0080] S6: Construct a building update type recognition model based on the improved Mask2Former architecture.
[0081] It should be noted that Mask2Former is an advanced unified full-image mask prediction architecture, originally used for image segmentation tasks, which can simultaneously support semantic segmentation, instance segmentation and panorama segmentation. The core idea of the architecture is to generate a set of image-level learnable queries through a Transformer-based mask decoder, each query corresponding to a target mask, and extracting the associated region information from the image features through a cross-attention mechanism. Mask2Former uses a decoding structure of Pixel Decoder + Transformer Decoder, effectively combining local pixel features and global context relationships, and has strong target expression ability and segmentation accuracy. In the present application, Mask2Former is further improved for identifying the update state changes of buildings in images at different times, and the precise region segmentation capability thereof is used to capture the differences in building appearance details.
[0082] In one possible implementation, S6 specifically includes sub-steps S601 and S602:
[0083] S601: replace the Swin Transformer backbone network in the original Mask2Former architecture with a hybrid convolution-Transformer backbone network.
[0084] S602: introduce a time-series image processing network after the intermediate module of the original Mask2Former architecture and before the mask decoder module of the original Mask2Former architecture, to construct a building update type identification model based on the improved Mask2Former architecture.
[0085] Optionally, the hybrid convolution-Transformer backbone network includes: a standard convolution layer (i.e. Figure 2 a traditional convolution layer in the original Mask2Former architecture), a depth separable convolution layer, a depth convolution layer, a depth convolution + residual connection layer, a first Transformer encoder layer, a second Transformer encoder layer, a third Transformer encoder layer, and a Transformer encoder + residual connection layer.
[0086] It should be noted that the mixed convolution-Transformer backbone network adopts a four-stage hierarchical structure (Stage 1 to Stage 4), which is responsible for feature extraction and semantic modeling at different levels, and takes into account local texture details and global context information. Stage 1 includes two convolution layers (standard convolution layer, depth separable convolution layer): first, a standard 3x3 convolution kernel (stride=2) is used to extract the basic features of the image, with an output channel number of 32. Then, a depth separable convolution is introduced to further refine the extraction of local texture information, with an output channel number still being 32, and Batch Normalization (Batch Normalization) and activation functions (such as ReLU) are used for nonlinear mapping and feature standardization. Stage 2 consists of two deep convolution (3x3) layers (deep convolution layer, deep convolution + residual connection layer): the first is a standard deep convolution layer, and the second is a deep convolution layer with a residual structure. The output channel number increases from 32 to 64. The residual connection mechanism is introduced in this stage to optimize information flow, effectively alleviate the gradient vanishing problem, and enhance the modeling ability of building edge contours and detail textures. Stage 3 includes two layers of Transformer encoder (first Transformer encoder layer, second Transformer encoder layer), each consisting of multi-head self-attention mechanism (Multi-Head Attention) and feed-forward neural network (Feed-Forward Network). Among them, 4 attention heads are set, each with a dimension of 64, to capture complex spatial relationships and context dependencies in the building area and improve the model's ability to express cross-scale features. Stage 4 is responsible for deeper aggregation and semantic enhancement of features, also including two layers of Transformer encoder structure (third Transformer encoder layer and Transformer encoder + residual connection layer), and Stage 4 adds cross-layer residual connection and layer normalization mechanism compared to Stage 3, ensuring the stability of deep information transmission and improving the model's discriminability and generalization performance in complex building appearance change scenarios.
[0087] The intermediate module includes a position encoding layer, a multi-scale fusion layer, an up-sampling layer, and a 1x1 convolution layer.
[0088] The time sequence image processing network includes a self-attention module and a cross-attention module.
[0089] Further, the time sequence image processing network further includes two Add&norm integrated layers.
[0090] In a possible implementation, the time sequence image processing network is specifically used for:
[0091] The input image sequence is processed by the backbone network and the intermediate module to generate a feature map sequence of the building at different periods. For example, if the building contains street view images of t1 and t2 eras, the feature map sequence will include building feature images from two different eras (t1 and t2) and wherein, and represent the building feature images of t1 and t2 eras, respectively, R represents the pixel value range of the image, H and W represent the height and width of the image, respectively, and C represents the number of channels of the image.
[0092] The self-attention module is used to receive the image feature sequence output from the feature extraction module, to realize the fusion of local and global features by modeling the attention relationship between each spatial position in the image, thereby enhancing the semantic consistency and spatial expression ability of the image, and providing a basis for the subsequent cross-attention module for cross-time sequence comparison.
[0093] The cross-attention module is used to introduce the cross-attention mechanism to align and fuse the features of images at different times. Specifically, the trainable projection matrices W Q and W K are used to perform linear transformation on and to align them to a shared low-dimensional space: , wherein, , is the projection matrix, Q and K are the projected feature representations (representing the transformed features of the t1 and t2 era images, respectively), and d is the dimension of the mapped features.
[0094] Specifically, Q and K are input into the cross-attention module to calculate the feature weight of era t1 to era t2, and to obtain new feature map :
[0095]
[0096] wherein, (i.e., the key and value in the self-attention mechanism are the same, V is a matrix identical to K, representing "Value"), and softmax represents the softmax activation function (used to normalize the weights to ensure the weighted combination of each feature).
[0097] The difference value of Q and is quantified using L1 norm:
[0098]
[0099] wherein, D(x) represents the difference value of each spatial position, which is used to measure the local change of the image.
[0100] The feature difference value D is converted into a similarity score, reflecting the significance of the change area:
[0101]
[0102] wherein, alpha represents an adjustable hyperparameter, controlling the response sensitivity of similarity to feature difference, represents the change significance of the image area, i.e. the similarity score, and exp represents the exponential function.
[0103] Finally, the similarity matrix S is multiplied with the feature map of the year t1 to obtain the weighted feature map, which is sent to the mask Decoder module for more detailed target segmentation, and finally the mask of the building change area is obtained.
[0104] The mask decoder module includes: a pixel decoder (i.e. Figure 2 the pixel decoder in the above formula) and a Transformer decoder.
[0105] In the embodiments of the present application, the improved Mask2Former architecture constructed by introducing a hybrid convolution-Transformer backbone network and a time series image processing network significantly improves the performance of the building update type recognition model. The hybrid convolution-Transformer backbone combines the local feature extraction capability of the convolution network and the global context modeling capability of the Transformer, effectively enhancing the sensitivity of the model to building changes of different scales. The time series image processing network improves the recognition accuracy of the model for subtle changes between images of different periods through self-attention and cross-attention modules. The intermediate module ensures accurate segmentation of the building update area through multi-scale feature fusion and up-sampling. At the same time, the use of depth separable convolution and residual connection optimizes the computational efficiency and robustness of the model, making it suitable for processing large-scale street image data. The method of the present application effectively improves the accurate recognition ability of building update state, and is suitable for application scenarios such as urban building update monitoring, urban planning and building safety management.
[0106] S7: Taking the building update type recognition dataset as training data, the building update type recognition model is trained.
[0107] In the embodiment of the present application, the building update type recognition dataset with high quality and fine labeling is used as training data to train the improved recognition model, which helps the model to fully learn the feature differences and change patterns between different update types, thereby significantly improving the recognition accuracy and classification accuracy of the model for minor changes in building appearance, enhancing the generalization ability and robustness of the model in actual complex urban street scenes, and providing a solid intelligent foundation for subsequent automatic building update detection.
[0108] S8: input the street view images of the same building in different periods in the initial building group geographic information database into the trained building update type recognition model, and output the building update type recognition result.
[0109] It should be noted that the update type recognition result is structural reinforcement (such as adding support members), appearance decoration (such as exterior coating or decoration reconstruction), demolition and reconstruction (building style and outline change significantly), and construction in progress (such as scaffolding or being blocked).
[0110] In the embodiment of the present application, by inputting the street view images of the same building in different periods in the initial building group geographic information database into the trained recognition model, intelligent discrimination of the target building update type is realized, which not only can quickly distinguish the specific update states such as structural reinforcement, appearance decoration, demolition and reconstruction, and construction in progress, but also can greatly improve the efficiency and automation level of update recognition, reduce manual intervention, and help the intelligent promotion of application scenarios such as urban renewal, supervision decision, and digital building archive management.
[0111] S9: store the update type recognition result into the initial building group geographic information database to obtain the target building group geographic information database.
[0112] In the embodiment of the present application, the update type recognition result is stored into the initial building group geographic information database to obtain the target building group geographic information database, which realizes the full-range dynamic management of the building from the spatial position, image metadata to the update state. This method not only constructs a traceable time sequence building archive, but also improves the integration and retrieval efficiency of data, avoids information dispersion and repeated processing. At the same time, the dynamic update capability of the database provides high-precision data support for urban planning, building safety supervision and digital twin city construction, and significantly enhances the practicality, expansibility and decision service capability of the system.
[0113] Figure 3 is a structural schematic diagram of a building update type recognition device based on time sequence street view images provided by the embodiment of the present application. Optionally, the building update type recognition device based on time sequence street view images 410 can include a first processor 2001.
[0114] Optionally, the building update type identification device based on time-series street view images 410 can further include a memory 2002 and a transceiver 2003.
[0115] The first processor 2001 is connected with the memory 2002 and the transceiver 2003, for example, through a communication bus.
[0116] The building update type identification device based on time-series street view images 410 will be described below in detail. Figure 3 The building update type identification device based on time-series street view images 410 will be described below in detail.
[0117] The first processor 2001 is the control center of the building update type identification device based on time-series street view images 410, which can be one processor or a plurality of processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), application specific integrated circuits (ASICs), or one or more integrated circuits configured to perform the functions of the present embodiments, such as one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0118] Optionally, the first processor 2001 can perform various functions of the building update type identification device based on time-series street view images 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0119] In a specific implementation, as an example, the first processor 2001 can include one or more CPUs, such as the CPU0 and the CPU1 shown in FIG. 10. Figure 3
[0120] In a specific implementation, as an example, the building update type identification device based on time-series street view images 410 can also include a plurality of processors, such as the first processor 2001 and the second processor 2004 shown in FIG. 10. Each of these processors can be a single-CPU or a multi-CPU. The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). Figure 3
[0121] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0122] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently, and may be connected via the interface circuit of the building update type recognition device 410 based on time-series street view images. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0123] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0124] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 3 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0125] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and can be connected to the interface circuit of the building update type recognition device 410 based on time-series street view images. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0126] It should be noted that, Figure 3 The structure of the building update type identification device 410 based on time-series street view images shown does not constitute a limitation on this router. Actual building update type identification devices based on time-series street view images may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0127] In addition, the technical effects of the building update type identification device 410 based on the time-series street view images can refer to the technical effects of the building update type identification method based on the time-series street view images described in the above method embodiments, which will not be repeated here.
[0128] It should be understood that the first processor 2001 in the embodiments of the present application can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor.
[0129] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (direct rambus RAM, DR RAM).
[0130] The above-described embodiments can be implemented in whole or in part by software, hardware (e.g., circuitry), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0131] It should be understood that the term "and / or" herein merely describes an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it, but it can also represent an "and / or" relationship, which can be understood in the context before and after it.
[0132] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0133] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined by their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0134] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0135] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0136] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0137] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0138] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0139] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0140] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for identifying building update types based on temporal street view images, characterized in that, The method comprises: S1: acquiring multiple street view images containing a building group taken at different time periods and different perspectives, the building group being composed of multiple buildings; S2: inputting the street view images into a YOLO target detection model to output a target detection result of the buildings; S3: determining geographical coordinates of the buildings in geographical space according to metadata information of the street view images and the target detection result; S4: constructing an initial building group geographical information database according to the geographical coordinates in combination with a target re-identification algorithm; S5: constructing a building update type identification data set according to the initial building group geographical information database; S6: constructing a building update type identification model based on an improved Mask2Former architecture; S7: training the building update type identification model by taking the building update type identification data set as training data; S8: inputting street view images of the same building at different time periods in the initial building group geographical information database into the trained building update type identification model to output an update type identification result of the building; S9: storing the update type identification result into the initial building group geographical information database to obtain a target building group geographical information database. 2.The building update type identification method based on time-sequential street view images according to claim 1, wherein, The target detection result comprises: bounding box coordinates, a class label and a confidence of the building; The metadata information comprises: a shooting time, a shooting perspective, a shooting position latitude and longitude coordinates and a unique identifier of the street view image. 3.The building update type identification method based on time-sequential street view images according to claim 2, characterized in that, After the S2 and before the S3, the method further comprises: S2A: determining whether the confidence of the building is greater than or equal to a preset confidence; if yes, retaining the street view image containing the building; otherwise, discarding the street view image containing the building; S2B: determining whether a pixel area of the building is greater than or equal to a preset pixel area; if yes, retaining the street view image containing the building; otherwise, discarding the street view image containing the building. 4.The building update type identification method based on time-sequential street view images according to claim 2, wherein, The S3 specifically comprises: S301: determining an azimuth angle of the building in the street view image according to the shooting perspective and the bounding box coordinates; S302: determining geographical coordinates of the building in geographical space according to the shooting position latitude and longitude coordinates and the azimuth angle of the building in the street view image. 5.The building update type identification method based on time-sequential street view images according to claim 1, wherein, The S4 specifically comprises: S401: acquiring a building contour at the geographical coordinates in an open block vector map; S402: calculating an included angle between the azimuth angle of the building in the street view image and an azimuth angle of the center of the building contour relative to a shooting device; S403: determining whether the included angle is less than or equal to a preset included angle; if yes, determining that the building and the building contour are matched successfully, adding a first identifier to the building, and entering S406; otherwise, determining that the building and the building contour are not matched successfully, and entering S404; S404: clustering each building that is not matched successfully by the target re-identification algorithm, and determining a clustering result as a potential building; S405: adding a second identifier to the potential building; S406: combining the street view image corresponding to the building after adding the identifier and the metadata information to construct the initial building group geographic information database. 6.The building update type identification method based on time-sequential street view images according to claim 1, wherein, The S5 specifically includes: S501: selecting the street view images corresponding to the buildings with the same identifier in the initial building group geographic information database, and combining two street view images with a difference in longitude and latitude coordinates less than or equal to a preset difference in longitude and latitude coordinates into an image pair; S502: performing an enhancement operation on the image pair to obtain an enhanced image pair; S503: labeling the enhanced image pair; S504: constructing the building update type recognition dataset according to the labeled enhanced image pair. 7.The building update type identification method based on time-sequential street view images according to claim 1, wherein, The S6 specifically includes: S601: replacing the Swin Transformer backbone network in the original Mask2Former architecture with a hybrid convolution-Transformer backbone network; S602: introducing a time sequence image processing network after the intermediate module of the original Mask2Former architecture and before the mask decoder module of the original Mask2Former architecture to construct the building update type recognition model based on the improved Mask2Former architecture. 8.The building update type identification method based on time-sequential street view images according to claim 7, wherein, The hybrid convolution-Transformer backbone network includes: a standard convolution layer, a depth separable convolution layer, a depth convolution layer, a depth convolution + residual connection layer, a first Transformer encoder layer, a second Transformer encoder layer, a third Transformer encoder layer, and a Transformer encoder + residual connection layer; The intermediate module includes: a position encoding layer, a multi-scale fusion layer, an up-sampling layer, and a 1x1 convolution layer; The time sequence image processing network includes: a self-attention module and a cross-attention module; The mask decoder module includes: a Pixel decoder and a Transformer decoder.
9. A building update type recognition device based on time-series street view images, characterized by, The building update type recognition device based on time sequence street view images includes: a processor; a memory having computer readable instructions stored thereon, the computer readable instructions being executed by the processor to implement the method of any one of claims 1-8.
10. A computer readable storage medium, characterized in that, The computer readable storage medium stores program code that can be called and executed by the processor to implement the method of any one of claims 1-8.
Citation Information
Patent Citations
High-resolution remote sensing image target state discrimination method and device, and storage medium
CN118691877A
Urban non-fortified building reinforcement identification method and device based on historical streetscape pictures
CN119418207A
Building shape recognition and classification method and system based on YOLO and CBAM attention mechanism
CN119863692A
Remote sensing image building identification method and system based on multi-neural network integration, and electronic equipment
CN120689764A
Dynamic image recognition model updates
US20160163029A1
Cited By
Target detection neural network accelerator based on compound eye array camera
CN121415358A
Building safety feature screening method and device based on large language model
CN121527555A
A street surface update identification method fusing time-series street view data and remote sensing data
CN122368611A
A street surface update identification method fusing time-series street view data and remote sensing data
CN122368611B