Historical building intelligent identification method and system based on multi-source spatio-temporal data

By using a unified spatiotemporal grid benchmark of multi-source spatiotemporal data and a dual-stream spatiotemporal neural network, the problem of insufficient spatiotemporal fusion in the identification of historical buildings is solved, and the accurate identification of building form, structure and semantic attributes is achieved, supporting the application of digital twin cities and smart cultural heritage protection platforms.

CN121033679BActive Publication Date: 2026-03-17GUANGZHOU URBAN PLANNING & DESIGN SURVEY RES INST
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511545893.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-03-17
Estimated Expiration
2045-10-28

AI Technical Summary

Technical Problem

Existing technologies for identifying historical buildings suffer from insufficient consideration of factors, inability to perform spatiotemporal fusion analysis, inadequate accuracy and effectiveness of identification results, inability to generate spatiotemporal databases that conform to geographic information standards, and inability to support the in-depth application of digital twin cities and smart cultural heritage protection platforms.

Method used

By collecting multi-source spatiotemporal data, a unified spatiotemporal grid benchmark is established, multi-dimensional features are extracted, and a dual-stream spatiotemporal neural network is used for building identification to generate a vector format distribution map and attribute database that conforms to geographic information standards.

Benefits of technology

It improves the reliability and accuracy of the historical evolution of buildings, adapts to the changes in building forms in different regions and historical periods, supports large-scale and efficient screening, ensures the consistency of identification standards, and can be directly applied to digital twin cities and smart cultural heritage platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033679B_ABST
    Figure CN121033679B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of building identification, and specifically discloses a historical building intelligent identification method and system based on multi-source space-time data. The method collects and fuses historical maps, multi-temporal remote sensing images and laser radar point cloud data to establish a unified space-time grid reference. Then, the three-dimensional shape structure features and space-time semantic attribute features of buildings are extracted from the fused data set, and cross-modal feature deep fusion is performed by using a double-flow space-time neural network. Finally, intelligent identification and multiple verification of historical buildings are completed based on the fused features, and the structured results including a distribution map and an attribute database are output. The application effectively solves the problems of single data dimension, dependence on artificial experience and insufficient practicality of the results in the traditional method, realizes full-process automation and intelligentization of historical building identification from data perception, feature analysis to result output, and significantly improves the accuracy and efficiency of identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of building recognition technology, and more specifically, relates to a method and system for intelligent recognition of historical buildings based on multi-source spatiotemporal data. Background Technology

[0002] As the core carriers of urban cultural heritage and regional culture, the accurate identification and protection of historical buildings is the foundation of urban and rural cultural heritage management. With the development of remote sensing technology, 3D laser scanning, and digital humanities technology, the methods for identifying historical buildings are gradually shifting from purely manual field surveys to a human-machine combined, data-driven model.

[0003] Existing technologies, such as the method for identifying typical cultural genes in historical blocks in southern Fujian disclosed in Chinese invention patent application No. 202411179994.8, achieve rapid and accurate identification and classification of macro, meso, and micro cultural genes in historical blocks in southern Fujian by constructing an integrated model that integrates building feature detection, street height-to-width ratio analysis, color recognition, and plan layout judgment.

[0004] Existing technologies, such as the Chinese invention patent application with application number 202010523895.2, disclose a method, device, equipment, and storage medium for the identification and detection of historical buildings based on deep learning and high-resolution images. This method achieves automatic identification and classification of historical buildings by segmenting building images into grid cells, utilizing bounding box detection and feature value matching techniques, and combining them with a pre-trained model, thereby effectively improving the accuracy and robustness of identification.

[0005] Based on the above existing technologies, it is clear that the current tendency to use a single data source, which is limited to two-dimensional image features, still has the following problems: 1. Insufficient consideration of factors, making it impossible to construct a complete closed-loop evidence chain of architectural form evolution, material aging and functional changes, which in turn leads to certain deviations in the accuracy of the recognition results.

[0006] 2. The current focus is on the planar distribution structure characteristics, without considering the spatial dynamic characteristics caused by time evolution. That is, no spatiotemporal fusion analysis is carried out, and it is impossible to effectively capture the propagation pattern of architectural style along the street network and its correlation with the spatiotemporal context.

[0007] 3. Existing methods rely on preset thresholds or fixed rules for parameters, and do not meet the adjustment requirements such as adaptive retrieval of point cloud segmentation parameters and dynamic weight interpolation. They are not effective in recognizing architectural variations in different regions and periods.

[0008] 4. Existing technologies mainly output classification labels or detection boxes, and none of them generate spatiotemporal databases and multi-dimensional attribute libraries that conform to geographic information standards, which cannot directly support in-depth applications of platforms such as digital twin cities and smart cultural heritage protection. Summary of the Invention

[0009] In view of this, in order to solve the above problems, a method and system for intelligent identification of historical buildings based on multi-source spatiotemporal data is proposed.

[0010] The objective of this invention can be achieved through the following technical solution: This invention provides a method for intelligent identification of historical buildings based on multi-source spatiotemporal data. The method includes: S1, collecting and preprocessing multi-source spatiotemporal data of the area where the target building is located, including historical map data, multi-temporal remote sensing image data and lidar point cloud data.

[0011] S2. Establish a unified spatiotemporal grid benchmark, and map the preprocessed multi-source spatiotemporal data to the same spatiotemporal grid to form a fused dataset.

[0012] S3. Extract multidimensional features of buildings from the fused dataset. The multidimensional features include three-dimensional form and structure features and spatiotemporal semantic attribute features.

[0013] S4. The multidimensional features are fused through a dual-stream spatiotemporal neural network, and the historical buildings are identified and verified based on the fused features. The verified historical building identification results are output, including a historical building distribution map and an attribute database.

[0014] The present invention also provides a historical building intelligent identification system based on multi-source spatiotemporal data. The system includes: a data acquisition and processing module, which acquires and preprocesses multi-source spatiotemporal data of the area where the target building is located, including historical map data, multi-temporal remote sensing image data and lidar point cloud data.

[0015] The data fusion module establishes a unified spatiotemporal grid benchmark, mapping preprocessed multi-source spatiotemporal data to the same spatiotemporal grid to form a fused dataset.

[0016] The multidimensional feature extraction module extracts multidimensional features of buildings from the fused dataset. These multidimensional features include three-dimensional form and structure features and spatiotemporal semantic attribute features.

[0017] The building identification output module uses the multi-dimensional features of the dual-stream spatiotemporal neural network to identify and verify historical buildings based on fused features, and outputs the verified historical building identification results, which include a historical building distribution map and attribute database.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention integrates heterogeneous spatiotemporal data such as historical maps, multi-period remote sensing images and laser point clouds, establishes a unified spatiotemporal grid benchmark for fusion processing, and constructs a complete evidence chain covering multiple dimensions such as architectural form evolution, material aging, and functional changes. This effectively overcomes the limitations of a single data source perspective, enabling the architectural historical change process to be cross-verified and quantitatively analyzed from different spatiotemporal scales, significantly improving the reliability, accuracy and integrity of the identification conclusions and the historical context.

[0019] (2) This invention replaces the traditional visual judgment that relies on expert experience by adopting automatic comparison based on the shape feature library, adaptive retrieval of point cloud segmentation parameters and three-dimensional reconstruction algorithm. It can automatically adapt to the architectural shape variations of different regional styles and historical periods, significantly improve the recognition ability when facing multiple samples and cross-temporal and spatial scenarios, and realize the objective quantification of the three-dimensional shape structure features of buildings.

[0020] (3) This invention also places the identification of a single building in a larger settlement context through spatial distribution analysis, which greatly reduces the randomness of subjective judgment, ensures the consistency of identification standards, and supports efficient screening of a large number of historical buildings.

[0021] (4) This invention utilizes a dual-stream spatiotemporal neural network architecture to fuse three-dimensional structural features and spatiotemporal semantic attributes, fully simulating the thought process of experts comprehensively considering form and documentary evidence during the judgment process. This ensures accurate identification of historical buildings in complex scenarios, especially for buildings that have been partially renovated or have blurred features.

[0022] (5) This invention automatically generates a vector format historical building distribution map that conforms to geographic information standards and a structured database that integrates multi-dimensional attributes, so that the identification results can be directly connected to modern digital planning platforms such as digital twin cities and smart cultural heritage protection, which greatly improves the engineering practicality and digital application depth of the results. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of the overall implementation process of the present invention.

[0025] Figure 2 This is a schematic diagram of the overall implementation framework of the present invention.

[0026] Figure 3 This is a schematic diagram of the system module connections of the present invention.

[0027] Figure 4 This is a simplified schematic diagram illustrating the three-dimensional shape and structural feature extraction process of this invention.

[0028] Figure 5 This is a simplified schematic diagram illustrating the spatial distribution feature analysis process of the present invention.

[0029] Figure 6 This is a simplified schematic diagram illustrating the form feature verification process of the present invention. Detailed Implementation

[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] Please see Figure 1 and Figure 2 As shown, the present invention provides a method for intelligent identification of historical buildings based on multi-source spatiotemporal data. The method includes: S1, collecting and preprocessing multi-source spatiotemporal data of the area where the target building is located, including historical map data, multi-temporal remote sensing image data and lidar point cloud data.

[0032] Understandably, the preprocessing of the historical map data includes, but is not limited to, digitization using a high-precision scanner, and the use of image enhancement algorithms to repair damaged areas and extract symbolic annotations. The preprocessing of the multi-temporal remote sensing image data includes, but is not limited to, radiometric correction, atmospheric correction, and multi-temporal image registration. The preprocessing of the lidar point cloud data includes, but is not limited to, point cloud filtering and 3D structural feature reconstruction. Furthermore, all the preprocessing methods mentioned above utilize existing technologies, and the specific execution process will not be elaborated upon here.

[0033] S2. Establish a unified spatiotemporal grid benchmark, and map the preprocessed multi-source spatiotemporal data to the same spatiotemporal grid to form a fused dataset.

[0034] Specifically, the establishment of a spatiotemporal grid benchmark requires standardized definitions across both spatial and temporal dimensions. The spatial grid employs equally spaced latitude and longitude grids or projected coordinate system grids, such as UTM grids, with the grid resolution and coverage area determined based on data characteristics and application requirements. The temporal grid defines time slices or time series nodes based on a unified time axis or fixed time intervals such as hourly, daily, or monthly intervals. During benchmark establishment, the uniqueness of grid encoding must be ensured; for example, by combining spatial coordinates and timestamps to assign a unique ID to each grid cell for subsequent data mapping and indexing. Furthermore, the benchmark should support metadata descriptions, including grid origin, resolution, and time range, to enhance scalability and interoperability.

[0035] Furthermore, the specific process of forming the fused dataset includes: using landmark buildings as control points, geometrically correcting historical map data using an elastic registration algorithm.

[0036] Timing alignment of multi-temporal remote sensing image data is performed using a phase correlation algorithm.

[0037] Coordinate transformation and spatial resampling are performed on the lidar point cloud data to spatially align it with the remote sensing image data.

[0038] All data is uniformly converted to a standard geographic coordinate system to form a spatiotemporally consistent fused dataset.

[0039] It should be added that the algorithms and processing methods mentioned above are all existing technologies, and the specific implementation process will not be described further.

[0040] S3. Extract multidimensional features of buildings from the fused dataset. The multidimensional features include three-dimensional form and structure features and spatiotemporal semantic attribute features.

[0041] Specifically, please refer to Figure 4 As shown, the specific extraction process of the three-dimensional shape and structure features mainly covers seven steps: constructing a shape feature library, extracting and comparing the initial outline of the building, extracting spatial distribution features, retrieving point cloud segmentation parameter combinations, generating a preliminary three-dimensional model and extracting geometric features, applying spatial constraints for reconstruction, and verifying and outputting shape features. These steps are specifically demonstrated in steps A1 and A7.

[0042] A1. Collect standard geometric features of different architectural styles from different historical periods, including roof slope range, length-to-width ratio of the plan, and height ratio of the facade, and construct a style feature library.

[0043] A2. Extract the initial outline of the building from multi-temporal remote sensing images, compare the extracted outline shape features with the shape feature database, and identify the outline segments that conform to all shape features.

[0044] A3. Extract the spatial distribution characteristics of the area where the target building is located. The spatial distribution characteristics include the number of buildings and the street network.

[0045] A4. Based on the spatial distribution feature analysis results, retrieve the corresponding point cloud segmentation parameter combination.

[0046] In this process, point cloud segmentation mainly adjusts multiple segmentation parameters of the point cloud segmentation synchronously based on the identified building form type. These segmentation parameters include, but are not limited to, the horizontal constraint angle of the normal vector, the point spacing threshold, and the growth threshold.

[0047] Based on empirical data, it can be observed that for gable-roof buildings, the normal vector horizontal constraint angle is typically set to 0-5 degrees, the point spacing threshold is 6-8 cm, and the growth threshold is 0.15. For hip-roof buildings, the normal vector constraint angle is typically set to 5-30 degrees, the point spacing threshold is 5-7 cm, and the growth threshold is 0.2. For palace-style buildings, the normal vector constraint angle is typically set to 30-60 degrees, the point spacing threshold is 8-10 cm, and the growth threshold is 0.25.

[0048] Furthermore, the normal vector constraint angle expands with increasing stylistic complexity, the point spacing threshold is negatively correlated with building density (the higher the density, the smaller the threshold), and the growth threshold is negatively correlated with structural complexity, while also considering chronological characteristics. For example, Qing dynasty buildings, due to their elaborate decoration, use lower thresholds, while Ming dynasty buildings, due to their simpler structure, can use higher thresholds. Taking a gable-roofed building as an example, the values ​​are shown below: For gable-roofed buildings, in high-density blocks, under multiple stylistic influences, and in the mid-Qing dynasty, the normal vector constraint angle is typically set at 0-8 degrees, the point spacing threshold at 5 cm, and the growth threshold at 0.12. In low-density blocks, under a single style, and in the early Ming dynasty, the normal vector constraint angle is typically set at 0-4 degrees, the point spacing threshold at 7 cm, and the growth threshold at 0.16.

[0049] In this embodiment, taking a hard-gable-style building in a certain historical block as an example, the building is located at the main intersection of a high-density traditional block, with a building spacing of less than 10 meters. Furthermore, spatial distribution feature analysis shows a similarity exceeding 80% with mid-Qing Dynasty buildings, thus confirming it as a mid-Qing Dynasty building. Specifically, for the example area, a building spacing of less than 10 meters indicates the building is located in a high-density traditional block. The matching point cloud segmentation parameter combination is: normal vector constraint angle of 0-8 degrees, point spacing threshold of 5 cm, and growth threshold of 0.12.

[0050] Further, please refer to Figure 5 As shown, the spatial distribution feature analysis includes: A4-1, extracting the axis map of the street and alley network, using spatial syntax to calculate the global integration degree and local connectivity degree of each street and alley, and outputting the network influence factor of each street and alley through linear weighted summation.

[0051] In practice, the extraction of street and alley network axis maps can be achieved using professional software such as Depthmap, QGIS with Axwoman plugin, or ArcGIS AxialMap tool, based on high-precision maps or satellite imagery. The axes of all streets and alleys within the study area can be drawn manually or semi-automatically. A line segment network map representing the spatial traffic potential of the streets and alleys is then generated, serving as the street and alley network axis map.

[0052] It is important to note that when a street has a significant bend, it should be cut off at the bend and represented as two intersecting axes to accurately reflect the change in traffic path. All axes should be cut off at intersections. For elevated roads, underpasses, and other grade-separated transportation systems, the need for connection should be determined based on the actual traffic flow; generally, grade-separated roads without interchanges are not considered connected.

[0053] Higher overall integration indicates a central location within the network, making it easier to reach other areas and typically indicating higher pedestrian and vehicular traffic and commercial activity. Local connectivity reflects the street's connectivity and accessibility within a specific local area. Streets with high connectivity are often the heart of local life.

[0054] Understandably, the specific implementation method operates as follows: First, using spatial syntax software such as Depthmap, the raw values ​​of global integration and local connectivity for each axis are calculated, and then min-max normalization is performed on them respectively. Simultaneously, weights are set according to the research objectives and the principle that the dissemination of historical architectural styles is mainly dominated by macroscopic street and alley structures. For example, the global integration weight can be set to 0.6, and the local connectivity weight to 0.4.

[0055] A4-2. Compare remote sensing images from different periods, calculate the annual rate of change in the number of buildings in each street and alley, and record the main orientation and average speed of the expansion of the building complex.

[0056] Understandably, the main direction of expansion can be determined by comparing the newly added building complex with earlier imagery to establish its spatial distribution relative to the core direction of the original building complex, such as primarily expanding eastward and southeastward. In practice, this can be determined by calculating the direction of the centroid of the new buildings or by observing their relative positions to streets, alleys, and natural features. The average speed of advancement can be obtained by measuring the farthest vertical distance from the edge of the original building complex to the leading edge of the new building complex along the determined main expansion direction, and then dividing by the time interval.

[0057] In one specific embodiment, taking Wangfujing Street as an example, based on remote sensing images from 2020 and 2023, buildings within the polygonal area of ​​the street were identified using image classification technology. The count revealed 155 buildings in 2020 and 168 buildings in 2023, resulting in an annual change rate of approximately 2.8%. Furthermore, by overlaying and comparing the building outlines from the two periods, it was found that the 13 newly added buildings were mainly concentrated at the northern end of the street, exhibiting a continuous, south-to-northward filling expansion pattern. Based on this, the main direction of expansion was recorded as northward, and the distance from the northern boundary of the buildings in 2020 to the northernmost new building in 2023 was measured to be 150 meters, calculating an average advancement speed of 50 meters per year.

[0058] A4-3. Based on the annual change rate of all streets and alleys, calculate the average change rate, and use the ratio of the annual change rate to the average change rate as the vitality coefficient. Identify the main channels for style dissemination based on the main orientation of the building complex expansion.

[0059] When identifying the main channels of style dissemination, the first step is to summarize the analysis results of the main directional expansion of building clusters in all streets and alleys. Subsequently, adjacent streets and alleys with the same or similar expansion directions are connected and merged in the GIS platform to form continuous, directional corridors. For example, if the analysis reveals that streets B, C, and D are sequentially adjacent and their expansion directions are northeast, east-northeast, and northeast respectively, they can be merged and identified as a unified northeast-southwest main channel. This main channel represents the main direction of construction activities and spatial development within the area. We assume it to be the physical path through which new buildings are most likely to imitate and spread the original style, and it serves as the carrier for subsequent calculations of dissemination attenuation.

[0060] A4-4. Set up sampling points at preset intervals along the extended trajectory of the main channel, and obtain the network influence factor and vitality coefficient of the street where each sampling point is located.

[0061] The preset interval can be a fixed distance such as 50 meters or 100 meters, or a fixed number of street segments at every intersection.

[0062] A4-5. Based on the network influence factor and vitality coefficient, establish an exponential decay model that includes the baseline decay rate, network influence adjustment coefficient, and vitality instability coefficient, and calculate the propagation decay rate.

[0063] This step aims to quantify the rate at which style influence diminishes with distance along the main channel. The implementation is as follows: Along the identified main channel, starting from the style source point, sampling points are set at regular intervals, such as 100 meters. For each sampling point, the network influence factor and vitality coefficient of the street / alley where it is located are used as independent variables to establish an exponential decay model: The propagation attenuation rate is calculated using this model.

[0064] in, This represents the baseline attenuation rate, with a value ranging from 0.08 to 0.12 per meter. This is the network influence moderating coefficient, determined through historical data fitting or expert experience, and its value ranges from 1.5 to 2.5. Network Impact Factor The vitality coefficient, The vitality instability coefficient, with a value ranging from 0.15 to 0.25, is used to quantify the positive correlation between vitality deviation from a steady state and decay rate. To ensure smoothness and avoid a zero denominator, the value is between 0.05 and 0.15. It is a natural exponential function. This represents the absolute deviation of the vitality coefficient from the ideal stable value of 1.

[0065] The item reflects the negative correlation effect of the network impact factor, namely The larger the value, the more complete the street and alley network, the smaller the attenuation rate, and the less resistance to style propagation. The item reflects the stabilizing effect of the vitality coefficient. The larger the vitality coefficient V deviates from 1, the more unstable the development of the neighborhood is, such as over-development or decline. The larger the value of this item, the greater the resistance to style dissemination and the more limited the dissemination.

[0066] A4-6. Based on the remote sensing image, identify style feature source points and calculate the matching degree between the target building and the style feature source points in three features: roof form, facade decoration, and building materials. The matching degree is quantified by the ratio of the number of matching features to the total number of features.

[0067] Understandably, a stylistic source point refers to a building with a pure form, a definite age, and good preservation, which can be selected based on historical documents, local chronicles, or authoritative surveying data. This building is the stylistic source point. For example, when analyzing a Qing Dynasty commercial district, the oldest and most standardized money exchange or bank building in the district can be designated as the source point. When analyzing a residential area next to it, the overall geometric center of a typical courtyard-style residential area or its most complete building can be designated as the source point.

[0068] It is important to note that in practical implementation, roof types may involve shape classification, such as hip roofs and gable roofs, and then shape similarity algorithms are used. Facade decoration may utilize image matching, which can be achieved using image recognition algorithms, while building materials involve texture analysis, which can be achieved using material property classification and recognition algorithms. For different feature dimensions, corresponding matching algorithms are used sequentially for matching calculations. All algorithms involved utilize existing technologies and will not be further elaborated upon.

[0069] A4-7. The matching degree of each feature is linearly weighted and summed to obtain the style similarity benchmark value, and the actual spatial distance between the target building and the source point of the style feature is measured.

[0070] The weighting process employs a combination of subjective and objective weighting methods: First, the analytic hierarchy process (AHP) is used to synthesize expert opinions and determine the theoretical importance of roof form, facade decoration, and materials in style assessment, forming subjective weights. Then, the entropy weighting method is used to analyze the discriminative power of each feature in the actual sample data, deriving objective weights. Finally, a preset preference coefficient linearly integrates the subjective and objective weights into a final combined weight. This weighting system can be dynamically adjusted according to specific building types and assessment scenarios, ensuring that the matching degree calculation is both professionally instructive and data-adaptable.

[0071] Preferably, the actual distance measurement can be based on a unified geographic coordinate system by calculating the three-dimensional Euclidean distance or geodetic distance between the geometric center of the target building and the source point of the style features to obtain its accurate actual spatial distance.

[0072] A4-8. Multiply the propagation attenuation rate by the actual spatial distance to obtain the attenuation intensity, and import it into the exponential attenuation function to output the attenuation factor.

[0073] This step aims to quantify the style influence loss caused by the combined effects of distance and decay rate. The exponential decay function uses an existing basic exponential function, where the decay coefficient adjusts the severity of the decay and can be calibrated according to the specific situation, for example, set to 0.1. When the decay intensity is 0, the decay factor is 1, indicating no decay. As the decay intensity increases, the decay factor decreases from 1 to 0. The decay factor directly represents the proportion of style similarity retained purely due to spatial distance and propagation path characteristics.

[0074] A4-9. Multiply the style similarity benchmark value by the attenuation factor to output the final style similarity.

[0075] A5. Combining lidar point cloud data and point cloud segmentation parameters, a preliminary 3D model is generated using a 3D reconstruction algorithm based on the construction rules of historical buildings. Geometric features of the building outline are then extracted based on the preliminary 3D model.

[0076] Specifically, the construction rules are matched and retrieved according to the historical style of the target building. For example, for walls: the segmented wall point cloud is fitted by plane, and then verticality constraints and coplanarity constraints are applied to generate a flat and vertical wall model.

[0077] For roofs: The segmented roof point cloud is fitted with a plane or quadric surface, and constraints such as ridge line horizontal constraint and roof slope consistency constraint are applied to generate regular double-slope, quadrilateral, or dome models.

[0078] For linear components such as columns and beams: the corresponding point cloud is fitted using a cylindrical or cuboid model to ensure that its axis is straight and its cross-section is uniform.

[0079] The specific process of generating the preliminary 3D model includes: First, under the guidance of geometric constraints, point cloud components are converted into accurate triangular mesh models or boundary representation models through parametric modeling operations such as least squares fitting, boundary contour extraction, stretching, rotation, and lofting. Then, the topological connections between components are automatically established, such as the intersection lines of the roof and walls, and the connection points of beams and columns, finally forming a complete preliminary 3D model. All the aforementioned operations utilize existing technologies and will not be elaborated upon here.

[0080] A6. Use the geometric features of the building outline as the spatial constraints for 3D reconstruction. During the point cloud surface reconstruction process, use the outline boundary as the boundary constraint for generating the triangular mesh.

[0081] A7. Verify the shape features of the reconstructed 3D model. Once the verification is successful, output the 3D shape and structure features that integrate spatial distribution features.

[0082] Further, please refer to Figure 6 As shown, the specific verification process of the shape feature verification includes: A71, calculating the relative deviation value between the contour shape feature of the reconstructed three-dimensional model and the corresponding feature in the shape feature library, and performing minimum-maximum normalization processing on the deviation value.

[0083] A72. Identify the forefront of style feature source points from multi-temporal remote sensing image data and generate architectural style propagation trajectories along the street network.

[0084] A73. Establish a building cluster expansion direction buffer in the GIS system, and overlay the architectural style propagation trajectory within the buffer.

[0085] A74. Calculate the coverage area ratio of the propagation trajectory within the expansion direction buffer using the area feature intersection analysis technique, and match the corresponding expansion direction weights based on the ratio.

[0086] Preferably, the matching example process for weighting the expansion direction based on empirical statistics is as follows: when the coverage area ratio is greater than 80%, it is considered high overlap, and the expansion direction weight can be 1.0. When the coverage area ratio is between 50% and 80%, it is considered moderate overlap, and the expansion direction weight can be 0.8. When the coverage area ratio is less than 50%, it is considered low overlap, and the expansion direction weight can be 0.5.

[0087] A75. Calculate the number of periods that differ between the target building period and the period with the highest concentration of buildings within a preset range, and match a preset age difference weight based on the number of periods that differ.

[0088] The specific implementation method based on matching the number of difference periods with the preset age difference weight is as follows: using an exponential decay function. Dynamically calculate the weight of age differences, where The baseline weight ranges from 0.9 to 1. The preset attenuation coefficient ranges from 0.2 to 0.5. The specific value should be selected based on the actual scenario requirements and historical experience. This represents the number of periods that differ between the target construction period and the period with the highest concentration of buildings within the preset range. is a natural constant. This function maps the differences in the number of periods to continuous weight values ​​in the interval [0.3, 0.9], so that buildings from similar periods receive higher weights, while the weights of buildings from different periods decrease significantly in an exponential manner, thereby quantifying the influence of the proximity of ages on the verification of morphological characteristics.

[0089] A76. Centered on the target building, match the preset distance weights based on the distances between each building within the preset range and the target building, and take the average distance weight as the final distance weight.

[0090] A77. The product of the weights of age difference, style similarity, spatial distance, and expansion direction is used as the spatiotemporal coordination coefficient.

[0091] The weighting of age difference, style similarity, spatial distance and expansion direction fully considers spatiotemporal consistency. The spatiotemporal coordination coefficient actually introduces spatiotemporal consistency constraints, which not only consider the static layout of the current block, but also its evolution process, development logic and inherent laws in multi-temporal images.

[0092] A78. Based on the spatiotemporal coordination coefficient, the weight ratio of each feature is dynamically adjusted using linear interpolation, and the final feature matching degree is obtained by linear weighted summation.

[0093] Preferably, when the spatial-temporal coordination coefficient is in the range of 0.8 to 1, the feature weights for the roof slope range are 0.4 to 0.5, the feature weights for the plan length-to-width ratio are 0.3 to 0.4, and the feature weights for the facade height ratio are 0.2 to 0.3. When the spatial-temporal coordination coefficient is in the range of 0.3 to 0.8, the feature weights are calculated using linear interpolation. When the spatial-temporal coordination coefficient is in the range of 0 to 0.3, the feature weights for the roof slope range are 0.2 to 0.3, the feature weights for the plan length-to-width ratio are 0.3 to 0.4, and the feature weights for the facade height ratio are 0.4 to 0.5. The linear interpolation method is an existing weight allocation method and will not be specifically described here.

[0094] A79. If the feature matching degree exceeds the preset verification threshold, the form feature verification passes; otherwise, the form feature verification fails. The preset requirement can be set to an empirical value for the feature matching degree.

[0095] It should be noted that the verification threshold is typically set between 0.7 and 0.85. In practice, it can be dynamically configured based on the different requirements for precision and recall in the identification task. For example, a higher threshold of 0.80 to 0.85 is used for high-value areas or key verification scenarios requiring extremely high accuracy; while a threshold of 0.7 to 0.75 is used to balance efficiency for large-scale general identification or scenarios allowing for subsequent manual review. This threshold is determined based on the verification results of historical labeled datasets, using ROC curve analysis to select the optimal classification performance range.

[0096] Specifically, the process of extracting the spatiotemporal semantic attribute features includes: E1, extracting handwritten text annotation information from historical maps using optical character recognition technology, and identifying building names, construction years, and building function semantic entities.

[0097] In practice, extraction can be performed using a pre-trained BERT+BiLSTM-CRF named entity recognition model.

[0098] E2. Perform spectral analysis on multi-temporal remote sensing images to extract the temporal characteristics of building material aging and vegetation cover changes.

[0099] The spectral analysis process is as follows: First, all images are radiometrically calibrated and atmospherically corrected to ensure the temporal comparability of spectral reflectance data. For monitoring the aging of building materials, the focus is on building rooftops and facades. The temporal changes in reflectance to visible and short-wave infrared bands are calculated, and specific spectral indices such as the corrosion index are used to quantify aging characteristics such as decreased gloss and compositional alteration. Simultaneously, pixel-level temporal NDVI analysis is used to extract the interannual and seasonal variations in vegetation cover and growth status in streets and alleys. Finally, the spectral indicators of building aging and vegetation indices are temporally coupled and analyzed to output the temporal characteristics of vegetation cover changes, thereby revealing the synergistic relationship between architectural style changes and ecological environment evolution within the study area.

[0100] E3. The identified architectural attributes are verified spatiotemporally with the database of architectural features and historical documents. The verified semantic attributes are converted into structured feature vectors, normalized and dimensionality reduced to generate unified spatiotemporal semantic attribute features.

[0101] Specifically, the structured feature vectors are of four types: time feature vectors containing time codes of the initial construction date and the dates of important renovations; function feature vectors using one-hot encoding to represent the main function types and functional change sequences; cultural association vectors with scores of the intensity of association with important historical events and figures; and spatiotemporal context vectors with relative codes of functional status and chronological status in the neighborhood.

[0102] In practice, the validated structured feature vectors are first processed using a min-max normalization method to eliminate differences in the dimensions and numerical ranges of different semantic attributes such as building age, architectural style, and decoration level, mapping them to the [0, 1] interval to ensure fairness in subsequent analysis. Then, for features that may exhibit high correlation, such as roof form and eaves style, dimensionality reduction algorithms such as principal component analysis are applied. While preserving most of the original information, the high-dimensional feature space is projected onto a low-dimensional, linearly independent principal component subspace, thereby generating a compact, unified set of spatiotemporal semantic attribute features that are easily processed by machine learning models. The specific algorithms involved are all existing algorithms and will not be described further.

[0103] S4. The multidimensional features are fused through a dual-stream spatiotemporal neural network, and the historical buildings are identified and verified based on the fused features. The verified historical building identification results are output, including a historical building distribution map and an attribute database.

[0104] Specifically, the dual-stream spatiotemporal neural network adopts a parallel processing architecture, including a historical map semantic processing stream and a remote sensing image feature processing stream.

[0105] The semantic stream receives structured spatiotemporal semantic attribute feature vectors after normalization and dimensionality reduction, and performs semantic encoding and temporal modeling.

[0106] In practice, the historical map semantic flow is integrated with temporal semantic association and multi-source verification constraints. First, semantic entities such as building names, construction years and functions are extracted based on optical character recognition technology. Then, text semantic encoding is completed by combining the BERT fine-tuning model. At the same time, the LSTM network is used to model temporal features such as the rate of evolution of building functions and the material aging index to form a basic semantic-temporal feature vector.

[0107] The feature stream receives building outlines from remote sensing images, lidar point clouds, and regional spatial distribution features, and then performs feature enhancement.

[0108] In practice, the Canny edge detection algorithm is first used to extract the initial building outline from the remote sensing image. The outline is then simplified using the Douglas-Peucker algorithm, with a threshold set at 2% of the outline perimeter. Control points are inserted at curvature extrema, and local fitting is performed using moving least squares with a fitting error less than or equal to 3 pixels. Finally, cubic B-spline curves are applied for smooth interpolation, while right-angle features are simultaneously... to Angle constraints were imposed, and parallel constraints of less than or equal to 2 pixels were applied to parallel edges. Geometric features unique to historical buildings were optimized by matching templates in a form feature library. This optimization was repeated up to 10 times until the Hausdorff distance between the contour and point cloud data improved by less than 1%. Then, based on LiDAR point cloud data, a region growing segmentation algorithm was used to separate individual buildings. A 3D model was generated through Poisson surface reconstruction, and geometric parameters such as roof slope and facade height were extracted. Next, a building spatial adjacency matrix was constructed, and a 3-layer graph convolutional network with a hidden layer dimension of 64 and a ReLU activation function was used to encode spatial topological relationships. Finally, spatial attention weights were calculated by linear weighted summation based on street integration and building density. Importance coefficients were generated using the Sigmoid function. Features in high-importance regions with coefficients greater than or equal to 0.7 were enhanced by 3 times, features in low-importance regions with coefficients less than or equal to 0.3 were attenuated by 0.5 times, and other regions were left unprocessed, completing the feature enhancement process for the feature flow.

[0109] In the linear weighting method, it is preferable to set the weight of street integration degree to 0.6 and the weight of building density to 0.4 to highlight the importance of the whole.

[0110] Deep fusion of dual-stream features through a cross-modal attention mechanism.

[0111] Furthermore, the cross-modal attention mechanism includes: constructing a three-dimensional shape and structure feature vector based on the three-dimensional shape and structure features fused with spatial distribution features.

[0112] L2 normalization was performed on the feature vectors of the three-dimensional shape structure and the feature vectors of spatiotemporal semantic attributes.

[0113] The dot product of two normalized feature vectors is calculated as the similarity, and the similarity is converted into an attention weight distribution by the softmax function.

[0114] Multiply each dimension feature value in the semantic feature vector by the corresponding attention weight, and output the reweighted feature value of each dimension.

[0115] Specifically, the process of identifying and verifying historical buildings includes: R1, calculating the probability value of each building belonging to a historical building through a multilayer perceptron based on the fused features, and generating a probability distribution map of historical buildings based on the probability values.

[0116] In practice, the fused feature vectors can be input into the input layer of a multilayer perceptron (MLP) network, and normalization can be used to eliminate the dimensional differences between features of different dimensions. The hidden layer adopts a 2-3 layer fully connected structure. The first layer uses the ReLU activation function to perform non-linear mapping on the fused features, selecting key features that conform to historical form and semantics and have been verified by multiple sources. The second layer introduces a Dropout layer, where the dropout rate is set to 0.3 to 0.5 to suppress overfitting and preserve the generalization ability of features. At the same time, BatchNormalization is used to accelerate training and stabilize the feature distribution. The output layer uses the Sigmoid activation function to map the features processed by the hidden layer to the interval [0, 1], and calculates the probability value of each building belonging to historical buildings. When the probability value is greater than or equal to 0.8, it is initially judged as a high-confidence candidate for historical buildings. The range of 0.5 to 0.8 requires further manual verification in conjunction with historical documents. If it is less than 0.5, it is judged as a non-historical building. The entire process uses the cross-entropy loss function to iteratively optimize the network parameters on the historical building annotation dataset to ensure the accuracy and reliability of probability calculation.

[0117] R2. Near-infrared spectral analysis is performed on buildings whose probability exceeds a set threshold in the probability distribution map to extract the spectral characteristics of the building materials.

[0118] The threshold is typically set to 0.8, and this 0.8 is not a fixed value; it can be adjusted based on the actual application.

[0119] For specific near-infrared spectroscopy analysis, an ASD FieldSpec4 spectrometer can be used to collect spectral data of building facade materials in the 900-2500 nm wavelength range. The spectrometer is set at a wavelength of 700 nm, with a spectral resolution of 3 nm and a sampling interval of 1 nm. Ten spectral curves are collected from five typical areas of each building. Dark current correction is used to eliminate instrument noise, and radiometric calibration is performed using a standard white board. Savitzky-Golay filtering is applied to obtain smooth spectral curves with a window size of 9. Finally, first-order derivative transformation is used to enhance spectral features, extracting 12 spectral characteristic parameters, including absorption depth, full width at half maximum (FWHM), and asymmetry, in characteristic bands such as 1350 nm, 1650 nm, and 2200 nm, forming a material spectral feature vector for each building.

[0120] R3. Compare and verify the spectral characteristics with a database of typical materials for historical buildings, and check the building's age information in conjunction with historical records.

[0121] In practice, age clues related to the target building can be extracted. The age information extracted from the literature can be cross-checked with the preliminary age range obtained by spectral material verification and form feature matching. If the age recorded in the literature falls within the preliminary age range, the age information is determined to be consistent. Otherwise, more complementary literature should be consulted for further determination.

[0122] For example, if the document records 1895, and the initial date range is 1880 to 1910, then the date information is considered consistent. If the document records 1920, and the initial date range is 1880 to 1900, then further consultation with cadastral archives, construction logs, or carbon-14 dating data of building components from the same period is needed to correct for any date discrepancies and finally determine the accurate range of the building's date. This accurate range is then used to determine whether the date information is consistent.

[0123] R4. If the spectral characteristics match the database of typical materials for historical buildings, and the dates recorded in the literature are consistent, generate a vector format distribution map that conforms to geographic information standards, and establish an attribute database that includes basic building attributes, form characteristics, and historical value assessment results.

[0124] Please see Figure 3 As shown, the present invention provides a historical building intelligent identification system based on multi-source spatiotemporal data. The system includes: a data acquisition and processing module, a data fusion module, a multi-dimensional feature extraction module, and a multi-dimensional feature extraction module.

[0125] In the above, the data fusion module is connected to the data acquisition and processing module and the multidimensional feature extraction module, respectively.

[0126] The data acquisition and processing module collects and preprocesses multi-source spatiotemporal data of the area where the target building is located, including historical map data, multi-temporal remote sensing image data, and lidar point cloud data.

[0127] The data fusion module establishes a unified spatiotemporal grid benchmark, maps preprocessed multi-source spatiotemporal data to the same spatiotemporal grid, and forms a fused dataset.

[0128] The multidimensional feature extraction module includes three-dimensional shape and structure features and spatiotemporal semantic attribute features.

[0129] The building identification output module fuses the multi-dimensional features through a dual-stream spatiotemporal neural network, performs historical building identification and verification based on the fused features, and outputs the verified historical building identification results, which include a historical building distribution map and an attribute database.

[0130] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.

Claims

1. A method for intelligent recognition of historical buildings based on multi-source spatio-temporal data, characterized in that, The method comprises: S1, collecting and preprocessing multi-source spatio-temporal data of the area where the target building is located, including historical map data, multi-temporal remote sensing image data and laser radar point cloud data; S2, establishing a unified spatio-temporal grid reference, mapping the preprocessed multi-source spatio-temporal data to the same spatio-temporal grid to form a fusion data set; S3, extracting multi-dimensional features of the building from the fusion data set, the multi-dimensional features including three-dimensional shape structure features and spatio-temporal semantic attribute features; The specific extraction process of the three-dimensional shape structure features comprises: Collecting standard shape geometric features of different building shape types in different historical periods, including roof slope range, plane length-width ratio and facade height ratio, to construct a shape feature library; Extracting the initial contour of the building from the multi-temporal remote sensing image, comparing the extracted contour shape feature with the shape feature library, and identifying the contour section conforming to all shape features; Extracting the spatial distribution features of the area where the target building is located, the spatial distribution features including the number of buildings and street network; Based on the analysis result of the spatial distribution features, the corresponding point cloud segmentation parameter combination is retrieved; Combining the laser radar point cloud data and the point cloud segmentation parameter combination, a preliminary three-dimensional model is generated by a three-dimensional reconstruction algorithm based on historical building construction rules, and the geometric features of the building contour are extracted based on the preliminary three-dimensional model; The geometric features of the building contour are used as the spatial constraint conditions for three-dimensional reconstruction, and the contour boundary is used as the boundary constraint for triangular mesh generation in the point cloud surface reconstruction process; The reconstructed three-dimensional model is subjected to shape feature verification, and the three-dimensional shape structure features with fused spatial distribution features are outputted after verification; The extraction process of the spatio-temporal semantic attribute features comprises: Extracting handwritten text annotation information in the historical map by optical character recognition technology to identify building name, construction year and building function semantic entities; Performing spectral analysis on the multi-temporal remote sensing image to extract building material aging degree and vegetation coverage change timing features; The identified building attributes are subjected to spatio-temporal logical verification with the shape feature library and historical documents, the verified semantic attributes are converted into a structured feature vector, normalized and dimensionally reduced to generate unified spatio-temporal semantic attribute features; S4, fusing the multi-dimensional features by a double-flow spatio-temporal neural network, identifying and verifying the historical building based on the fused features, and outputting the verified historical building identification result, the result including a historical building distribution map and an attribute database.

2. The historical building intelligent recognition method based on multi-source spatio-temporal data according to claim 1, characterized in that: The specific formation process of the fusion data set comprises: Using landmark buildings as control points, performing geometric correction on historical map data by an elastic registration algorithm; Aligning the multi-temporal remote sensing image data in time sequence by a phase correlation algorithm; Performing coordinate conversion and spatial resampling on the laser radar point cloud data to align it with the remote sensing image data in space; Converting all data to a standard geographic coordinate system to form a spatio-temporally consistent fusion data set. 3.The method of claim 1, wherein: The spatial distribution feature analysis comprises: Extracting the axis graph of the street network, calculating the global integration degree and local connection degree of each street by spatial syntax, and outputting the network influence factor of each street by linear weighted summation; Calculate the annual change rate of the number of buildings in each street and lane by comparing remote sensing images of different periods, and record the main direction and average advancing speed of the expansion of the building group; Calculate the average change rate based on the annual change rate of all streets and lanes, take the ratio of the annual change rate to the average change rate as the vitality coefficient, and identify the main channel of style propagation based on the main direction of the expansion of the building group; Set sampling points along the expansion track at preset intervals in the main channel, obtain the network influence factor and vitality coefficient of the street and lane where each sampling point is located; Based on the network influence factor and vitality coefficient, an exponential decay model containing a benchmark decay rate, a network influence adjustment coefficient, and a vitality instability coefficient is established, and a propagation decay rate is calculated; Based on the remote sensing image, identify the style feature source point, calculate the matching degree of the target building and the style feature source point in the roof form, facade decoration, and building material, and the matching degree is quantified by the ratio of the number of matching features to the total number of features; Linearly weight and sum the matching degrees of each feature to obtain a style similarity benchmark value, and measure the actual spatial distance between the target building and the style feature source point; Multiply the propagation decay rate by the actual spatial distance to obtain the decay intensity, and import it into the exponential decay function to output the decay factor; Multiply the style similarity benchmark value by the decay factor to output the final style similarity. 4.The method of claim 1, wherein: The specific verification process of the shape feature verification includes: Calculate the relative deviation value of the contour shape feature of the reconstructed three-dimensional model and the corresponding feature of the shape feature library, and normalize the deviation value; Identify the front of the style feature source point from the multi-temporal remote sensing image data, and generate the building style propagation track along the street and lane network; Establish a building group expansion direction buffer zone in the GIS system, and superimpose the building style propagation track in the buffer zone; Calculate the coverage area ratio of the propagation track in the expansion direction buffer zone by face element intersection analysis technology, and match the corresponding expansion direction weight based on the ratio; Calculate the difference period number between the target building period and the period with the most concentrated number of buildings within a preset range, and match a preset age difference weight based on the difference period number; Based on the distance between each building within a preset range and the target building, match a preset distance weight, and take the average distance weight as the final distance weight; Take the product of the age difference weight, the style similarity, the spatial distance weight, and the expansion direction weight as the spatiotemporal coordination coefficient; According to the spatiotemporal coordination coefficient, dynamically adjust the weight ratio of each feature by linear interpolation method, and calculate the final feature matching degree by linear weighted summation; If the feature matching degree exceeds the preset verification threshold, the shape feature verification is passed, otherwise the shape feature verification is not passed.

5. The historical building intelligent recognition method based on multi-source spatio-temporal data according to claim 1, characterized in that: The double-flow spatiotemporal neural network adopts a parallel processing architecture, including a historical map semantic processing flow and a remote sensing image feature processing flow; The semantic processing flow receives normalized and dimensionally reduced structured spatiotemporal semantic attribute feature vectors, and performs semantic encoding and time series modeling; The feature processing flow receives remote sensing image building contours, laser radar point clouds, and regional spatial distribution features, and performs feature enhancement; Deep fusion of dual-flow features is performed through a cross-modal attention mechanism.

6. The historical building intelligent recognition method based on multi-source spatio-temporal data according to claim 5, characterized in that: The cross-modal attention mechanism comprises: a three-dimensional shape structure feature vector is constructed based on the fused spatial distribution feature and the three-dimensional shape structure feature; L2 normalization is performed on the three-dimensional shape structure feature vector and the spatio-temporal semantic attribute feature vector respectively; the dot product of the two normalized feature vectors is calculated as a similarity, and the similarity is converted into an attention weight distribution through a softmax function; each dimension feature value in the semantic feature vector is multiplied by the corresponding attention weight, and each dimension feature value after reweighting is output.

7. The historical building intelligent recognition method based on multi-source spatio-temporal data according to claim 1, characterized in that: The specific process of identifying and verifying the historical building comprises: based on the fused features, a multi-layer perception network is used to calculate the probability value of each building belonging to a historical building, and a historical building probability distribution map is generated based on the probability value; near-infrared spectrum analysis is performed on the buildings in the probability distribution map whose probability exceeds a set threshold, and spectral features of the building materials are extracted; the spectral features are compared and verified with a historical building typical material database, and the building age information is checked in combination with historical records; if the spectral features are consistent with the historical building typical material database and the literature record age check is consistent, a vector format distribution map conforming to the geographic information standard is generated, and an attribute database containing building basic attributes, shape features, and historical value evaluation results is established. 8.The method of claim 1, wherein: The method also applies a historical building intelligent identification system based on multi-source spatio-temporal data when executed in detail, and the system comprises: a data acquisition and processing module that acquires and pre-processes multi-source spatio-temporal data of the area where the target building is located, including historical map data, multi-temporal remote sensing image data, and laser radar point cloud data; a data fusion module that establishes a unified spatio-temporal grid reference, maps the pre-processed multi-source spatio-temporal data to the same spatio-temporal grid, and forms a fused data set; a multi-dimensional feature extraction module, wherein the multi-dimensional features include three-dimensional shape structure features and spatio-temporal semantic attribute features; a building identification output module that fuses the multi-dimensional features through a dual-flow spatio-temporal neural network, identifies and verifies historical buildings based on the fused features, and outputs the verified historical building identification results, wherein the results include a historical building distribution map and an attribute database.

Citation Information

Patent Citations

  • Historical building identification and detection method based on deep learning and high-resolution images

    CN111709346A

  • Southern Fujian historical block typical culture gene identification method

    CN119169624A

  • Intelligent recognition system for urban and rural building styles and features based on multi-source data fusion

    CN120472353A

  • Remote sensing image building extraction method fusing semantic and edge features

    CN120635710A