A system and method for health prognosis of historic buildings

By performing initial regional unit division, regional unit coupling matrix construction, and health prediction on historical buildings, the problem of insufficient consideration of regional health differences and coupling effects within historical buildings was solved, enabling the location and interpretation of regional risks and improving the accuracy and interpretability of predictions.

CN121599239BActive Publication Date: 2026-05-19LONGYAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LONGYAN UNIV
Filing Date
2026-01-28
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies are insufficient to accurately reflect the differences in health status between different areas inside historical buildings, and do not fully consider the influence of material and structural transmission coupling between areas, resulting in reduced accuracy of health assessment results.

Method used

The system employs a historical building initial area unit division module, an area unit determination module, an area unit coupling matrix construction module, and a health prediction module. By acquiring the original information set, it constructs an initial area unit set, calculates the influence weights between area units, constructs a neighborhood aggregation representation of area units, and performs health prediction.

Benefits of technology

It enables the location and interpretation of regional risks within historical buildings, reduces the inaccuracy of health prediction results, improves the interpretability of engineering projects and the ability to identify cascading degradation and spreading diseases, and reflects the true degradation mechanism and evolution path of buildings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121599239B_ABST
    Figure CN121599239B_ABST
Patent Text Reader

Abstract

The application provides a health prediction system and method for historical buildings, and relates to the technical field of building health prediction, and comprises a historical building initial region unit division module, a historical building region unit determination module, a region unit coupling matrix construction module and a historical building health prediction module. The application solves the problem that the prior art only faces the overall evaluation of the building, is difficult to depict the health difference of the internal region of the historical building, and does not explicitly model the coupling influence of the material and structure conduction between regions, resulting in insufficient prediction accuracy and interpretability at the region level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of building health prediction technology, and in particular to a health prediction system and method for historical buildings. Background Technology

[0002] Most existing building health assessment methods collect building condition data through sensors and non-destructive testing technologies, and then combine this data with various algorithmic models for evaluation. These methods rely on deploying sensors (such as strain gauges, crack monitors, and temperature and humidity sensors) at key locations within the building to periodically collect building health data. This data is then input into the health assessment model, which assesses and predicts the building's health status based on its geometric structural characteristics, condition data, and historical monitoring data.

[0003] For example, Chinese invention patent application CN115063040B discloses a method and system for collaborative assessment and prediction of building structural health. The method includes collecting data on structural components; the data includes geometric structural data, state data, and condition data; constructing a building feature map based on the geometric structural data, thereby obtaining a structural feature vector; obtaining a state data matrix based on the state and condition data; and obtaining a current health assessment value and a predicted health assessment value based on the structural feature vector and the state data matrix. The current health assessment value is used to assess the current health status of the building; the predicted health assessment value is used to predict the future health status of the building. The system includes a data intelligent sensing module and an assessment and prediction module.

[0004] Specifically, existing technologies primarily focus on the overall health of a building, with less consideration for the differences between different areas within the building. Unlike modern buildings, historical buildings often undergo multiple renovations or alterations, resulting in a high degree of diversity and complexity in their materials (such as wood, stone, and brick) and structural forms (such as vaults and wooden beams). This can lead to significant variations in the health status of different areas. For example, the basement of a historical building may experience wood decay or stone weathering due to long-term exposure to high humidity, while the roof and facade are more susceptible to the effects of temperature changes, weathering, and ultraviolet radiation, causing material degradation and structural damage in different areas. These differences are often overlooked in health assessments, making it difficult for overall health prediction models to accurately reflect the specific health status of each area.

[0005] Furthermore, the structural layout and functional zoning of historical buildings are often not as standardized and modular as modern buildings. In many historical buildings, the connections between structural components are complex and unique. Walls, columns, beams, and other components may be tightly connected using traditional methods (such as mortise and tenon joints, stone masonry joints, etc.), and the health of these components can affect each other. For example, groundwater may affect the upper floor slabs through the capillary action of the walls, and the expansion of cracks may spread to other areas through stone joints. If this complex coupling relationship between areas is not fully considered, it will also lead to a decrease in the accuracy of health assessment results. Summary of the Invention

[0006] In view of this, embodiments of the present invention provide a health prediction system and method for historical buildings, which solves the problems of existing technologies that only focus on the overall assessment of the building, are difficult to characterize the health differences in the internal areas of the historical building, and do not explicitly model the coupling effects of material and structural transmission between areas, resulting in insufficient regional prediction accuracy and interpretability.

[0007] This application provides a health prediction system for historical buildings, including:

[0008] The initial regional unit division module for historical buildings is used to obtain the original information set of the historical buildings to be predicted, construct each candidate spatial unit of the historical buildings to be predicted, introduce zoning constraints, divide the historical buildings to be predicted into regions, and obtain the initial regional unit set.

[0009] The historical building area unit determination module is used to obtain the state dataset within the coverage area of ​​the initial area unit set, construct the internal heterogeneity index of the initial area unit set and the difference distance of each adjacent initial area unit pair, and thereby perform subdivision and / or merging of the initial area unit set to obtain each area unit.

[0010] The regional unit coupling matrix construction module is used to establish a set of directed candidate edges based on each regional unit, calculate the influence weights between regional units in the set of directed candidate edges and normalize them to obtain the regional unit coupling matrix used to characterize the influence relationship between regions.

[0011] The historical building health prediction module is used to collect the time series of multidimensional health status parameters of each regional unit, construct the neighborhood aggregation representation of each regional unit based on the coupling matrix of the regional unit, use the time series of multidimensional health status parameters of each regional unit and the neighborhood aggregation representation of each regional unit as input to the health prediction model, output the health prediction results of each regional unit, and summarize the health prediction results of each regional unit to obtain the total health prediction result of the historical building to be predicted.

[0012] This application also provides a method for predicting the health of historical buildings, including:

[0013] The original information set of the historical building to be predicted is obtained, candidate spatial units of the historical building to be predicted are constructed, zoning constraints are introduced, the historical building to be predicted is divided into regions, and an initial set of regional units is obtained.

[0014] Obtain the state dataset within the coverage area of ​​the initial region unit set, construct the internal heterogeneity index of the initial region unit set and the difference distance of each adjacent initial region unit pair, and perform subdivision and / or merging of the initial region unit set to obtain each region unit.

[0015] A set of directed candidate edges is established based on each regional unit. The influence weights between regional units in the set of directed candidate edges are calculated and normalized to obtain the regional unit coupling matrix used to characterize the influence relationship between regions.

[0016] The time series of multidimensional health status parameters of each regional unit are collected. Based on the coupling matrix of the regional unit, the neighborhood aggregation representation of each regional unit is constructed. The time series of multidimensional health status parameters of each regional unit and the neighborhood aggregation representation of each regional unit are used as inputs to the health prediction model. The health prediction results of each regional unit are output. The health prediction results of each regional unit are summarized and processed to obtain the overall health prediction result of the historical building to be predicted.

[0017] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0018] 1. This invention provides a health prediction system and method for historical buildings. It adopts a series of steps, including original information collection, candidate spatial units, constrained area division, area subdivision / merging iteration, area coupling matrix, neighborhood aggregation prediction, and summary output, to transform the prediction object from a single state of the entire building into a traceable, calculable, and observable regional unit. This enables the location and interpretation of regional risks and reduces the inaccuracy of health prediction results caused by differences within the building.

[0019] 2. After the initial regional unit is formed, this invention uses internal heterogeneity index to trigger subdivision and adjacent difference distance to trigger merging. This avoids the internal state fluctuation and dispersion of the divided regional unit, the mixing of different health behaviors, or the separation of adjacent areas that are highly consistent in state. In the difference distance calculation, a decomposition and grading assignment rule of topological penalty is introduced to suppress the erroneous merging of areas that are surface adjacent but weakly connected or cross key boundaries of historical buildings. This stabilizes and preserves the structural boundaries such as floor boundaries and construction transition layers that are common in historical buildings, and improves the engineering interpretability and verifiability of regional unit boundaries.

[0020] 3. This invention, through the construction and normalization constraints of the regional unit coupling matrix, enables the mutual influence of different parts inside the historical building to enter the prediction calculation at a unified scale, thereby avoiding the subjectivity and non-reproducibility caused by relying solely on experience to judge the influence path. This makes the prediction no longer limited to the monitoring curve of a single area, reduces the regional independent misjudgment caused by insufficient local monitoring points, and improves the ability to identify chain degradation and spread-type diseases in advance.

[0021] 4. This invention integrates structural / material / environmental transmission with historical state correlation, enabling the coupling relationship to reflect the inherent construction logic of historical buildings and to be dynamically corrected by monitoring data. This avoids the problems of excessive coupling caused by relying solely on structural topology or accidental resonance being mistaken for an influence by relying solely on short-term data correlation. At the boundaries of areas with differences in repair boundaries, material replacement, or environmental exposure, the fusion mechanism can reduce the weight of unrealistic cross-boundary influences and suppress the spread of erroneous information. In continuous segments with evidence of long-term synchronous degradation, the fusion mechanism can increase the weight of effective influences, enhance the ability to capture trend risks, and make predictions more consistent with the true degradation mechanism and evolution path of historical buildings. Attached Figure Description

[0022] Figure 1 This is a framework diagram of a health prediction system for historical buildings provided in an embodiment of the present invention;

[0023] Figure 2 This is a schematic diagram of the health prediction model structure involved in an embodiment of the present invention;

[0024] Figure 3 This is an overall framework diagram of the training and deployment of the health prediction model involved in the embodiments of the present invention;

[0025] Figure 4 This is a flowchart of a health prediction method for historical buildings provided by an embodiment of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. Embodiments of the present invention provide a health prediction system and method for historical buildings, such as... Figure 1The diagram shows a framework for a health prediction system for historical buildings. The system includes: a module for initial regional unit division of historical buildings, a module for determining regional units of historical buildings, a module for constructing a regional unit coupling matrix, and a module for predicting the health of historical buildings.

[0028] Example 1: Collect original information on historical buildings and construct candidate spatial units to provide a unified spatial carrier and computable feature input for subsequent zoning and regional health prediction. Specifically, the following is done: Obtain a set of original information related to the historical building to be predicted, including at least one or more of the following: design data, historical archives, inspection records, monitoring system data, and on-site survey results. Specifically, for design data, electronic files such as historical building survey drawings, as-built / renovation design drawings, detailed structural drawings, and structural calculation instructions are received and registered; for historical archives, archives of the protection unit, previous renovation approvals, historical photos / videos, written records, and management ledgers are received and registered, and file indexes are established for scanned and image files; for inspection records, disease survey forms, non-destructive testing records, material sampling test records, structural appraisal reports, and disease image records for the corresponding historical buildings are received and registered; for monitoring system data, time-series data such as crack width, tilt, settlement, vibration, temperature, and humidity, as well as their point layout tables, are obtained from the monitoring platform or data export files; for on-site investigation results, on-site record forms, location information, photos / panoramic images, point cloud / 3D reconstruction results, etc., collected by survey personnel are received. During the acquisition process described above, a unique carrier identifier is generated for each source carrier, and the source type, acquisition method, acquisition time, identifier of the providing unit or collector (which can be anonymized), and clues related to the historical building components / areas (such as component number, room / axis / facade number, measurement point number or photo annotation) are recorded to form a set of original information that can be used for subsequent association and conflict determination.

[0029] The original information set was structured and organized, and fields related to spatial geometry, component type, material construction, repair records, damage records, and monitoring point layout were extracted as structured entries. Specifically, a field dictionary and coding rules for historical buildings were first established. The field dictionary should include at least: spatial geometry fields (grid / room / facade / elevation, component location, geometric dimensions, outline / centerline / control points); component type fields (beams, columns, walls, arches, floor slabs, roofs, foundations, and optional extended types such as brackets, purlins, corbels, platform bases, and piers commonly found in historical buildings). Material construction fields (material type such as brick / stone / wood / rammed earth / mortar, masonry / mortise and tenon / stitching, connection and construction level); Repair record fields (repair scope, process, material replacement, reinforcement method, responsible unit, repair date and acceptance conclusion); Defect record fields (defect types such as cracks, weathering, salt erosion, peeling, insect infestation, mold, leakage, deformation, etc., and their location, scale, grade, and development trend); Monitoring point layout fields (measuring point number, attached component, coordinates / elevation, sensor type, range, sampling frequency, installation date and calibration information). During field extraction, corresponding parsing and mapping processes are invoked for different source types: For drawings / surveying results, component boundaries, centerlines, elevations, and dimensions are parsed according to layer, block attributes, annotation text, and component numbering rules, and the parsing results are written into spatial geometry entries; For point cloud / 3D reconstruction results, the outer contour of historical buildings, key facades, and component control points are extracted under a unified coordinate benchmark, and they are associated with component numbers or regional partitions; For inspection reports and survey forms, component types, disease types, disease levels, quantitative indicators (length / width / area / depth, etc.), and photo indexes are extracted from table fields and chapter templates to generate disease entries; For repair archives and ledgers, repair dates, repair locations, construction techniques, and material replacement information are extracted to generate repair entries; For monitoring system data and point tables, measuring point numbers, deployment locations (associated with component / region / spatial geometry references), sensor types, and sampling parameters are extracted to generate measuring point entries and monitoring data index entries. Each entry is stored in the form of object identifier (component / measuring point / disease event / repair event) - attribute key - attribute value - unit / dimension - spatial reference - time reference, thus forming a structured set of entries that can be directly read and calculated by subsequent algorithms.

[0030] The physical spatial boundaries of the historical building to be predicted are determined based on the original information set, and the set of structural component types and topological connections of the components are determined as the basis for generating candidate spatial units. Specifically, the building's outer contour and floor boundary information are obtained from design data or survey results to form a spatial boundary description that includes the building's horizontal boundaries, elevation range, and number of floors. When design drawings are lacking, a corresponding spatial boundary description is generated based on on-site surveys or point cloud modeling results. Relevant personnel further determine the set of structural component types to be included in the modeling from archives and on-site survey results, and establish topological connections of components based on their connections, intersections, supports, or continuity to characterize the interaction paths and spatial constraints between components such as beams, columns, walls, arches, floor slabs, roofs, and foundation structures.

[0031] A candidate spatial unit set is generated using a set of structural component types and their topological connections as the framework. Component units include at least one or more of the following: beams, columns, walls, arches, floor slabs, roofs, and foundation structures. Each component serves as the minimum modeling unit, establishing a unique identifier and geometric extent. Based on the component's topological connections, the components are organized into a queryable spatial network. When the spatial scale of a component exceeds a preset scale threshold, or when there are obvious material partitions, repair boundaries, or defect boundaries within the component, the component is subdivided according to preset geometric subdivision rules, generating at least two component sub-units. These sub-units are then incorporated into the candidate spatial unit set. Geometric subdivision rules include at least one or more of the following: equidistant segmentation along the component's length direction, grid division, or segmentation based on the component's natural boundaries (such as door / window opening boundaries, arch segment boundaries, and floor slab cross-zone boundaries).

[0032] The spatial scale of the aforementioned component unit is used to measure the size of the candidate spatial unit in a geometric sense, in order to determine whether the component unit needs to be further subdivided to obtain sub-units with relatively consistent internal features. The spatial scale is not a single fixed index, but rather the corresponding scale measurement item is selected from the preset rules according to the geometric type of the component unit, and a comparable scale value is formed uniformly. Specifically, for linearly dominant component units (such as beams, purlins, rafters, connecting components, etc.), the spatial scale is defined as the length of its geometric centerline or effective span; for planarly dominant component units (such as walls, floor slabs, roof surfaces, arch panels, etc.), the spatial scale is defined as its outer contour area, unfolded area, or length in the main direction (e.g., a combination of wall segment length and height, where the wall segment length is preferred for threshold determination, supplemented by height when necessary); for volume-dominant component units (such as platform foundations, foundation blocks, piers, masonry column bases, etc.), the spatial scale is defined as the volume or the maximum side length of the outer enclosure box; when the component unit is a composite form or has significant geometric complexity such as openings, concavities, and convexities, the spatial scale can be further defined as a combination of the main dimension and complexity correction, where the complexity correction term is calculated by one or more of the following (weighted average): number of openings, proportion of opening perimeter, number of boundary curvature changes, or number of panel divisions.

[0033] The preset scale threshold is used to delineate between units that should remain as single units and those that should be subdivided. Its preset nature adheres to the feasibility and engineering rationality constraints of historical building modeling, and is determined in conjunction with data resolution, spatial heterogeneity of defects, and component load-bearing / structural segmentation characteristics. Specifically, the preset scale threshold is generated in the initialization phase by relevant experts in the field based on experience and grouped according to component categories. The determination of the threshold value must at least meet the following constraints: First, it should not be less than the minimum identifiable size corresponding to the effective resolution of available spatial data, to avoid the sub-units being unable to be stably positioned under surveying / point cloud / drawing accuracy; Second, it should not be greater than the upper limit of the spatial scale of common defects and repair boundaries in historical buildings, enabling subdivision to separate local differences such as areas with dense cracks, salt erosion zones, and leakage-affected areas from large components; Third, it should be compatible with structural logic. For example, wall segments should prioritize subdivision points such as door and window opening boundaries, corners, buttresses, and arch feet, while beams and purlins should prioritize subdivision points such as support points, mid-span, and node areas, to avoid creating sub-unit boundaries that violate structural understanding. To facilitate project implementation, the preset scale threshold can be automatically calibrated according to the statistical distribution of the same historical building: after the candidate spatial units are initially generated, the quantiles of the scale values ​​of similar components are calculated (for example, P75 or P80 is taken as the initial threshold value), and truncation is performed in combination with manually configurable upper and lower limits. This ensures that the threshold can adapt to the component scale characteristics of the building and avoid threshold distortion caused by extreme components. When higher resolution data is introduced later or the scale of the disease boundary is found to be significantly smaller than the original threshold, the threshold configuration can be updated and the subdivision can be retried without changing the overall process, so as to keep the spatial carrier granularity of the candidate spatial unit set matching the spatial scale of the actual health differences.

[0034] For each candidate spatial unit in the candidate spatial unit set, a feature vector for partitioning is constructed. This feature vector includes at least one or more of the following: spatial or geometric features, material or structural features, functional or load-bearing role features, environmental transmission features, and defects or repair features. Specifically, spatial or geometric features such as floor number, elevation, area or volume, and distance from external walls / roof / ground are extracted for each candidate spatial unit; material or structural features such as material type, construction method, connection type, age, or repair version are extracted; functional or load-bearing role features such as load-bearing / enclosure / minor components are extracted; environmental transmission features such as distance from moisture sources, sunlight / rain and wind exposure level, and groundwater impact level are extracted; and defects or repair features such as defect type code, defect density, number of repairs, and most recent repair time are extracted. These features are converted into numerical or comparable parameter representations according to preset coding rules, forming the feature vector of the candidate spatial unit, which is then used by the subsequent partitioning algorithm to perform region division processing under the premise of satisfying partitioning constraints.

[0035] The preset coding rules are used to uniformly convert the heterogeneous features of candidate spatial units of historical buildings into numerical or comparable parameter representations, so as to ensure that features from different sources, with different dimensions and types can be measured, similar to and constrained within the same computing framework during subsequent partitioning. The preset coding rules are generated and solidified by computer equipment during the offline configuration or initialization phase, and include at least a field dictionary, value range constraints, category coding table, binning rules, normalization rules, missing and conflict handling rules, and time feature conversion rules. The pre-setting process of the pre-defined coding rules includes at least the following sub-steps: First, construct a feature field dictionary, defining spatial or geometric features, material or structural features, functional or load-bearing role features, environmental transmission features, and defects or repair features as field groups, and assigning a field identifier, data type label (continuous / discrete / ordinal / set / time), and allowed unit and dimension label to each field; among them, constraints are imposed on common objects of historical buildings, including at least component type extension items (such as brackets, purlins, corbels, platform bases, etc.), material category extension items (such as blue bricks, rammed earth, stone, wood, mortar, etc.), and defect type extension items (such as weathering, salt erosion, insect infestation, mold, leakage, freeze-thaw peeling, etc.) to ensure that the coding rules are consistent with the object genealogy of historical buildings. For continuous spatial geometric features (such as floor number, elevation, area or volume, distance from exterior walls / roof / ground), a unit unification-normalization coding rule is preset: First, based on the field dictionary, the length is unified to meters, the area to square meters, the volume to cubic meters, and the elevation to the elevation value under the same reference plane; then, based on the statistical distribution of historical building sample database or current engineering data, a reasonable value range and truncation threshold for each field are preset (e.g., using quantile upper and lower bounds or engineering experience boundaries), and out-of-bounds values ​​are truncated to the threshold range; subsequently, it is converted into a comparable parameter representation according to a preset normalization method. The normalization method includes at least one or more of min-max normalization, mean-variance standardization, or logarithmic compression. Among them, long-tailed distributed fields such as area and volume are preferentially compressed on a logarithmic scale before normalization to reduce the dominant effect of maximal units on distance measurement. For categorical material or structural features (e.g., material category, construction method, connection type, age, or repair version), a pre-defined enumeration dictionary-multi-value encoding rule is used: enumeration sets are established for material category, construction method, and connection type, and a stable code is assigned to each enumeration item; when candidate spatial units have multiple material levels or composite structures, they are encoded according to multi-value sets, and the multi-value set encoding includes at least multi-hot encoding or hierarchical expansion encoding; for age or repair version ordinal features, they are converted into comparable parameter representations on the time axis, including at least one of the year difference from the base year, the time difference from the most recent repair, or the version number, and can be further converted into weight values ​​according to the time decay function, so that newer repair versions contribute more to the partition similarity.For functional or load-bearing role characteristics (e.g., load-bearing / enclosure / secondary components), pre-defined ordinal level coding rules are used: load-bearing and functional roles are mapped to ordered levels or weight coefficients. For example, load-bearing components are assigned a higher structural importance coefficient, enclosure components a medium coefficient, and secondary components a lower coefficient. When there are composite roles in the same unit, a single role coefficient is generated according to pre-defined priority or linear combination rules, and the combination source is recorded for traceability. For environmental transmission characteristics (e.g., distance to moisture source, solar / rain exposure level, groundwater impact level), pre-defined binning and level mapping rules are used: for measurable distance features, continuous normalization or binning according to engineering thresholds is used (e.g., mapping distance ranges to levels from 0 to K). For level fields such as exposure level and groundwater impact level, pre-defined level enumeration and ordinal coding are used, and correction factors are allowed to be introduced based on historical building characteristics such as building orientation, shading conditions, and roof drainage organization, so that the same exposure level can produce different numerical results under different orientations or different shading conditions. For defects or repair characteristics (such as defect type coding, defect density, number of repairs, and most recent repair time), a preset event aggregation-time decay coding rule is established: defect types are coded according to a preset defect dictionary; defect density is normalized and aggregated according to the area or number of components of the candidate spatial unit (e.g., number of defects per unit area, number of cracks per unit length); the number of repairs is directly counted or weighted by repair level; the most recent repair time is converted into the time difference from the current analysis period, and further converted into an influence coefficient through a preset decay function to reflect the historical building pattern of repair effect decay over time. In summary, the preset coding rule is not arbitrarily set, but is formed and solidified by computer equipment in the initialization stage based on constraints such as the genealogy of historical building objects, field data types, consistency of unit dimensions, reasonable value range of engineering projects, event time effects, and traceability of missing and conflicting data. This stably converts various characteristics of candidate spatial units into computable and comparable feature vector representations.

[0036] Example A: Take a typical two-story historical building A, which is a brick-wood mixed structure, as the object. The building is a two-story brick-wood mixed structure. The first floor is a hall and side rooms, and the second floor is a corridor and a wooden roof structure. The building’s outer outline is an irregular polygon. The first floor’s reference elevation is 0.00m, and the highest elevation of the roof ridge is about 9.80m. The collected original information specifically includes: ① 1998 survey results (one set each of scanned paper map and vector file generated by later digitization, including site plan, plan, elevation, section and grid annotations); ② 2013 repair design documents (including detailed drawings of structural nodes, material descriptions and reinforcement instructions); ③ 2022 structural appraisal report (including description of the current status of components and conclusions on load-bearing capacity); ④ 2024 disease survey form (table file, with disease photo numbers and orientation annotations); ⑤ Time series data exported from the monitoring platform (crack width, tilt, settlement, temperature and humidity, time span from January 2025 to October 2025, sampling interval 10 minutes) and point layout table; ⑥ 2025 site survey record form (including photos and text records with component number annotations) and local point cloud model (used to complete local geometry). Each of the above-mentioned data was assigned a carrier identifier and its source type and formation time were recorded. The carrier identifiers for surveying results are “MAP-1998-01”, repair design carrier identifiers are “DES-2013-02”, structural appraisal carrier identifiers are “REP-2022-01”, disease survey carrier identifiers are “SUR-2024-03”, monitoring time series carrier identifiers are “MON-2025-TS”, point layout table carrier identifiers are “MON-2025-PT”, and reconnaissance record carrier identifiers are “VIS-2025-08”. Generate building outline boundary entries (horizontal boundary polygon coordinate sequence + floor boundary elevation 0.00m / 3.60m / 7.20m + roof area), and parse the grid (A~F axes, 1~6 axes) and room numbers (101~116, 201~208); after parsing the component geometry and numbering rules, form a set of component geometry entries: columns "C-1F-001~C-1F-032, C-2F-001~C-2F-024", beams "B-1F-001~B-1F-048, B-2F-001~B-2F-036", and walls "W-EXT-0 01~W-EXT-014 (exterior wall section), W-INT-001~W-INT-022 (interior wall section)", arch “AR-001~AR-008”, floor slab “SL-1F-001~SL-1F-006, SL-2F-001~SL-2F-005”, roof timber structure “RF-TR-001~RF-TR-012 (roof truss / truss), RF-PU-001~RF-PU-026 (purlin), RF-RA-001~RF-RA-120 (rafter)”, platform / foundation “FD-001~FD-006”.Defect entries were extracted from the defect survey form and structural assessment report: For the exterior wall section W-EXT-007, a defect event entry “DEF-2024-112 (object identifier W-EXT-007, defect type = salt erosion, grade = II, area = 1.8㎡, location = 1st floor south facade A-3 axis, photo index = P-2024-331~P-2024-336)” was generated; for column C-1F-014, a defect event entry “DEF-2024-087 (defect type = insect infestation, grade = III, affected range = column base 0.0~0.6m, photo index = P-2024-210~P-2024-214)” was generated; for arch AR-004, a crack entry “DEF-2024-156 (defect type = crack - radial, maximum crack width = 0.62mm, length = 1.9m)” was generated. Extract repair entries from the repair archive: Generate repair event “FIX-2013-05 (object identifier W-EXT-007, process method = grouting + partial brick replacement, material replacement = blue brick, repair date = 2013-06-30, acceptance conclusion = qualified)”; Generate reinforcement event “FIX-2013-11 (object identifier B-1F-021, reinforcement method = steel hoop + grouting reinforcement, repair date = 2013-07-12)”. Extract monitoring point entries from the point layout table and associate them with the components: Crack measuring point “P-CR-013 (attached component = W-EXT-007, elevation = 1.25m, coordinates = (x,y,z), range = ±5mm, sampling frequency = 10min)”; Inclination measuring point “P-TI-002 (attached component = C-1F-014, elevation = 2.80m)”; Settlement point “P-SE-004 (attached object = FD-003)”; Temperature and humidity point “P-TH-006 (attached area = northwest corner of the first floor)”. Each structured entry simultaneously includes the source type, formation time, and credibility marker: "DEF-2024-112" has a source type of inspection record (disease survey) and a formation time of 2024-09-05; the credibility marker "source reliability" is assigned a value according to "identification report > measurement record > repair archive > surveying results > on-site survey record > historical text record". In this example, the disease survey entry is assigned 0.80, the identification report entry is assigned 0.90, the monitoring time series entry is assigned 0.95, and the historical text record entry is assigned 0.50; if the same object has multiple values ​​for the same attribute (e.g., the material is recorded as "adobe" in the historical record, but the sampling record is "rammed earth with blue brick skin"), then multiple material entries are retained in parallel and the same conflict group identifier "CON-MAT-W-INT-006" is set, without overwriting the original value.Based on the above spatial boundary description and component set, the component topological connection relationships are established as follows: C-1F-014 has a "support" relationship with B-1F-021 and B-1F-022; W-EXT-007 has a "support" relationship with FD-003 and an "intersection / opening boundary" relationship with AR-004; RF-TR-006 has a "connection" relationship with RF-PU-014 to RF-PU-018, and RF-PU-016 has a "support / continuity" relationship with RF-RA-041 to RF-RA-052. Subsequently, a candidate spatial unit set was generated using components as the minimum modeling carrier: the initial number of candidate units was (56 columns + 84 beams + 36 walls + 8 arches + 11 floor slabs + 158 roof timber structures + 6 foundations), totaling 359. Components that met the subdivision conditions were subdivided: exterior wall segments longer than 6m were divided into 3m segments at equal intervals along the length direction; W-EXT-007 (total length 9.2m) was subdivided into three segments: W-EXT-007-A / B / C. Wall segments containing door and window openings were divided into three sub-units based on the opening boundaries: upper lintel, side wall limbs, and lower wall. Arch AR-004 was divided into AR-004-L / AR-004-M / AR-004-T based on the natural segmentation boundaries of arch foot—arch shoulder—arch crown. Beam B-1F-021, which showed obvious defect boundaries, was further divided into B-1F-021-1 / 2 based on the boundaries between dense and non-dense crack areas. After subdivision, the total number of candidate spatial units is expanded to 612, while the parent-sub-unit relationship is retained for traceability. A feature vector is formed for each candidate spatial unit and numericalized: taking sub-unit W-EXT-007-B as an example, its spatial / geometric features are: layer number = 1; elevation range = 0.00~3.60m; area = 3.0m×3.6m=10.8m. 2Distance from exterior wall = 0m; Distance from roof ≈ 6.2m; Distance from ground moisture source (based on the center line of the damp zone identified by on-site survey) = 1.2m. Material / structural characteristics: Material type = blue brick + mortar; Construction method = masonry; Connection type = intersecting with column + supporting with foundation; Year / Repair version = 2013 repair version exists. Functional / load-bearing role characteristics: Load-bearing + enclosure composite (load-bearing is the primary role based on priority). Environmental transmission characteristics: Sunlight level = low (north-facing obstruction); Wind and rain exposure level = high (windward facade); Groundwater impact level = medium (near the outer edge of the platform). The characteristics of the disease / repair are as follows: the disease type code includes "salt erosion", "surface weathering" and "cracks - vertical"; the disease density = number of cracks / wall segment length = (3 cracks / 3.0m) = 1.0 cracks / m; the number of repairs = 1; the time from the most recent repair to the analysis period (2025-10) = 12.3 years. When quantifying, the following fixed codes are used: In the material category coding table, "Blue Brick = 03, Mortar = 11, Timber = 01, Stone = 02, Rammed Earth = 05"; in the construction method coding table, "Masonry = 21, Mortise and Tenon = 31, Corrugated = 22"; in the load-bearing role coding table, "Bearing = 3, Enclosure = 2, Secondary = 1"; in the exposure level coding table, "Low = 1, Medium = 2, High = 3"; in the disease type coding table, "Salt Erosion = 104, Weathering = 101, Insect Infestation = 203, Leakage = 301, Cracks - Vertical = 401"; continuous quantities are normalized to 0-1 after unit unification (based on the minimum-maximum interval obtained from the statistics of all candidate units of this building, for example, the area normalization interval [0.8m]). 2 18m 2 [、Damp source distance range [0m, 8m]). Accordingly, the feature vector of W-EXT-007-B in this example falls as: Geometric vector: [Layer number = 1, elevation value = 1.8m, area_norm = 0.58, distance from outer wall_norm = 0.00, distance from roof_norm = 0.63, distance from damp source_norm = 0.15]; Material / construction multi-value encoding: [Material (03) = 1, Material (11) = 1, Construction (21) = 1, Connection "Intersection / Support" flag = 1, Repair version exists] [Flag bit = 1]; Role / environment ordinal code: [Force-bearing role = 3, Sunlight = 1, Wind and rain exposure = 3, Groundwater influence = 2]; Disease / repair code: [Disease (104) = 1, Disease (101) = 1, Disease (401) = 1, Crack density_norm = 0.72, Repair times = 1, Time difference between most recent repairs (years) = 12.3]; and store this vector together with the object identifier W-EXT-007-B as direct input for subsequent partitioning.

[0037] Partitioning constraints are set based on the candidate spatial unit set. These constraints include at least: minimum and / or maximum spatial scale constraints for regional units, minimum discriminative constraints between regional units, topological connectivity constraints for regional units, and observability constraints. Specifically, the minimum spatial scale constraint ensures that the composite geometric scale of the candidate spatial units contained in any regional unit is not less than a preset minimum scale threshold (based on empirical presets, the preset rules for the following thresholds are the same and will not be repeated here), to avoid regional units being too fragmented, leading to unstable regional-level predictions. The maximum spatial scale constraint ensures that the composite geometric scale of any regional unit is not greater than a preset maximum scale threshold, to avoid regional units being too large and masking local differences. The minimum discriminative constraint ensures that the difference between any two adjacent or candidate merged regional units in the feature vector space is not less than a preset discriminative threshold, to avoid mistakenly merging units with significantly different properties into the same region. The topological connectivity constraint ensures that the candidate spatial units within each initial regional unit form a connected subgraph under the topological connection relationship or spatial adjacency relationship of historical building components. The observability constraint ensures that the observability index of each initial regional unit is not less than a preset minimum observability threshold, to ensure that regional-level health predictions have sufficient observational support.

[0038] A coverage mapping from monitoring points to a set of candidate spatial units is established, and the observability index of each candidate spatial unit is calculated based on the coverage mapping. Specifically, for each monitoring point, based on its deployment location, attached component identification, spatial coordinates, and influence radius / effective range, the set of candidate spatial units covered by it is determined, forming a monitoring point-candidate spatial unit coverage mapping relationship. When a monitoring point is directly attached to a candidate spatial unit, the coverage relationship includes at least that candidate spatial unit; when the effective range of a monitoring point spans multiple candidate spatial units, the coverage relationship includes all candidate spatial units falling within its effective range, and the coverage weight is recorded. Subsequently, the observability index is calculated for each candidate spatial unit. The observability index reflects at least one or more of the following factors: the number of monitoring points covering the unit, the richness of monitoring point types, sampling frequency, data availability / missing rate, and the cumulative value of the coverage weight. Furthermore, a minimum observability threshold is set, and it is stipulated that the observability index of the initial regional unit is obtained by the observability index of the candidate spatial units it contains according to a preset aggregation rule. The aggregation rule includes at least one or more of the following: summation, mean, weighted mean, minimum value constraint, or quantile aggregation. When dividing the region, the observability constraint is limited to: the observability index of any initial regional unit is not less than the minimum observability threshold.

[0039] Under the premise of satisfying the partitioning constraints, a hard-constrained region growth partitioning algorithm based on adjacency graphs is invoked. Using the feature vectors of candidate spatial units as the similarity basis, spatial scale, discriminability, topological connectivity, and observability are constrained and verified during each absorption / merging process to obtain an initial set of regional units. Each initial regional unit consists of one or more candidate spatial units. Specifically, the region partitioning process uses candidate spatial units as the basic objects and their feature vectors as input. Adjacency relationships are formed by combining topological connectivity constraints, and the determination of candidate allocation, merging, or splitting is iteratively executed: if adding a candidate spatial unit to a regional unit causes that regional unit to violate the minimum / maximum spatial scale constraint, topological connectivity constraint, or observability constraint, the addition operation is revoked or the regional unit is re-splitted; if merging two adjacent regional units results in a discriminability lower than the minimum discriminability threshold or an excessive maximum spatial scale, merging is prohibited; if the observability index of a regional unit is lower than the minimum observability threshold, candidate spatial units connected to it and contributing the most to improving observability are preferentially absorbed until the threshold is met or the maximum spatial scale constraint is triggered, at which point a re-partitioning strategy is adopted. Thus, under the premise that the constraints are continuously satisfied, the region division is completed, and the initial set of region units is obtained.

[0040] Output an initial set of regional units and establish a mapping relationship between the initial set of regional units and the set of candidate spatial units. Specifically, generate a regional unit identifier for each initial regional unit and record its list of candidate spatial unit identifiers, synthetic spatial boundary / geometric extent, aggregated regional feature vector, regional observability index, and constraint satisfaction verification results; at the same time, generate a mapping table of candidate spatial unit identifiers and initial regional unit identifiers, which is used for regional aggregation and tracing of monitoring data, disease events, and repair events in the subsequent regional health prediction stage.

[0041] Before the region partitioning begins, only the basic unit-level quantities of candidate spatial units are calculated and stored, including the spatial scale S(u), eigenvector x(u), and observability index O(u) of the candidate spatial unit. Region-level indices are generated in real-time during the region partitioning process according to aggregation rules: when a candidate spatial unit u is selected as a seed, it is treated as an initial candidate region unit R={u}, and the spatial scale of region R is set... Aggregated feature vector of region R Observability index of region R Agg = Aggregation. S Rules / functions for aggregating scales (e.g., taking the maximum side length of the bounding box, or calculating after taking the union), Agg x Rules / functions for aggregating feature vectors (e.g., weighted mean, most-hot count, mode, etc.).O Rules / functions for aggregating observability (e.g., weighted average + weakest link protection).

[0042] When a region grow, merge, or migrate operation updates the member set of a region unit to... (u' represents a newly added region unit) or When (i,j are the numbers of the region units, i / j=1,2,...,n,n is the number of region units), S(R'), x(R') and O(R') are incrementally updated according to the preset aggregation rules, and the minimum / maximum spatial scale constraints, minimum discriminative constraints, topological connectivity constraints and observability constraints are checked accordingly; if the check fails, the update is rolled back or other candidate operations are used.

[0043] For the initial set of regional units, based on the attribution mapping relationship between the initial set of regional units and the candidate set of spatial units, a state dataset covering the area of ​​the initial set of regional units is obtained. Specifically, let the candidate set of spatial units be U, and the initial set of regional units be {R}. i}, the attribution mapping is For each candidate spatial cell u, the coverage mapping between monitoring point p and candidate spatial cell u is performed. And the time series data of the monitoring points, first form a unit-level state parameter sequence s u (t), where the state parameters include at least one or more of the following: crack width, tilt, settlement, vibration response, temperature, and humidity; when multiple monitoring points cover the same candidate spatial unit, they are ranked according to coverage weight. Weighted fusion of similar parameters, for example, setting parameter m as... , where y u,m (t) represents the observed value at monitoring point p at time t, where P m To provide a set of monitoring points that can provide parameter m, the unit-level state sequences are then aggregated into a region-level state dataset according to the attribution mapping: for each initial region unit R, its state dataset S R (t) is satisfied by all S of candidate spatial units u The data is composed of (t) and aggregated according to area / volume weight w(u) to obtain a regional-level data sequence that can be used for statistical calculation. In this embodiment, the preferred approach is to... .

[0044] After aggregation, the status dataset undergoes time alignment and missing tagging to eliminate incomparability caused by asynchronous, discontinuous, and drifting historical building monitoring data. Specifically, a unified time grid is selected ( The time step Δt is determined by the main sampling interval of the monitoring system or is an approximation of the least common multiple of multiple sampling intervals, and upper and lower limits (e.g., 1 min to 60 min) are allowed to be set (empirical setting); for any parameter sequence S R,m (t), employing an alignment strategy combining nearest-neighbor sampling and window aggregation: when in (where 'a' is the label of each sampling time point in the time grid, a = 1, 2, ..., N, where N is the number of sampling time points) If there are multiple observations within a window, the median or weighted mean is used as the alignment value; if there are no observations within a window, they are marked as missing and a missing marker is written. If there is observation Simultaneously, when an observation exceeds the sensor's range, or when its amplitude changes within a short period exceeds a preset threshold (determined by historical building monitoring experience or statistical quantiles), the point is marked as an anomaly and treated as a missing value for subsequent robust statistical processing. This results in a regional-level state dataset that performs time alignment and missing value labeling. .

[0045] For each initial region unit, robust center values ​​for each state parameter within the initial region unit are calculated based on its state dataset, and normalization is performed on state parameters of different dimensions. Specifically, for each parameter m, after removing missing and outliers, a robust center value C is calculated. R,m The median is preferred as the robust central value: the median is defined as follows: To achieve comparability across parameters, robust normalization is performed for each parameter m, preferably using standardization based on MAD (Median Absolute Deviation): Let Then the normalized value ,in This is a small constant set based on experience to prevent division by zero.

[0046] The internal heterogeneity index of the initial regional unit is calculated based on the normalized state dataset to characterize the degree of fluctuation and dispersion of the state parameters within the initial regional unit. Specifically, the internal heterogeneity index is preferably obtained by fusing multiple parameter dispersions: the standard deviation of the normalized sequence for each parameter is calculated and then weighted and fused according to preset weights, where the weights can be configured and normalized according to sensor reliability, contribution to health prediction, or inverse proportion to the missing rate.

[0047] Adjacent initial region unit pairs are identified, defined as initial region units directly connected in terms of topological connectivity. Specifically, an adjacency table is first established between candidate spatial units, obtained from geometric adjacency and / or component topological connections (such as support, intersection, continuity). Then, based on the attribution mapping relationship, the adjacency relationships of candidate spatial units are elevated to region-level adjacency relationships: when there is a pair of mutually adjacent candidate spatial units, and they belong to two different initial region units respectively, these two initial region units are determined to be adjacent at the region level, and they are added to the set of adjacent initial region unit pairs. For any adjacent initial region unit pair, a center value difference term for each state parameter is constructed based on the difference in their respective robust center values. The center value difference terms are then fused to obtain the difference distance. A topological penalty term is introduced to correct the difference distance, resulting in the difference distance for each adjacent initial region unit pair. Specifically, for a pair of adjacent areas, robust central values ​​of each state parameter are read (including, but not limited to, the median crack width, median settlement, median tilt, median temperature and humidity, etc.). A central value difference term is calculated for each state parameter, representing the difference in typical levels between the two areas for that parameter. To allow comparison of parameters with different dimensions (millimeters, millimeters / meter, degrees Celsius, etc.), the difference is divided by the robust fluctuation scale of that parameter across the entire historical building (e.g., the overall median absolute deviation of that parameter) when calculating the central value difference term, thus converting the difference into the degree of deviation relative to the parameter's natural fluctuation range. Subsequently, the central value difference terms of each state parameter are fused according to preset fusion weights to obtain the basic difference distance. The fusion weights can be set and normalized based on sensor reliability, data missing rate, or contribution to health prediction, so that parameters with more reliable data and higher contributions have a larger proportion in the difference distance.

[0048] When introducing a topology penalty term to correct for basic difference distances, the topology penalty coefficient is determined by decomposing and assigning values ​​in tiers based on the penalty coefficient. The topology penalty coefficient is a value greater than or equal to 1 and not greater than a preset upper limit. It is obtained by multiplying the cross-regional adjacency strength penalty sub-coefficient, the cross-floor penalty sub-coefficient, the cross-critical structural boundary penalty sub-coefficient, and the critical component type difference penalty sub-coefficient in sequence. To prevent the penalty from amplifying infinitely, an upper limit for the penalty coefficient is preset by experts in the field based on experience, and the product result is truncated to this upper limit. The cross-regional adjacency strength penalty sub-coefficient is determined using a boundary strength tiering mapping table preset by experts in the field based on experience. Boundary strength is defined as the number of cross-regional candidate spatial unit adjacency pairs (or the number of equivalent shared boundary segments) between two regions. For example, when the boundary strength is 1-2, it is determined to be a weak boundary and a penalty sub-coefficient of 1.40 is assigned; when the boundary strength is 3-5, the penalty sub-coefficient is 1.20; when the boundary strength is 6-10, the penalty sub-coefficient is 1.10; and when the boundary strength is greater than 10, the penalty sub-coefficient is 1.00. This grading rule significantly increases the difference distance for adjacent areas supported only by a few adjacency relationships during the merging determination, while adjacent areas with wide boundaries and strong connections are not penalized additionally. The cross-floor penalty sub-coefficient is determined by combining floor relationship discrimination with a fixed multiplier. If all candidate spatial units of two areas are located on the same floor, the cross-floor penalty sub-coefficient is 1.00; if the two areas are located on different floors or the candidate spatial unit sets of the two areas cross floor boundaries (i.e., merging will form a cross-floor area), the cross-floor penalty sub-coefficient is 1.25; if the cross-floor relationship also occurs at typical boundaries of historical buildings (such as platform-first floor, first floor-second floor, eaves-roof timber structure transition layer), the cross-floor penalty sub-coefficient is 1.35 to reflect the higher merging risk at structural level transition points. The cross-critical structural boundary penalty sub-coefficient is determined by combining boundary type enumeration with a multiplier table, and the common structural logic of historical buildings is used as the fixed boundary type set. If the boundary of adjacent areas does not cross the key structural boundary, the penalty coefficient is 1.00; if it crosses the boundary between the exterior and interior walls, it is 1.20; if it crosses the boundary between the arch / wall segment or the control zone of the opening boundary, it is 1.25; if it crosses the boundary between the roof timber structure / lower enclosure (such as the boundary between the truss / purlin system and the wall / column grid system), it is 1.35; if it crosses the boundary between different material systems (such as timber and masonry, masonry and rammed earth, stone and brick) and the boundary is supported by repair records or structural node evidence in the original information, it is 1.40. This coefficient table makes it more difficult to mistakenly merge adjacent areas that cross the key structural boundaries of historical buildings, even if the differences in their conditions are small. The penalty coefficient for key component type differences is determined by combining the dominant component type determination with the difference level mapping.The dominant component type is determined by the highest proportion of candidate spatial units within the region, weighted by area / volume. When two regions have the same dominant component type (e.g., both are predominantly masonry wall segments or both are predominantly timber roof components), the penalty coefficient is 1.00. When two regions have different dominant component types but belong to the same major category (e.g., both belong to masonry but are wall segments and arches respectively, or both belong to timber but are beams and purlins respectively), the penalty coefficient is 1.10. When two regions have dominant component types belonging to different major categories (e.g., timber and masonry, masonry and foundation / pile categories), the penalty coefficient is 1.25. When different major categories are accompanied by different load-bearing roles (e.g., load-bearing timber and enclosing masonry), the penalty coefficient is 1.35. This mapping makes it less likely for regions with significant differences in component lineage and load-bearing roles to be merged. The final topology penalty coefficient is obtained by multiplying the above four types of penalty sub-coefficients, and the product result is compared with the upper limit of the penalty coefficient and the smaller value is taken as the final penalty coefficient; the final difference distance is obtained by "basic difference distance × final topology penalty coefficient", and the final difference distance and the merging threshold are used together for the determination of merging adjacent regions.

[0049] When the internal heterogeneity index of any initial regional unit exceeds a preset subdivision threshold, the initial regional unit is subdivided according to preset subdivision rules, generating at least two sub-regional units. The subdivision threshold is set using a data-driven approach: after obtaining the initial set of regional units, the distribution of internal heterogeneity indices across all regions is statistically analyzed, and a higher quantile interval (e.g., 70%–85% quantile) is selected as a candidate threshold. Upper and lower limits are allowed to accommodate different historical building types and monitoring densities. When the internal heterogeneity of a region exceeds this threshold, it indicates inconsistent internal state performance within the region, posing a risk of mixing different health behaviors within the same region, necessitating subdivision. The subdivision rules are executable rules, including at least: extracting the induced adjacency subgraph of the candidate spatial units within the region to ensure connectivity in the subdivision results; then, segmenting based on the differences in the state performance of candidate spatial units within the region, with segmentation criteria including at least one or more of the following: making the state within sub-regions more consistent, prioritizing segmentation along natural structural boundaries, and segmenting at weaker internal boundaries. In practice, candidate spatial units within the region can first be binary clustered according to their unit-level robust central values ​​(or normalized state statistics) to obtain two sets of candidate subsets. Then, a connectivity check is performed on each set of candidate subsets. If a set has multiple connected components, it is split into multiple sub-regions according to the connected components. Subsequently, each sub-region is checked to see if it meets the partitioning constraints (such as minimum / maximum spatial scale, observability, topological connectivity, and minimum distinguishability from adjacent regions). If it does not meet the constraints, the current split is rolled back or a suboptimal split point is used (such as prioritizing the split along natural boundaries such as door and window openings, arch foot / arch shoulder segment boundaries, and floor slab cross-region boundaries) until at least two sub-region units that meet the constraints are obtained.

[0050] When the difference distance between any two adjacent initial region units is less than a preset merging threshold, the adjacent initial region unit pairs are merged to obtain merged region units. The merging threshold is preset using a data-driven approach: in one iteration, the difference distance between all adjacent region pairs is statistically analyzed, and a lower quantile interval (e.g., 10%–30% quantile) is selected as a candidate merging threshold, allowing for adjustments based on monitoring sparsity (the sparser the monitoring, the more relaxed the threshold; the denser the monitoring, the tighter the threshold). During merging, the adjacent region pairs are first considered as candidate merged regions. After temporary merging, it is immediately checked whether hard constraints are violated: for example, whether the spatial scale after merging exceeds the maximum spatial scale constraint, whether the observability after merging is still not lower than the minimum observability threshold, whether merging with surrounding adjacent regions will lead to insufficient distinguishability, and whether topological connectivity is still maintained after merging. If any hard constraint is violated, merging is prohibited; if all are satisfied, merging is confirmed and the original region pair is replaced.

[0051] After performing subdivision and / or merging processes, the mapping relationship between the regional units and their corresponding candidate spatial unit sets is updated, and each regional unit is output when a preset iteration stopping condition is met. Specifically, during each subdivision, the original regional identifier is replaced with two or more sub-region identifiers, and the candidate spatial units within that region are reassigned to the corresponding sub-regions according to the subdivision results; during each merging, adjacent regional pairs are replaced with the merged regional identifier, and the candidate spatial units within both regions are uniformly assigned to the same merged region; simultaneously, the regional-level adjacency relationship and cross-regional adjacency strength are updated to ensure that adjacent regional pairs can still be correctly identified and the difference distance calculated in the next round. The iteration stopping condition includes at least: no subdivision and no merging are triggered in an iteration, indicating that the regional structure has stabilized; or the number of iterations reaches a preset upper limit (e.g., 3 to 20 times configurable) to prevent abnormal data from causing infinite loops. To prevent subdivision and merging from oscillating repeatedly in adjacent iterations, a stability stopping condition can also be set: when the number of regions remains unchanged for several consecutive iterations, or the proportion of candidate spatial unit affiliation changes is lower than a preset proportion threshold, the iteration is terminated early and the regional unit set is output.

[0052] The process of constructing a regional unit coupling matrix is ​​used to explicitly characterize the directional relationships of state propagation / influence between regions at the regional unit level, and to provide computable coupling inputs for subsequent regional-level health prediction. Specifically, a regional unit adjacency graph is used as the basis for candidate edge screening. The adjacency graph is obtained by improving the adjacency relationships of candidate spatial units through a membership mapping, and spatial adjacency discrimination is supplemented to cover the case of geometrically close neighbors but not directly sharing boundaries. Spatial adjacency discrimination is jointly determined by a minimum geometric distance threshold and a floor consistency threshold. That is, when the minimum geometric distance between two regional units is not greater than the preset adjacency distance threshold and meets the condition of being on the same floor or adjacent floors, they are judged as a candidate association pair. For each screened regional unit pair, two or one directed edge is generated using a direction discrimination rule: only a unidirectional edge is generated when there is clear evidence of propagation direction, and a bidirectional edge is generated when there is a lack of directional evidence. The direction determination rules include at least the following: for moisture damage / leakage, the transmission direction is from high to low elevation, and from upstream to downstream along the drainage path; for settlement / foundation impact, the direction is from the foundation unit to the upper load-bearing unit; for roof rainwater infiltration impact, the direction is from the roof area to the lower enclosure and beam / column area; for wind and rain exposure impact, the direction is from the windward side of the exterior wall to the inner enclosure area. When none of the above rules are satisfied, bidirectional candidate edges are used, and the dominant direction is distinguished by historical correlation components in the subsequent weight calculation. This results in a set of directed candidate edges, and for each directed edge, the edge endpoint area identifier, edge type label (topology / spatial adjacency / environmental transmission, etc.) and direction source reason code are recorded.

[0053] For each directed edge in the set of directed candidate edges, the original influence weight between regional units is calculated based on the structural or functional association information, material continuity information, environmental transmission information, and / or historical state data association information between the two regional units at the two ends of the directed edge. Specifically, two types of components are set for each directed edge: static association component and historical association component, and the original influence weight is obtained by using a preset fusion rule. The static association component includes at least one or more of the following: spatial adjacency component, material continuity component, functional or structural dependence component, and environmental transmission component. The historical association component is calculated from the historical state data sequence of the two regional units at the two ends. The following defines the acquisition method, value range, and fusion method of each component: the spatial adjacency component is determined by combining the shared boundary strength grading with distance attenuation. Specifically, the boundary contact strength between the two regions is first calculated: when the two regions share a boundary, the number of cross-regional candidate spatial unit adjacency pairs is used as the contact strength; when the two regions only meet the distance adjacency requirement and do not share a boundary, the contact strength is set to the lowest level and recorded as distance adjacency type. The spatial adjacency component values ​​are then obtained by mapping according to the contact strength: 0.20 for contact strength of 1-2, 0.35 for 3-5, 0.55 for 6-10, and 0.70 for greater than 10. If it is a distance adjacency type, the value is adjusted according to the distance level based on the minimum distance of 0.20: 0.20 for minimum distance less than or equal to 1m, 0.15 for (1m, 3m], 0.10 for (3m, 5m], and edges exceeding 5m are not considered candidate edges. The spatial adjacency component values ​​are limited to between 0 and 0.70.

[0054] The material continuity component is determined by a joint judgment of material consistency and structural continuity. Specifically, the dominant material type set and dominant structural method set of the two regions are read, and these are used as the joint query key to extract the material continuity component from the pre-stored material continuity component mapping table in the database. For example, when the dominant materials and structural methods of the two regions are completely consistent (such as both being brick masonry with consistent masonry methods, or both being timber mortise and tenon systems with consistent connection methods), the material continuity component is 0.60; when the materials are consistent but the structural methods are different (such as both being brick masonry but with different pointing / grouting / sandwich wall methods), it is 0.40; when the materials are different but there is a provable continuity interface (such as timber purlins and timber rafters, or upstream and downstream of the same timber system), it is 0.30; when both the materials and structures are significantly different and there is no evidence of a continuity interface, it is 0.10. The value of the continuous component of the material is limited to between 0.10 and 0.60, and is adjusted downward according to the source credibility mark: when the material source is only from historical text records or low credibility entries, the component is multiplied by 0.85; when the material source is from sampling and testing or repair and acceptance records, it remains unchanged.

[0055] Functional or structural dependency components are determined using a strength grading system based on force path / support relationship. Specifically, region units are projected onto load-bearing force transmission / constraint paths based on component topological connections. Using this as the query key, the corresponding functional or structural dependency components are extracted from the database's predicted and stored functional or structural dependency component mapping table. For example, when the source region unit contains dominant components such as foundations or pedestals and the target region unit contains upper load-bearing components with evidence of support relationships, the functional or structural dependency component is set to 0.70; when the source region contains dominant components such as beams / arches / trusses and the target region contains columns / walls or secondary components supported by them with evidence of direct connection, it is set to 0.55; when the two regions are only in the same force system but lack evidence of direct support relationships, it is set to 0.35; when the two regions are mainly enclosures or decorations and lack evidence of structural dependency, it is set to 0.10. The component values ​​are limited to between 0.10 and 0.70, and directionality is constrained: only when there is a clear support / force transmission direction are values ​​assigned along that direction, and 0.10 is used on the opposite side.

[0056] The environmental conduction component is determined by combining conduction type enumeration with a fixed multiplier table, and a fixed direction rule is applied to its directionality. Specifically, the candidate edge direction is used as the query key value to extract the corresponding environmental conduction component from the environmental conduction component mapping table. For example, when the candidate edge direction conforms to the direction of moisture damage / leakage conduction (elevation from high to low, or roof to lower enclosure) and the wind and rain exposure level of the source area is higher than that of the target area, the environmental conduction component is set to 0.60; when the candidate edge direction conforms to the direction of groundwater / humidity source influence (near the outer edge of the foundation or a more intimate area of ​​the moisture source pointing to a more inward or higher area), it is set to 0.45; when the candidate edge direction conforms to temperature and humidity diffusion or thermal bridging conduction (windward side / high exposure side of the exterior wall pointing inward or adjacent area), it is set to 0.35; when there is a lack of clear evidence of environmental conduction, it is set to 0.10. The value of the environmental conduction component is limited to between 0.10 and 0.60, and the direction is required to be consistent with the conduction type; otherwise, the component is treated as 0.10.

[0057] The historical correlation component is determined by time-delay alignment of historical state data sequences from both ends of the region, and satisfies a preset minimum sample size condition. Specifically, for the source and target regions of each directed edge, their regional-level historical state data sequences on the same state parameter are read. First, a unified time grid alignment is performed and missing / outlier points are removed. Then, time-delay alignment is performed within a preset time delay range: the preset time delay range uses a discrete time delay set, with values ​​of at least a portion of 0, 1, 2, 3, 6, 12, 24, and 48 hours, and is adaptively converted to the corresponding time steps according to the monitoring frequency. For each time delay, a set of aligned sample pairs is formed by "shifting the source sequence forward / aligning the target sequence backward," and the number of valid sample pairs is counted. The preset minimum sample size condition is set as follows: the number of valid sample pairs is not less than 30% of the number of unified time grid points and not less than a fixed lower limit, which is 48 or 96 (corresponding to a 10-minute sampling level of 2 days or 4 days, configurable). When a certain time delay meets the minimum sample size condition, the robust correlation strength between the source and the target under that time delay is calculated and the correlation direction is recorded. Robust association strength is achieved using an association metric insensitive to anomalies, employing Spearman rank correlation or truncated correlation as a fixed metric. The result with the largest absolute value of association strength among all time delays satisfying the sample size condition, and whose sign aligns with the propagation direction, is selected as the historical association result for that parameter. When none of the time delays satisfy the sample size condition or the association is insignificant, the historical association result for that parameter is set to 0 and written into the reason code. Finally, the historical association results of multiple parameters are fused according to preset weights to obtain historical association components. The weights are determined and normalized according to the following principles: lower missing rate, higher reliability, and higher matching degree with edge type. To ensure that the historical association components can be used for weight fusion, they are truncated to the range of 0-0.70, and negative correlations are either set to zero or recorded as separate suppression channels when considered as suppressive effects. Preferably, only non-negative influence weights are retained in the coupling matrix to maintain consistency in propagation interpretation.

[0058] The original influence weight is obtained by fusing historical correlation components with at least one static correlation component. Specifically, a fixed fusion structure is adopted, with historical correlation as the primary component and static correlation as the secondary component, and a fixed weight is set: when a directed edge has a historical correlation component that meets the sample size condition, the original influence weight is obtained by fusing the historical correlation component and the static correlation component. The fusion ratio is an empirically preset value. In this embodiment, the historical correlation component accounts for 0.6 and the static correlation component accounts for 0.4. The static correlation component is obtained by fusing spatial adjacency components, material continuity components, functional or structural dependence components, and environmental transmission components according to equal weight or edge type weight. Among them, the topology / structure dependence edge increases the weight of the functional or structural dependence component to the preset maximum allocation weight. (In this embodiment, a weight of 0.4 is preferred, with the maximum weight preset empirically.) The remaining weights are then evenly distributed among the other three categories. For environmental transmission edges, the weight of the environmental transmission component is increased to the preset maximum weight, and the remaining weights are evenly distributed among the other three categories. For pure spatial adjacency edges, the weight of the spatial adjacency component is increased to the preset maximum weight, and the remaining weights are evenly distributed among the other three categories. When there are no historical correlation components that meet the sample size condition for the directed edge, the original influence weight is determined entirely by the static correlation component, and its upper limit is preset empirically. In this embodiment, a weight of 0.55 is preferred to avoid excessive coupling of static information in the absence of historical evidence. The original influence weight is finally truncated to between 0 and 1 and is simultaneously written into the main contribution component type label for traceability.

[0059] Using each target region unit as the normalization object, normalization is performed on all original influence weights pointing to that target region unit, ensuring that the sum of the influence weights pointing to that target region unit satisfies a preset normalization condition. Specifically, for each target region unit, the original influence weights of all incoming edges are collected, and a normalization condition of summing to 1 is applied: the original influence weights of each incoming edge are scaled proportionally so that the sum of the scaled incoming edge weights equals 1; when the number of incoming edges is large, causing noise diffusion from low-weight edges, the original influence weights of the incoming edges are first truncated and filtered, retaining the top K incoming edges by weight and discarding the rest, where K is 3-8 and configurable, and the weight of the discarded edges is considered 0; then, normalization is performed on the retained edges; when the weight of an incoming edge exceeds the preset maximum single-edge weight limit after normalization, it is truncated and the remaining weight is proportionally distributed to the remaining incoming edges to prevent a single edge from completely dominating the target region.

[0060] When a region cell lacks incoming edge influence weights that meet certain conditions, a self-loop influence weight is set for the region cell to ensure that the coupling matrix of the region cell satisfies the normalization constraint and can be used for subsequent health prediction. Specifically, when the set of incoming edges of the target region cell is empty or all incoming edge original influence weights are 0 after screening, a self-loop edge of the region cell is set and assigned an original self-loop influence weight of 1; when there are incoming edges but their normalized sum is less than 1 (e.g., due to truncation or upper limit restrictions), the difference weight is compensated to the self-loop edge so that the final sum of incoming edge weights satisfies the normalization condition and the self-loop edge weight is not less than the preset minimum self-loop weight lower limit (e.g., 0.10), to ensure that the region retains at least its own state continuation contribution and avoids non-normalization or numerical instability in the coupling matrix.

[0061] Based on the normalized influence weights, a regional unit coupling matrix is ​​generated and output to characterize the influence relationship between regions. Specifically, the normalized influence weight of each directed edge is written into the corresponding row and column positions of the coupling matrix, where the matrix rows correspond to the source region and the matrix columns correspond to the target region being influenced; for region pairs without candidate edges, their matrix elements are set to zero; for each target region unit, its column vector (or set of incoming edge weights) is guaranteed to satisfy the preset normalization condition. During output, the candidate edge set, the component source label of each edge, the time delay selection result, and the sample size satisfaction are output simultaneously to form a traceable coupling matrix to construct an evidence chain.

[0062] The historical building health prediction process outputs health status prediction results at the regional unit level for future target time points or future target time windows, and further summarizes them to obtain the overall health prediction result for the historical buildings to be predicted. This process takes as input a set of regional units, a regional unit coupling matrix, a time series of regional-level multidimensional health status parameters, missing data markers, and quality markers, and outputs the health prediction results for each regional unit and the overall health prediction result. Figure 2 The illustrated health prediction model structure uses a regional-level input sequence within a time window as input. A time-series encoder composed of gated cyclic units encodes the sequence step-by-step, and outputs the encoded representation at the final time to drive multi-output head prediction. Specifically, the x below the figure... t-3 x t-2 x t-1 x tThis represents a series of consecutive time grid points aligned to a unified time grid. Each input consists of three types of information: the local state sequence channel, the neighborhood aggregation sequence channel, and the missing / quality marker channel (represented as self-region / neighborhood / marker in the figure), and is sent to the gated recurrent unit at the corresponding time. The gated recurrent units in the middle are connected in series along the time direction to form an encoding link. Solid arrows indicate the transmission of hidden states between adjacent time grid points; dashed horizontal arrows are used to emphasize that the transmission of hidden states belongs to the conceptual link of time recursion / sequence expansion, to distinguish it from input injection and output branch. At the rightmost time, the encoded representation of the last time moment is extracted as a sequence-level feature, and multi-task output is achieved through a fully connected output head: Output head 1 outputs a health index (risk / health level in the range of 0 to 1) via Sigmoid; Output head 2 outputs a health level (optional) via Softmax; Output head 3 outputs an uncertainty level (optional), used to provide graded hints on prediction confidence. The above multiple output heads share the same temporal encoder, thereby realizing joint modeling and inference of the index, level, and uncertainty under the same input caliber. The following provides specific guidelines for time alignment, missing / outlier labeling rules, neighborhood aggregation rules, input feature combination methods, health prediction model format, output definition, and the determination and attenuation rules for aggregated weights: For each regional unit, a multidimensional health status parameter sequence is collected within a unified analysis period. The health status parameters include at least one or more of crack width, tilt, settlement, vibration response, temperature, and humidity. Candidate spatial unit / monitoring point level observations are aggregated to the regional level using regional unit attribution mapping. Time alignment is achieved using a unified time grid method, i.e., a unified time step is determined during the initialization phase. The unified time step is taken as the main sampling interval of the monitoring system or an integer multiple thereof and limited to the range of 1 min to 60 min, preferably consistent with the output of the monitoring platform. Subsequently, window alignment is performed on each parameter sequence according to the unified time grid, i.e., for each time grid point, a robust center value is taken as the grid point value within the corresponding time window. Missing values ​​and / or outliers in the multidimensional health status parameter time series are generated with missing labels and / or quality labels to form a regional-level input sequence for health prediction.Specifically, the missing data marker adopts a binary labeling format. For each region, each parameter, and each time grid point, if there is no valid observation for that grid point within the aligned time window, the missing data marker is set to 1; otherwise, it is set to 0. The quality label adopts a multi-level labeling format and includes at least three categories: normal, suspected abnormal, and unusable. The suspected abnormal is used to identify situations where there are observations but the confidence is insufficient, including at least: the observed value exceeds the sensor range, the difference between adjacent grid points exceeds the preset jump threshold, several consecutive grid points maintain constant values ​​and are inconsistent with environmental changes, or there is a physical consistency conflict with other related parameters in the same region. For suspected abnormal grid points, they are used as missing data participation model inputs, and the quality label is synchronously written into the quality channel for use. Model differentiation: For parameter channels that are unusable for identifying the proportion of valid observations below a preset lower limit, the lower limit is set to 30% and is configurable. When the proportion of valid points of a parameter in a certain region within the entire input window is lower than this lower limit, the entire parameter channel is set to unusable and replaced with a fixed fill value and quality channel identifier during subsequent feature stitching. For each region unit, the multidimensional health status parameter time series is extracted from other region units associated with the non-zero influence weight corresponding to the region unit in the region unit coupling matrix. The extracted multidimensional health status parameter time series is then weighted and aggregated according to the preset neighborhood aggregation rules based on the non-zero influence weight to obtain the neighborhood aggregation representation corresponding to the region unit. Specifically, for any target region cell, the non-zero incoming edge weights pointing to that target region cell in the coupling matrix are read, and its neighborhood region set is determined accordingly. For each neighborhood region, the multidimensional health status parameter sequence and its missing / quality markers on the same time grid as the target region are extracted. Neighborhood aggregation adopts a determination rule of weighted averaging, missing renormalization, and quality decay: for each time grid point and each parameter dimension, the effective observations of the neighborhood region at that grid point are first screened, and only observations with normal quality markers are retained to participate in aggregation. When the number of neighborhood regions participating in aggregation is less than the preset minimum neighborhood number limit (e.g., less than 2), suspected abnormal observations are allowed to participate in aggregation with decay weights, and the decay ratio is 0.5. Then, the weighted average of the observations participating in aggregation is performed according to the incoming edge weights to obtain the neighborhood aggregation value, and the weights and sums caused by missing values ​​are renormalized so that the weights and sums participating in aggregation return to 1. If all neighborhood regions of that grid point are missing or unavailable, the neighborhood aggregation value is set to missing and the corresponding missing neighborhood marker is set to 1. This yields a neighborhood aggregation representation of the target region, whose dimensions are consistent with the multidimensional health status parameters of the target region. Additional quality features such as the neighborhood missing rate and the number of effective neighbors can be included as supplementary channels. The time series of multidimensional health status parameters of each regional unit are concatenated or combined with the corresponding neighborhood aggregation representation to construct the input features of the health prediction model.Specifically, a fixed splicing structure of local sequence channel - neighborhood aggregated sequence channel - marker channel is adopted as the input feature: the local sequence channel includes the aligned numerical sequence of each parameter; the neighborhood aggregated sequence channel includes the neighborhood aggregated value sequence of the corresponding parameter; the marker channel includes local missing markers, quality marker codes, and neighborhood missing markers. To enhance feasibility, the marker channel is numerated using a fixed coding table: missing markers are kept at 0 / 1; the quality markers encode normal, suspected abnormal, and unusable as 0, 1, and 2 respectively; and the numerical sequence of unusable channels is filled with the robust central value of the parameter across the entire building, while a quality marker of 2 explicitly indicates that the model ignores the numerical contribution of that channel. The input window length is implemented using a fixed backtracking window, that is, the most recent L time grid points before the target prediction time point are selected as the model input, where L is any fixed value of 24, 48, 72, or 144 and is determined by the time step to determine the backtracking duration, preferably covering at least 24 hours or at least 7 days of monitoring information to adapt to the gradual change process of historical buildings. The input features are input into the health prediction model, and the health prediction results of each regional unit are output. The health prediction model is implemented using a trainable sequence prediction model, with a fixed structure of a temporal encoder combined with an output head. The temporal encoder encodes the multi-channel sequence within the input window to extract time-dependent features, and the output head maps the encoded result to the prediction target. Specifically, the health prediction model can be implemented using a gated recurrent network, a temporal convolutional network, or an attention-based temporal network. In this embodiment, a gated recurrent network is used as the temporal encoder, and a fully connected layer is used as the output head to ensure stable training and inference even with missing and quality labels. The prediction target is output as a regional health index or risk value. The output type includes at least one of the following: a predicted health index point at a future target time, a trend of health index changes within a future time window, or a health level classification result. The numerical domain of the health index is fixed at 0-1 during the initialization phase, where 0 represents health and 1 represents high risk. The output result also includes a prediction confidence level or uncertainty label. The uncertainty label is determined by the input missing rate, the number of effective neighbors, and the model output stability, and is output with multi-level labels.

[0063] The health prediction results of each regional unit are weighted and summarized based on the regional criticality weight to obtain the overall health prediction result of the historical building to be predicted. The regional criticality weight is used to reflect the importance of different regions to the overall safety and function, and is generated by the method of determining the load-bearing role weight × disease level weight × quality decay factor. Specifically, the load-bearing role weight is determined by the load-bearing role of the dominant component in the region and adopts an empirically preset fixed mapping table. In this embodiment, it is preferred that the load-bearing area is 1.00, the enclosure area is 0.70, and the secondary area is 0.40. When there are multiple roles in the region, the weight of the largest role is taken and the composite reason code is recorded. The disease level weight is determined by the historical disease level of the region and adopts an empirically preset fixed mapping table. In this embodiment, it is preferred that Level I is 0.60, Level II is 0.80, Level III is 1.00, and Level IV is 1.20. When there are multiple types of diseases in the region, the weight corresponding to the highest level is taken and a small gain is added to the number of multiple diseases. The upper limit of the gain is limited to 0.10 to avoid extreme amplification. The quality decay factor is used to attenuate the aggregation weights of regions with low data quality during the aggregation process. The quality decay factor is determined by the proportion of valid observations and the proportion of suspected anomalies within the region's input window, using an empirically preset tiered mapping. In this embodiment, a preferred example is: 1.00 when the valid proportion is not less than 80% and the suspected anomaly proportion is not more than 10%; 0.80 when the valid proportion is between 50% and 80% or the suspected anomaly proportion is between 10% and 30%; 0.60 when the valid proportion is between 30% and 50% or the suspected anomaly proportion is between 30% and 50%; and 0.40 when the valid proportion is less than 30% or the suspected anomaly proportion is more than 50%, simultaneously marking the prediction result for that region as low confidence. The final regional criticality weight is obtained by multiplying the above three factors and performing normalization across all regions to ensure the weight sum is 1, guaranteeing that the overall health prediction result is interpretable and comparable. The overall health prediction results of the historical buildings to be predicted are obtained by weighted aggregation, that is, the overall health index is obtained by weighting and summing the health prediction results of each regional unit according to the regional key weight.

[0064] The health prediction model training process learns the mapping relationship between the region's own state, neighborhood coupling effects, and health outcomes from historical monitoring data and known health labels / surrogate labels. This ensures that stable and interpretable regional-level health prediction results can be output for future target time points during the inference phase. For example... Figure 3As shown, the overall framework diagram for the training and deployment of the health prediction model is presented, organized into three columns: training data and sample organization, model training, and training output (deployment package). The training data and sample organization includes sample construction (windowing), label generation (supervision signal), and standardization and solidification: windowing is performed using a fixed input window length L and prediction step size H, and sample validity conditions are set (the effective proportion of key parameters is not less than 50% and the average effective neighborhood number is not less than 2); health index or health level labels are generated from disease level / density / change rate and identification conclusions, and the sample weights are adjusted according to the credibility of the source. Robust centers and scales are statistically determined for the training set and solidified into parameters for inference reuse. Missing / quality labels are used as independent inputs to the model training process. The local state sequence, neighborhood aggregation sequence, and missing / quality labels are concatenated to form input features, which are then passed through a gated recurrent network as a temporal encoder and connected to a fully connected layer to output at least one of the following: predicted value, trend, or rank. During training, a loss function is constructed based on the task type, including the mean absolute error (MAE) and high-risk weighting for exponential prediction, and the cross-entropy and sample weights for rank classification. The loss is then decayed according to a missing rate mapping table. Optimization is then achieved through mini-batch iteration, gradient pruning, learning rate decay, and early stopping strategies. Training outputs and deployment constraints: Evaluation and selection are based on metrics such as MAE and trend consistency, macro F1 and high-risk recall. Model parameters, standardized parameters, quality label encoding table, L (input window length) / H (prediction step size) configuration, and neighborhood aggregation and normalization rules are solidified into a deployment package. Simultaneously, consistency constraints are set to ensure that the inference end strictly reuses the time alignment, label generation, aggregation / concatenation, and standardized parameters from the training end, guaranteeing consistency between training and inference. The following provides a defined approach to training data construction, label generation, sample organization, training objectives, loss function, training strategy, early stopping, and deployment consolidation: S1, Constructing the training dataset. Specifically, select at least one historical building to be predicted or multiple similar historical buildings as training sources. Collect regional-level multidimensional health status parameter time series within a unified analysis period, and complete time alignment, missing label generation, and quality label generation. Perform neighborhood aggregation on each region based on the regional unit coupling matrix to obtain the neighborhood aggregation representation. Organize training samples using a fixed input window length L and a fixed prediction step size H. That is, for each target regional unit, the input window is formed by the L nearest time grid points back from time t, and the health result corresponding to the Hth time grid point after time t (or the target time window after t) is used as the supervision label, thus forming a sample pair of input features and output labels. To ensure that the samples are trainable, the following sample validity conditions are set: the effective observation ratio of key parameters in the input window is not less than 50%, and the time average of the effective neighborhood number of neighborhood aggregation is not less than 2; sample pairs that do not meet the sample validity conditions are removed or used only for unsupervised pre-training but not for supervised loss calculation.

[0065] S2, Generate Training Labels. Training labels adopt one or a combination of two forms: health index labels and / or health level labels. Firstly, health index labels are generated from historical damage records, repair records, assessment conclusions, and monitoring anomalies. Damage levels, damage density, crack growth rate, settlement rate, tilt changes, and vibration anomalies in the area within the same time period are mapped to a 0-1 range according to preset rules to obtain a health index. Damage levels are assigned values ​​according to a fixed mapping table and serve as the dominant term, while monitoring anomalies serve as correction terms. When a structural assessment report provides a component or area safety level, the assessment conclusion is prioritized for mapping the health index, and monitoring anomalies are used for minor corrections, with the correction range limited to ±0.10 to maintain label stability. Secondly, health level labels are obtained by discretizing the health index according to a fixed threshold preset by experience. For example, a health index less than 0.25 is Level I, 0.25-0.50 is Level II, 0.50-0.75 is Level III, and greater than or equal to 0.75 is Level IV. To address label noise, labels with lower source credibility are assigned lower sample weights. The sample weights are determined in the following order: identification conclusion > testing record > repair and acceptance > long-term monitoring inference > subjective record of on-site inspection. The sample weight of the lowest credibility label is reduced to 0.5.

[0066] S3 performs data standardization and feature solidification on the training samples. Specifically, robust centers and robust scales of each state parameter are statistically analyzed within the training set and solidified into a standardized parameter file. The same set of standardized parameters is used to scale the input sequences during both the training and inference phases. Missing labels and quality labels are directly input into the model as independent channels without standardization. For unusable channels, a fixed padding strategy is adopted, that is, the parameter is filled with the robust center value within the training set, and the unusability is indicated by the quality label at the same time to avoid the model misidentifying the padding value as the real observation.

[0067] S4. Determine the training objective and loss function. If the output is a health index point prediction, the training objective is to minimize the error between the predicted value and the label value. The mean absolute error is used as the main loss, and a risk weighting factor is added to high-risk samples to improve sensitivity to degraded regions. The risk weighting factor is categorized according to the size of the health index. That is, the database pre-stores a mapping table corresponding to the health index and the risk weighting factor. When using it, the health index is used as the query key to extract the risk weighting factor. If the output is a health level classification, the training objective is to minimize the classification error. Cross-entropy is used as the main loss, and a weighting coefficient is introduced for sample weight and class imbalance. When both the health index and health level are output, multi-task training is used, and the two types of losses are weighted and summed according to a fixed ratio preset by experience. In this example, 0.6 and 0.4 are preferred. To suppress the interference of missing and outliers on training, a quality masking mechanism is used to attenuate the loss at the sample level. Based on the missing rate of the input window, the corresponding weight scaling factor is extracted from a pre-defined missing rate-weight scaling factor mapping table, and the loss weight is reduced according to the weight scaling factor. For example, when the missing rate of the input window exceeds 50%, the loss weight of the sample is multiplied by 0.7, and when the missing rate exceeds 70%, it is multiplied by 0.4. The sample is then marked as a low-confidence sample to reduce its impact on parameter updates.

[0068] S5 executes the training and validation partitioning and iterative training. Specifically, a time-sequential partitioning method is used to construct the training and validation sets, with samples from earlier time periods used as the training set and samples from later time periods used as the validation set to avoid information leakage. When there is data from multiple historical buildings, a "reserve by building" cross-validation method is used, reserving all data from some buildings for validation to test the model's cross-building generalization ability. During training, mini-batch iterative updates of model parameters are used, and gradients are pruned to improve training stability. A fixed learning rate plan is adopted with a learning rate decay rule: if the validation set metric shows no improvement within 3 consecutive evaluation periods, the learning rate is multiplied by 0.5; if the validation set metric shows no improvement within 6 consecutive evaluation periods, early stopping is triggered, and the model parameters are rolled back to the optimal validation metric.

[0069] S6 sets up model selection and overfitting suppression strategies. Specifically, a comprehensive validation set index is used as the basis for model selection: health index prediction is jointly evaluated using the validation set mean absolute error and trend consistency index, while health level classification is jointly evaluated using the macro-average F1 score and high-risk recall rate; a "high-risk priority" constraint is set, meaning that when the comprehensive indices are similar, the model with the higher high-risk recall rate is selected first. To suppress overfitting, a weight decay and random deactivation strategy is used during training, and random deactivation is applied only to the numerical channels, not to the missing label and quality label channels, to avoid destroying the label semantics.

[0070] S7, after training, solidify the model and inference configuration. Specifically, the trained model parameter file, standardized parameter file, quality label encoding table, input window length L, prediction step size H, neighborhood aggregation rule parameters, and coupling matrix normalization method are all solidified into a model configuration package. During the inference phase, input features are generated according to the same time alignment, label generation, neighborhood aggregation, and feature concatenation process as in the training phase, and the model configuration package is loaded to output the prediction results. To ensure the interpretability of the output, sample-level quality assessment results are output simultaneously, including the input window missing rate, the number of effective neighbors, the prediction uncertainty level, and a list of the neighborhood regions with the highest contribution, used to indicate and trace the credibility of the prediction results.

[0071] It should be noted that the specific values ​​in the above-mentioned confirmation criteria are only specific examples of this embodiment. In actual implementation, specific settings can be made based on different application buildings or different application purposes. This invention does not limit these settings.

[0072] Based on Example 1, to ensure that the encoding rules are implementable and maintainable, missing and conflict handling rules are also generated during the encoding rule pre-setting stage: when a field is missing, a default code is selected according to the field data type (for example, continuous types use missing indicator bit-neutral value, categorical types use "unknown / not recorded" enumeration code, and time types use "cannot be determined" mark and reduce confidence); when there are conflicting values ​​in the same field, they are not directly forcibly merged during the encoding stage, but multiple candidate codes are generated in parallel or a dual-channel representation of principal value-conflict intensity is generated, where the principal value is selected according to the source confidence and formation time, and the conflict intensity is calculated by the number of conflicting entries, the difference magnitude and the degree of confidence divergence, so that the subsequent partitioning algorithm can avoid the misclassification of high conflict units under constraints.

[0073] Based on the unchanged aspects of Embodiment 1 or Embodiment 2, and under the premise of satisfying the partitioning constraints, the partitioning algorithm used to obtain the initial set of regional units can be implemented by, in addition to the aforementioned hard-constrained region growth based on adjacency graphs, partitioning algorithms including but not limited to constraint hierarchical clustering partitioning algorithms, constraint graph partitioning partitioning algorithms, constraint spectrum clustering partitioning algorithms, and partitioning algorithms based on multi-objective constraint optimization. All of these algorithms use the feature vectors of candidate spatial units as similarity criteria and the adjacency graphs of candidate spatial units as connectivity constraint carriers. During the partitioning process, the minimum / maximum spatial scale constraints, minimum discriminative constraints, topological connectivity constraints, and observability constraints are verified and constrained.

[0074] The constraint graph partitioning algorithm uses candidate spatial units as graph nodes, adjacency relationships as edges, and assigns edge weights. Edge weights are determined by the similarity of the candidate spatial unit's feature vectors (higher similarity results in higher edge weights or higher cutting costs). The graph is divided into several connected subgraphs as initial region units using minimum cut / normalized cut or multi-way partitioning strategies. During partitioning, upper and lower limits of the spatial scale are used as capacity constraints for each subgraph, and an observability threshold is used as a feasibility constraint for the subgraph. Infeasibility partitioning is addressed through backtracking and pairing. Figure 2 Subgraphs that violate constraints are eliminated by repairing and splitting or merging and then splitting subgraphs, thereby obtaining an initial set of region cells that satisfy the partitioning constraints.

[0075] The constrained spectral clustering partitioning algorithm constructs a similarity matrix based on the feature vectors of candidate spatial units and uses an adjacency graph to mask or reduce the weight of the similarity of non-adjacent units, making spectral clustering naturally tend to form connected regions. After obtaining the initial cluster labels, a connectivity test is performed on the candidate spatial units within each cluster. If non-connected components are found, they are split into multiple connected sub-clusters according to the adjacency graph. The spatial scale and observability of the split sub-clusters are calculated. If the constraints are not met, inter-cluster migration is performed or the sub-cluster is merged with the nearest and mergeable adjacent sub-cluster. At the same time, the discriminability between adjacent sub-clusters is checked, and merging of sub-clusters with discriminability below the threshold is prohibited, thus outputting the initial set of region units.

[0076] like Figure 4 As shown, this application also provides a flowchart of a health prediction method for historical buildings, including: obtaining the original information set of the historical building to be predicted, constructing each candidate spatial unit of the historical building to be predicted, introducing zoning constraints, dividing the historical building to be predicted into regions, and obtaining an initial set of regional units.

[0077] Obtain the state dataset within the coverage area of ​​the initial region unit set, construct the internal heterogeneity index of the initial region unit set and the difference distance of each adjacent initial region unit pair, and perform subdivision and / or merging of the initial region unit set to obtain each region unit.

[0078] A set of directed candidate edges is established based on each regional unit. The influence weights between regional units in the set of directed candidate edges are calculated and normalized to obtain the regional unit coupling matrix used to characterize the influence relationship between regions.

[0079] The time series of multidimensional health status parameters of each regional unit are collected. Based on the coupling matrix of the regional unit, the neighborhood aggregation representation of each regional unit is constructed. The time series of multidimensional health status parameters of each regional unit and the neighborhood aggregation representation of each regional unit are used as inputs to the health prediction model. The health prediction results of each regional unit are output. The health prediction results of each regional unit are summarized and processed to obtain the overall health prediction result of the historical building to be predicted.

[0080] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0081] It should be understood that determining B based on A does not mean determining B solely based on A; it also means determining B based on A and / or other information.

[0082] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0083] The above description is only an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A health prediction system for historical buildings, characterized in that, The system includes: The initial regional unit division module for historical buildings is used to obtain the original information set of the historical buildings to be predicted, construct each candidate spatial unit of the historical buildings to be predicted, introduce partitioning constraints, divide the historical buildings to be predicted into regions, and obtain the initial regional unit set. The historical building area unit determination module is used to obtain the state dataset within the coverage area of ​​the initial area unit set, construct the internal heterogeneity index of the initial area unit set and the difference distance of each adjacent initial area unit pair, and thereby perform the subdivision and / or merging of the initial area unit set to obtain each area unit; The regional unit coupling matrix construction module is used to establish a set of directed candidate edges based on each regional unit, calculate the influence weights between regional units for the regional units in the set of directed candidate edges and normalize them to obtain the regional unit coupling matrix used to characterize the influence relationship between regions. The historical building health prediction module is used to collect the time series of multidimensional health status parameters of each regional unit, construct the neighborhood aggregation representation of each regional unit based on the regional unit coupling matrix, take the time series of multidimensional health status parameters of each regional unit and the neighborhood aggregation representation of each regional unit as input to the health prediction model, output the health prediction results of each regional unit, summarize and process the health prediction results of each regional unit, and obtain the total health prediction result of the historical building to be predicted. The regional unit coupling matrix construction module includes: obtaining the topological connectivity and / or spatial adjacency relationships between each regional unit, filtering regional unit pairs based on the topological connectivity and / or spatial adjacency relationships, and organizing the filtered regional unit pairs into a set of directed candidate edges according to the influence direction; For each directed edge in the set of directed candidate edges, a historical correlation component is calculated based on the historical state data sequence of the corresponding two end region units. The historical correlation component is used as the historical state data correlation information. The historical correlation component is determined by time delay alignment within a preset time delay range and meets the preset minimum sample size condition. The historical correlation components are fused with at least one static correlation component to obtain the original influence weight. The static correlation components include one or more of the following: spatial adjacency components, material continuity components, functional or structural dependence components, and environmental transmission components. Using each target region unit as the normalization object, normalization processing is performed on all the original influence weights pointing to that target region unit, so that the sum of the influence weights pointing to that target region unit satisfies the preset normalization condition. Generate and output the regional unit coupling matrix to characterize the influence relationship between regions based on the normalized influence weights; When a region cell does not have an incoming edge influence weight that meets the conditions, a self-loop influence weight is set for the region cell to ensure that the coupling matrix of the region cell meets the normalization constraint and can be used for subsequent health prediction.

2. The health prediction system for historical buildings as described in claim 1, characterized in that, The process of obtaining the original information set of the historical building to be predicted and constructing candidate spatial units for the historical building to be predicted includes: Obtain a set of original information related to the historical building to be predicted, wherein the set of original information includes at least one or more of the following: design data, historical archives, inspection records, monitoring system data, and on-site survey results; Based on the original information set, the physical spatial boundary of the historical building to be predicted is determined, and the set of structural component types and component topological connection relationships on which the candidate spatial unit generation is based are determined. A candidate spatial unit set is generated based on the set of structural component types and the topological connection relationship of the components. The candidate spatial unit includes component units and / or component sub-units. The component units include at least one or more of the following: beams, columns, walls, arches, floor slabs, roofs, and foundation structures. For each candidate spatial unit in the candidate spatial unit set, a feature vector for partitioning is constructed to obtain the feature vector of each candidate spatial unit. The feature vector includes at least one or more of the following: spatial or geometric features, material or structural features, functional or stress-bearing role features, environmental conduction features, and disease or repair features.

3. The health prediction system for historical buildings as described in claim 2, characterized in that, For component units with a spatial scale greater than a preset scale threshold, the component units are subdivided according to preset geometric subdivision rules to obtain component sub-units, and the component sub-units are incorporated into the candidate spatial unit set.

4. A health prediction system for historical buildings as described in claim 1, characterized in that, The introduction of zoning constraints involves dividing the historical buildings to be predicted into regions to obtain an initial set of regional units. The specific steps are as follows: Partitioning constraints are set based on the candidate spatial unit set. The partitioning constraints include at least the minimum spatial scale constraint and / or maximum spatial scale constraint of the regional unit, the minimum distinguishability constraint between regional units, the topological connectivity constraint of the regional unit, and the observability constraint. Establish a coverage mapping from monitoring points to a set of candidate spatial units, and calculate the observability index of each candidate spatial unit based on the coverage mapping. Set a minimum observability threshold. The observability index of the initial regional unit is obtained from the observability index of the candidate spatial units it contains according to a preset aggregation rule, and the observability constraint is limited to the observability index of the initial regional unit being no less than the minimum observability threshold. Under the premise of satisfying the partitioning constraints, the partitioning algorithm is invoked to perform region partitioning on the candidate spatial unit set based on the feature vector of the candidate spatial unit set to obtain an initial regional unit set, wherein each initial regional unit is composed of one or more candidate spatial units. Output the initial set of regional units and establish the attribution mapping relationship between the initial set of regional units and the candidate set of spatial units.

5. A health prediction system for historical buildings as described in claim 1, characterized in that, The historical building area unit determination module specifically includes: For the initial set of regional units, based on the attribution mapping relationship between the initial regional units and the candidate spatial unit set, a state dataset within the coverage area of ​​the initial set of regional units is obtained. The state dataset is a regional-level data sequence that has undergone time alignment and missing label processing and can be used for statistical calculations. For each initial region unit, the robust central value of each state parameter within the initial region unit is calculated based on its state dataset, and normalization is performed on the state parameters of different dimensions. Based on the normalized state dataset, the internal heterogeneity index of the initial region unit is calculated. The internal heterogeneity index characterizes the degree of fluctuation and dispersion of the state parameters within the initial region unit. Determine adjacent initial region unit pairs, wherein the adjacent initial region unit pairs are initial region units that are directly connected in terms of topological connectivity; For any pair of adjacent initial region units, a center value difference term for each state parameter is constructed based on the difference in their robust center values. The center value difference terms are then fused to obtain the difference distance. A topology penalty term is introduced to correct the difference distance, thus obtaining the difference distance for each pair of adjacent initial region units. When the internal heterogeneity index of any initial region unit is greater than the preset subdivision threshold, the initial region unit is subdivided according to the preset subdivision rule to generate at least two sub-region units. When the difference distance between any two adjacent initial region units is less than a preset merging threshold, the adjacent initial region units are merged to obtain merged region units. After performing the subdivision and / or merging processes, the region units and their affiliation mapping with the candidate spatial unit set are updated, and each region unit is output when the preset iteration stop condition is met. The iteration stopping conditions include at least one of the following: no subdivision processing and no merging processing is triggered in one iteration, or the number of iterations reaches a preset upper limit.

6. A health prediction system for historical buildings as described in claim 1, characterized in that, The historical building health prediction module specifically includes: Collect time series of multidimensional health status parameters for each regional unit and perform time alignment; Missing values ​​and / or outliers in the time series of the multidimensional health status parameters are generated with missing labels and / or quality labels to form a regional input sequence for health prediction. Based on the region unit coupling matrix, for each region unit, the time series of multidimensional health status parameters of other region units corresponding to the region unit in the region unit coupling matrix are extracted. The extracted multidimensional health status parameter time series are then weighted and aggregated according to the non-zero influence weights based on the preset neighborhood aggregation rules to obtain the neighborhood aggregation representation corresponding to the region unit. The multidimensional health status parameters time series of each regional unit are spliced ​​or combined with the corresponding neighborhood aggregation representation to construct the input features of the health prediction model. The input features are fed into the health prediction model, and the health prediction results for each regional unit are output.

7. A health prediction system for historical buildings as described in claim 6, characterized in that, The historical building health prediction module also includes: The health prediction results of each regional unit are weighted and summarized based on the regional key weights to obtain the overall health prediction result of the historical building to be predicted. The regional criticality weight is determined by the component stress role, historical disease level and / or quality mark of the regional unit, and the aggregation weight of regional units with low data quality is attenuated during the weighted aggregation process.

8. A health prediction method for historical buildings, applied to a health prediction system for historical buildings as described in any one of claims 1-7, characterized in that, The method includes: Obtain the original information set of the historical buildings to be predicted, construct each candidate spatial unit of the historical buildings to be predicted, introduce zoning constraints, divide the historical buildings to be predicted into regions, and obtain the initial set of regional units. Obtain the state dataset within the coverage area of ​​the initial region unit set, construct the internal heterogeneity index of the initial region unit set and the difference distance of each adjacent initial region unit pair, and perform subdivision and / or merging of the initial region unit set to obtain each region unit; A set of directed candidate edges is established based on each regional unit. The influence weights between regional units in the set of directed candidate edges are calculated and normalized to obtain the regional unit coupling matrix used to characterize the influence relationship between regions. The time series of multidimensional health status parameters of each regional unit are collected. Based on the coupling matrix of the regional unit, the neighborhood aggregation representation of each regional unit is constructed. The time series of multidimensional health status parameters of each regional unit and the neighborhood aggregation representation of each regional unit are used as inputs to the health prediction model. The health prediction results of each regional unit are output. The health prediction results of each regional unit are summarized and processed to obtain the overall health prediction result of the historical building to be predicted.