A building crack time sequence detection method

CN122637007BActive Publication Date: 2026-09-29TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611063608.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-29
Estimated Expiration
2046-07-17

AI Technical Summary

Technical Problem

[0006]根据本发明的上述技术方案,针对建筑表面复杂纹理易误检裂缝、多时相检测精度差的问题,先完成多时相图像配准,通过像素级材质分区实施差异化纹理抑制,从源头减少与裂缝相似性高的复杂纹理的假阳性干扰;再根据材质分区自适应开展材质感知裂缝检测,并融合检测框、材质分类信息、历史时相裂缝掩码等多模态特征做像素级分割

Benefits of technology

[0006]根据本发明的上述技术方案,针对建筑表面复杂纹理易误检裂缝、多时相检测精度差的问题,先完成多时相图像配准,通过像素级材质分区实施差异化纹理抑制,从源头减少与裂缝相似性高的复杂纹理的假阳性干扰;再根据材质分区自适应开展材质感知裂缝检测,并融合检测框、材质分类信息、历史时相裂缝掩码等多模态特征做像素级分割。可精准区分固有纹理与真实裂缝,降低误检漏检。同时,基于累积置信度跟踪裂缝演化,可有效提升古建筑多时相裂缝检测精度、分割精细度与检测可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122637007B_ABST
    Figure CN122637007B_ABST
Patent Text Reader

Abstract

The application provides a building crack timing detection method, and relates to the technical fields of image processing and structural health detection. The method comprises the following steps: acquiring a multi-temporal image sequence of a building surface; performing spatial registration on non-reference temporal images based on the spatial coordinates of the reference temporal image; performing pixel-by-pixel material classification on the reference temporal image based on semantic segmentation to generate a material partition map; performing differentiated texture suppression preprocessing on different material regions in the registered temporal images according to the material partition map, wherein the preprocessing is matched with the texture interference type of the material region; performing material-aware crack detection on the preprocessed temporal images based on the material partition map, and combining the detection threshold determined for each known region divided from the preprocessed temporal images to obtain a crack detection frame set of the preprocessed temporal images; and performing pixel-level segmentation on the current temporal crack in the current temporal image in the preprocessed current temporal image according to the multi-modal fusion prompt feature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing and structural health detection technology, and more specifically, to a method for time-series detection of building cracks. Background Technology

[0002] As important cultural heritage, the structural safety of ancient buildings needs to be assessed through long-term inspections and change monitoring. Cracks are one of the most common and obvious forms of structural damage to ancient buildings, and multi-temporal tracking and monitoring of cracks is a key technical aspect of preventive protection for ancient buildings.

[0003] With the development of computer vision technology, image-based crack detection methods have made significant progress, mainly including methods based on traditional image processing, end-to-end methods based on semantic segmentation, and methods combining object detection and instance segmentation. General crack detection methods are primarily designed for relatively simple backgrounds such as roads, bridges, or ordinary building walls. Ancient building surfaces typically possess complex textures such as wood grain, brick joints, mortar joints, weathered and peeling edges, gray plaster pattern boundaries, and painted pattern boundaries. These textures have a strong visual similarity to real cracks in images. When general crack detection methods are directly applied to ancient building surfaces, these inherent textures are easily misdetected as cracks, leading to decreased detection performance and making it difficult to meet the accuracy and reliability requirements of multi-temporal detection. Summary of the Invention

[0004] In view of this, the present invention provides a method for time-series detection of building cracks.

[0005] One aspect of the present invention provides a method for temporal detection of building cracks, comprising: acquiring a multi-temporal image sequence of a building surface, the multi-temporal image sequence including a reference temporal image and non-reference temporal images; spatially registering the non-reference temporal images based on the spatial coordinates of the reference temporal images to obtain a registered multi-temporal image sequence; performing pixel-by-pixel material classification on the reference temporal images based on semantic segmentation to generate a material partition map; performing differential texture suppression preprocessing matching the texture interference type of the material region on different material regions in each registered temporal image according to the material partition map, and outputting a preprocessed multi-temporal image sequence; performing material-aware crack detection on each preprocessed temporal image based on the material partition map, and combining the detection threshold determined for each known region obtained by dividing each preprocessed temporal image to obtain a set of crack detection boxes for each preprocessed temporal image, each crack detection box having a crack detection confidence level; and performing multimodal fusion. The system provides a feature that performs pixel-level segmentation of the current-phase cracks within the current-phase crack detection box of the preprocessed current-phase image, generating a current-phase crack instance mask to obtain the current-phase crack detection result on the building surface. The multimodal fusion feature incorporates at least the features of the current-phase crack detection box and the material classification information of the corresponding region of the current-phase crack detection box, determined based on the material partitioning map. If the current-phase crack detection box is associated with historical-phase crack instances, the multimodal fusion feature also incorporates the morphological features of the corresponding historical-phase crack instance mask. Historical-phase crack instances have cumulative confidence scores, and these instances and their corresponding cumulative confidence scores are stored in a historical instance database. This database stores cross-temporal detection information for all crack instances and is updated based on the current-phase crack detection result for subsequent temporal threshold adjustment and pixel-level segmentation prompt generation.

[0006] According to the above-mentioned technical solution of the present invention, addressing the problems of false detection of cracks due to complex textures on building surfaces and poor multi-temporal detection accuracy, multi-temporal image registration is first completed. Differential texture suppression is implemented through pixel-level material partitioning to reduce false positive interference from complex textures highly similar to cracks at the source. Then, material-aware crack detection is adaptively performed based on material partitioning, and multi-modal features such as detection boxes, material classification information, and historical crack masks are fused for pixel-level segmentation. This can accurately distinguish between inherent textures and real cracks, reducing false detections and missed detections. Simultaneously, tracking crack evolution based on cumulative confidence can effectively improve the detection accuracy, segmentation fineness, and detection reliability of multi-temporal cracks in ancient buildings. Attached Figure Description

[0007] The above and other objects, features and advantages of the present invention will become more apparent from the following description of embodiments of the invention with reference to the accompanying drawings, in which:

[0008] Figure 1 An exemplary system architecture for applying a temporal detection method for building cracks according to an embodiment of the present invention is shown;

[0009] Figure 2A A main flowchart of a time-series detection method for building cracks according to an embodiment of the present invention is shown;

[0010] Figure 2B A general flowchart of a time-series building crack detection method based on material adaptive preprocessing and historical instance library feedback according to an embodiment of the present invention is shown.

[0011] Figure 3 An architecture diagram of a material-aware crack detection model according to an embodiment of the present invention is shown;

[0012] Figure 4 A schematic diagram of an adaptive detection threshold mechanism according to an embodiment of the present invention is shown;

[0013] Figure 5 An architecture diagram of a cross-modal cue fusion network according to an embodiment of the present invention is shown;

[0014] Figure 6 A schematic diagram of a closed-loop feedback mechanism for a historical instance library according to an embodiment of the present invention is shown;

[0015] Figure 7 A block diagram of a building crack timing detection device according to an embodiment of the present invention is shown;

[0016] Figure 8 A block diagram of an electronic device suitable for implementing a time-series detection method for building cracks according to an embodiment of the present invention is shown. Detailed Implementation

[0017] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0020] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0021] The surface of ancient buildings usually contains areas of multiple materials, and the interference caused by different material areas is different.

[0022] The main interference on the surface of wood structures is the wood grain, which appears as directional dark stripes in the image and is visually highly similar to real cracks extending along the grain direction. Wood grain has a clear main directional feature, and regular textures along this direction can be directionally suppressed; however, relevant general crack detection procedures usually do not combine the wood structure material identification results to estimate the wood grain orientation field and select the corresponding directional suppression strategy, which leads to false positives in crack detection due to wood grain.

[0023] The main disturbances on masonry surfaces are brick joints and mortar joints, which exhibit periodicity, continuity, and regular arrangement, closely resembling crack morphology in edge detection results. While the periodicity and equidistant arrangement of brick joints are identifiable structural features, relevant crack detection procedures typically fail to consider the masonry material region to identify and suppress these regular linear structures, making it difficult to reliably distinguish them from genuine irregular crack structures.

[0024] The main interferences on plaster / plaster / painted surfaces are weathered and peeling edges and painted pattern boundaries. These non-crack edge interferences are visually difficult to distinguish from low-contrast fine cracks. True cracks on such surfaces often have low contrast and require targeted enhancement for effective detection; however, general enhancement strategies typically do not incorporate the edge morphology characteristics of the plaster / plaster / painted material areas, making it difficult to simultaneously filter out non-crack edge interferences and improve the detectability of low-contrast cracks.

[0025] In summary, different material areas exhibit different dominant interference mechanisms: wood grain represents directional stripe interference, requiring directional suppression; brick and mortar joints represent periodic linear structure interference, requiring elimination based on regularity identification; weathered and peeling edges and pattern boundaries in stucco / plaster / painted areas represent non-crack edge interference, requiring filtering based on morphological characteristics. When these multiple material areas coexist on the same building surface, if they are treated uniformly without distinguishing material types, it becomes difficult to simultaneously address these significantly different interference types. This is one of the important reasons for the low accuracy of crack detection in ancient buildings.

[0026] Furthermore, the detection of cracks in ancient buildings requires repeated testing over a long period of time, and this temporal requirement further exposes the shortcomings of relevant methods in utilizing historical detection information.

[0027] Detecting cracks in ancient buildings requires repeated, long-term inspections of the same building surface, with significant temporal correlations between adjacent time phases. However, current detection methods perform detection and segmentation independently in each time phase, without utilizing historical observation results as historical detection information, which presents the following three problems.

[0028] First, fixed detection thresholds are difficult to adapt to the needs of temporal detection. For example, single-stage real-time object detection algorithms (such as the YouOnly Look Once (YOLO) series) detect independently in each time phase, using a uniform detection confidence threshold and not utilizing detection results from historical time phases. There are two contradictory needs in the temporal detection of cracks in ancient buildings. On the one hand, for cracks already confirmed in historical time phases, a lower detection threshold is needed for finer searching to capture minor expansions, new branches, or subtle width changes. On the other hand, for areas where no cracks were found in historical time phases, a standard detection threshold needs to be maintained; lowering the threshold in such areas increases the risk of misidentifying background textures such as wood grain and brick seams as cracks. A uniform threshold cannot simultaneously satisfy the opposing needs of "high sensitivity in known areas" and "low false detection rate in unknown areas."

[0029] Secondly, the segmentation cues are simplistic and fail to utilize the temporal continuity of crack morphology. For example, cue-driven instance segmentation models such as the SegmentAnything Model (SAM) series perform segmentation independently in each temporal phase, using only the current bounding box as a spatial cue. Ancient building cracks exhibit significant temporal continuity; that is, the morphological changes of the same crack between adjacent temporal phases are usually gradual, and the segmentation results of historical phases contain crucial morphological information such as the crack's precise location, direction, and branching structure. However, these methods segment independently in each temporal phase, failing to utilize this historical morphological information.

[0030] Third, false detections are difficult to correct automatically. In multi-phase detection, if a certain phase misidentifies brick joints or wood grain as cracks, the false detection record will remain in the detection data because the relevant methods do not perform reliability assessments on recorded crack instances. When subsequent phases are associated and matched, they will continue to associate based on this erroneous record, causing the error to accumulate with the number of detections rather than being gradually eliminated.

[0031] The core idea of ​​this invention is to optimize the detection of cracks in ancient buildings in multiple time phases. On the one hand, it achieves differentiated texture suppression through material partitioning to reduce interference from the source. On the other hand, it uses the detection results of historical time phases (including which locations have cracks and the precise shape of the cracks) as reference information to guide the detection and segmentation of the current time phase.

[0032] 1. Layered anti-interference mechanism for material perception.

[0033] (1) Material identification and zoning. Using the reference time-phase image as input, the semantic segmentation model is used to classify the material pixel by pixel, and the surface of the ancient building is divided into three material regions: wooden structure surface, brick and stone masonry surface, and gray plastic / plaster / painted surface, generating a material zoning map.

[0034] (2) Differentiated texture suppression. Based on the material partitioning map, suppression strategies matching the texture interference type are applied to different material regions to reduce the interference of the inherent texture of various materials on crack detection from the source.

[0035] (3) Layered anti-interference and material reuse. In the preprocessing stage, textures are suppressed in a targeted manner according to the material type. In the detection stage, material information is further utilized to make the model more resistant to residual interference, forming a layered anti-interference. The material partition map is used multiple times in the whole process (texture suppression, detection enhancement, segmentation hints), with one partition and multiple reuses.

[0036] 2. Historical information-driven detection and segmentation closed loop.

[0037] After each detection and segmentation is completed, the results are recorded in a historical instance database. During the next phase of processing, historical information is retrieved from the database to guide the current detection and segmentation. This involves three aspects.

[0038] (1) Adjust the detection sensitivity based on historical detection results. For crack areas that have been confirmed to exist in historical time phases, lower the detection threshold to improve sensitivity and capture small expansions; for areas where no cracks have been found, maintain the standard threshold to control false detections. Maintain a cumulative confidence level for each crack instance, and dynamically adjust the detection threshold for each area accordingly.

[0039] (2) Utilizing historical morphology to assist current segmentation. In the instance segmentation stage, the detection box of the current phase is used as a spatial range indicator. For known cracks, the precise segmentation contour of the crack in the historical phase is also introduced as a morphological reference, while the material information of the region is integrated. The detection box determines the spatial range of the segmentation, the historical contour provides the expected shape of the crack, and the material information tells the segmentation model what material the region belongs to, so that the model can refer to the interference patterns unique to that material when segmenting.

[0040] (3) Confidence evolution and false detection elimination. If a crack is detected and a corresponding record can be found in the historical instance database, the confidence of the instance is increased; if a historical crack is not detected in the current phase, its confidence is decreased. Instances that have not matched for a long time and have low confidence are marked as false detections and eliminated, so that the detection system gradually improves itself as the number of observations increases. New detection results are then updated back to the historical instance database, forming a phase-by-phase feedback optimization.

[0041] In view of this, the present invention constructs a temporal detection method for building cracks based on the above core ideas. It includes: acquiring a multi-temporal image sequence of a building surface, the multi-temporal image sequence including a reference temporal image and non-reference temporal images; spatially registering the non-reference temporal images based on the spatial coordinates of the reference temporal image to obtain a registered multi-temporal image sequence; performing pixel-by-pixel material classification on the reference temporal image based on semantic segmentation to generate a material partition map; performing differential texture suppression preprocessing matching the texture interference type of different material regions in each registered temporal image according to the material partition map, outputting a preprocessed multi-temporal image sequence; performing material-aware crack detection on each preprocessed temporal image based on the material partition map, and combining the detection thresholds determined for each known region obtained from the partitioning of each preprocessed temporal image to obtain a set of crack detection boxes for each preprocessed temporal image, each crack detection box having a crack detection confidence level; and performing material-aware crack detection on the preprocessed temporal images based on multi-modal fusion prompt features. The current-phase cracks within the current-phase crack detection bounding box of the previous-phase image are segmented at the pixel level to generate a current-phase crack instance mask, thus obtaining the current-phase crack detection result on the building surface. The multimodal fusion prompt feature incorporates at least the features of the current-phase crack detection bounding box and the material classification information of the corresponding region of the current-phase crack detection bounding box, determined based on the material partitioning map. If the current-phase crack detection bounding box is associated with historical-phase crack instances, the multimodal fusion prompt feature also incorporates the morphological features of the corresponding historical-phase crack instance mask. Historical-phase crack instances have cumulative confidence scores, and the historical-phase crack instance mask and the corresponding historical-phase crack instance cumulative confidence scores are stored in a historical instance database. This database stores cross-phase detection information for all crack instances and is updated based on the current-phase crack detection result for subsequent phase detection threshold adjustment and pixel-level segmentation prompt generation.

[0042] Figure 1 An exemplary system architecture for applying a temporal detection method for building cracks according to an embodiment of the present invention is shown. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to embodiments of the present invention, in order to help those skilled in the art understand the technical content of the present invention, but do not mean that embodiments of the present invention cannot be used in other devices, systems, environments or scenarios.

[0043] like Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0044] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).

[0045] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0046] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0047] It should be noted that the time-series building crack detection method provided in this embodiment of the invention can generally be executed by server 105. Correspondingly, the time-series building crack detection device provided in this embodiment of the invention can generally be located in server 105. The time-series building crack detection method provided in this embodiment of the invention can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the time-series building crack detection device provided in this embodiment of the invention can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Alternatively, the time-series building crack detection method provided in this embodiment of the invention can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the building crack timing detection device provided in the embodiments of the present invention can also be installed in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.

[0048] For example, the multi-temporal image sequence of the building surface can be originally stored in any one of the first terminal device 101, the second terminal device 102, or the third terminal device 103 (e.g., the first terminal device 101, but not limited thereto), or it can be stored on an external storage device and imported into the first terminal device 101. Then, the first terminal device 101 can locally execute the building crack temporal detection method provided in the embodiments of the present invention, or send the multi-temporal image sequence to other terminal devices, servers, or server clusters, and have the other terminal devices, servers, or server clusters that receive the multi-temporal image sequence execute the building crack temporal detection method provided in the embodiments of the present invention.

[0049] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0050] Figure 2A A main flowchart of a time-series detection method for building cracks according to an embodiment of the present invention is shown.

[0051] like Figure 2A As shown, the method includes operations S201 to S206.

[0052] In operation S201, a multi-temporal image sequence of the building surface is acquired. The multi-temporal image sequence includes a reference temporal image and non-reference temporal images.

[0053] In operation S202, based on the spatial coordinates of the reference time phase image, spatial registration is performed on the non-reference time phase image to obtain a registered multi-time phase image sequence.

[0054] In operation S203, the reference time-phase image is classified pixel by pixel based on semantic segmentation to generate a material partition map.

[0055] In operation S204, based on the material partition map, differential texture suppression preprocessing matching the texture interference type of the material region is performed on different material regions in each registered phase image, and the preprocessed multi-phase image sequence is output.

[0056] In operation S205, based on the material partition map, material-aware crack detection is performed on the preprocessed temporal images. Combined with the detection threshold determined for each known region obtained from the division of the preprocessed temporal images, a set of crack detection boxes for each preprocessed temporal image is obtained, and each crack detection box has a crack detection confidence level.

[0057] In operation S206, based on the multimodal fusion prompt features, the current-phase cracks within the current-phase crack detection box of the preprocessed current-phase image are segmented at the pixel level to generate a current-phase crack instance mask, thus obtaining the current-phase crack detection result of the building surface. The multimodal fusion prompt features fuse at least the features of the current-phase crack detection box and the material classification information of the corresponding region of the current-phase crack detection box determined based on the material partitioning map. If the current-phase crack detection box is associated with historical-phase crack instances, the multimodal fusion prompt features also fuse the morphological features of the corresponding historical-phase crack instance mask. Historical-phase crack instances have cumulative confidence scores, and the historical-phase crack instance mask and the corresponding historical-phase crack instance cumulative confidence scores are stored in a historical instance library. This historical instance library stores cross-temporal detection information for all crack instances and is updated based on the current-phase crack detection result for subsequent temporal threshold adjustment and pixel-level segmentation prompt generation.

[0058] According to embodiments of the present invention, the above-described method for detecting temporal building cracks can be implemented based on material adaptive preprocessing and feedback from a historical instance library. The historical instance library is mainly used to provide masks of historical time-phase crack instances and the cumulative confidence levels of the historical time-phase crack instances they represent.

[0059] Figure 2B A general flowchart of a time-series building crack detection method based on material adaptive preprocessing and historical instance library feedback according to an embodiment of the present invention is shown.

[0060] like Figure 2B As shown, the method includes steps S1 to S7.

[0061] S1. Multi-temporal image registration: Aligning a multi-temporal image sequence to a unified coordinate system of a reference temporal image, providing a spatial basis for cross-temporal comparison. Specifically, this corresponds to operations S201 to S202 above.

[0062] For example, the input of multiple time-phase images is time-phase images T0, T1, T2, ..., Tn. T0 is determined as the reference time-phase image, and T1, T2, ..., Tn are non-reference time-phase images. During spatial registration, each time-phase image can be aligned to the coordinate system of time phase T0.

[0063] S2. Material Recognition and Partitioning: Based on the baseline temporal image, a semantic segmentation model is used to classify materials pixel by pixel and generate a material partitioning map (generated once and reused throughout the entire sequence).

[0064] For example, a material partition map can be generated based on the T0 phase image. Material feature modulation is performed based on this material partition map to drive the material-aware crack detection in step S4. The material context cues determined based on this material partition map drive the multimodal enhanced cue segmentation in step S5.

[0065] S3. Material-Adaptive Texture Suppression: Based on the material partitioning map, perform differentiated texture suppression on different material regions. Specifically, this corresponds to operation S204 above.

[0066] S4. Material-aware crack detection: Corresponding to the above operation S205, based on the target detection model, standard threshold full-image detection can be used for the reference time phase, and adaptive threshold detection can be performed for non-reference time phases based on the cumulative confidence of the historical instance library.

[0067] For example, the detection threshold can be adaptively adjusted based on the cumulative confidence level.

[0068] S5. Multimodal Enhanced Segmentation: Corresponding to the above operation S206, based on the instance segmentation model, standard bounding boxes and material context can be used for segmentation of the reference time phase; multimodal enhanced segmentation can be performed on the associated crack fusion detection boxes, historical segmentation masks and material context in non-reference time phases; and detection boxes and material context can be used for segmentation of newly discovered cracks.

[0069] For example, multimodal enhanced cue segmentation can be achieved through cross-modal cue fusion.

[0070] S6. Confidence Management: Based on the correlation matching results of the current time phase, the cumulative confidence of each crack instance in the historical instance library is updated by incrementing or decreasing, and false detection elimination is performed.

[0071] S7. Historical Instance Library Update: The crack detection bounding box, crack instance mask, material classification information, association status, and updated cumulative confidence score of the current time phase are written into the same historical instance library. The historical instance library also serves as the basis for adjusting the detection threshold and the source of segmentation morphology cues for the next time phase, providing historical detection information for detection and segmentation in the next time phase. Specifically, the cumulative confidence score in the historical instance library is mapped to a spatially known region mask in the current time phase image. By determining the detection threshold for each puzzle region, the detection threshold in step S4 is adaptively adjusted. Historical time phase crack instances in the historical instance library are encoded as morphology cues, and the historical masks drive the enhanced segmentation cues in step S5. The crack detection bounding box, crack instance mask, and updated cumulative confidence score of the current time phase are updated in reverse to the same historical instance library, forming a closed-loop temporal detection mechanism with dual-path, bidirectional feedback.

[0072] Corresponding to, for example Figure 2B The overall process shown includes the following inputs: a sequence of multi-temporal images of the ancient building surface, such as a baseline temporal image T0 and subsequent non-baseline temporal images T1, T2, ..., Tn. Outputs may include: a set of crack detection boxes for each temporal phase, an instance-level pixel segmentation mask for each crack, the cumulative confidence score for each crack instance, preprocessed images of each temporal phase (texture interference has been specifically suppressed), a material partitioning map, and a historical instance library (containing cross-temporal records of crack instances under management), etc.

[0073] Table 1 shows the data flow transmission relationship between steps S1 to S7.

[0074] Table 1:

[0075]

[0076] The core closed loop: The historical instance library plays two roles in each phase of processing: guiding the adjustment of the detection threshold in S4 with cumulative confidence, and guiding the generation of segmentation hints in S5 with historical segmentation masks. After detection and segmentation are completed, the new results are updated to the historical instance library for continued use in the next phase, forming a phase-by-phase feedback optimization. In addition, the material partitioning map is independent of the above closed loop and is reused in the three stages of texture suppression, detection enhancement, and segmentation hints.

[0077] Through the above embodiments of the present invention, the problems of false crack detection due to complex textures on building surfaces and poor multi-temporal detection accuracy are addressed. First, multi-temporal image registration is completed. Differential texture suppression is implemented through pixel-level material partitioning to reduce false positive interference from complex textures highly similar to cracks at the source. Then, material-aware crack detection is adaptively performed based on material partitioning, and multi-modal features such as detection boxes, material classification information, and historical crack masks are fused for pixel-level segmentation. This accurately distinguishes inherent textures from real cracks, reducing false positives and false negatives. Simultaneously, tracking crack evolution based on cumulative confidence effectively improves the detection accuracy, segmentation fineness, and detection reliability of multi-temporal cracks in ancient buildings.

[0078] The following describes specific embodiments. Figure 2B The method shown will be further explained.

[0079] S1. Multi-temporal image registration.

[0080] S1.1, Input.

[0081] The multi-temporal image sequence of the ancient building surface includes a reference temporal image T0 and subsequent non-reference temporal images T1, T2, ..., Tn.

[0082] S1.2 Registration method.

[0083] Using the reference time phase image T0 as a spatial reference, the subsequent non-reference time phase images T1, T2, ..., Tn are spatially registered so that each time phase image is aligned to the unified coordinate system where the T0 image is located.

[0084] Spatial registration methods can be any of the following or a combination thereof.

[0085] Registration methods based on feature point matching: extract feature points between images, such as Scale-Invariant Feature Transform (SIFT), Speeded Up Robust Features (SURF), and Oriented Fast and Rotated BRIEF (ORB), and estimate spatial transformation parameters through feature point matching.

[0086] Mutual information-based registration method: using mutual information between images as a similarity measure, and maximizing mutual information by optimizing transformation parameters.

[0087] Deep learning-based registration method: Uses deep neural networks to learn the spatial transformation relationship between images.

[0088] Orthophoto-based registration method: When there is a 3D model or orthophoto data, registration is performed using geometric constraints.

[0089] According to embodiments of the present invention, the spatial registration method is not limited to the above implementation.

[0090] S1.3, Processing Rules.

[0091] For the reference time phase image T0, no spatial registration operation is required; it can be directly used as a spatial reference.

[0092] For each subsequent non-reference time phase image T1~Tn, register them one by one and transform them into the coordinate system of T0.

[0093] The registration accuracy should meet the crack-level pixel alignment requirements to support subsequent cross-temporal crack morphology comparisons.

[0094] S1.4, Output.

[0095] After registration, the multi-temporal image sequence is spatially aligned in a unified coordinate system of T0.

[0096] S2, Material Identification and Zoning.

[0097] S2.1, Purpose of the steps.

[0098] Since the surface of ancient buildings usually includes a variety of materials such as wood structure, brick and stone, stucco / plaster / painting, etc., different materials have different texture structures and interference characteristics, the first step is to identify and partition the surface of ancient buildings to support the subsequent material adaptive preprocessing and material perception segmentation.

[0099] S2.2, Detailed Implementation Method.

[0100] Using the baseline time-phase image T0 as input, a semantic segmentation model is employed to perform pixel-by-pixel material classification, generating a material partition map. The semantic segmentation model can employ DeepLabV3+, U-Net, a semantic segmentation model based on a self-attention transformer (SegFormer), or other encoder-decoder structures. Material categories must include at least three types: wooden surfaces, masonry surfaces, and plaster / painted / decorated surfaces. These three material categories cover the most common crack detection scenarios on ancient building surfaces. For special materials such as adobe walls, stone carvings, and metal components, corresponding material categories can be extended into the semantic segmentation model; the material adaptive preprocessing and material-aware detection framework of this invention are also applicable.

[0101] During the model training phase, images of ancient building surfaces were collected to construct a material recognition dataset. This dataset covers different material categories and their common appearance variations. Wooden structure samples include logs, painted wooden structures, and decayed / discolored wooden structures; brick and stone samples include blue bricks, flagstones, rubble, and mortar joints; and plaster / painted samples include whitewashed surfaces, plaster decorations, murals, and weathered gray underlayers. Pixel-level material annotations were applied to these samples, and difficult examples with similar material appearances or unclear boundaries were added. During model training, transfer learning was used to fine-tune the pre-trained weights, and higher weights were assigned to pixels in material boundary regions and easily confused areas in the loss function to improve the model's ability to recognize material boundaries and complex textures.

[0102] In the post-processing stage, after obtaining the initial material label map, isolated noise regions with an area smaller than a preset threshold are removed by connected component analysis, and these isolated noise regions are merged into the material category with the dominant area in the neighborhood to eliminate isolated misclassified pixels. For pixel regions with similar class probability values ​​on both sides at the material boundary, they are marked as transition zones, and the width of the transition zone is adaptively determined according to the class probability gradient. Pixels in the transition zone are finally assigned to the material category with the dominant area or the highest probability response in the neighborhood to smooth the partition boundary.

[0103] S2.3, Output.

[0104] Material partition diagram.

[0105] Each pixel in the material partitioning map corresponds to a material category. Since the material partitioning map is generated based on the reference temporal image T0, and subsequent temporal images have been registered and aligned to the T0 coordinate system via S1, and in the crack detection scenario addressed by this invention, crack development typically does not change the main material category of the surface region, the material partitioning map can be generated once and reused in subsequent temporal images. The material partitioning map is reused in three steps: texture suppression in S3, material feature modulation in S4, and material context hints in S5, achieving single partitioning and full-sequence reuse.

[0106] S3, Material Adaptive Texture Suppression.

[0107] S3.1, Purpose of the steps.

[0108] Based on the material partition map obtained in S2, differential texture suppression is performed on different material regions in each time phase image to reduce the interference of the inherent texture of ancient buildings on crack detection. S2 has generated a material partition map based on the T0 reference time phase. Since the images of each time phase are registered and aligned by S1 and the main material category remains stable within the detection period (see S2.3), the material partition map can be directly used for texture suppression in each time phase without repeating material recognition.

[0109] S3.2, Detailed Implementation Method.

[0110] (1) Treatment of wooden structure areas.

[0111] For the wood-structure region, the processing flow consists of three sub-steps: grayscale normalization, wood grain orientation field estimation, and directional suppression. The core implementation method is as follows: The wood grain orientation field is estimated for the grayscale-normalized wood-structure region, and the main extension direction of the wood grain is extracted. Based on the main extension direction, an adaptive Gaussian kernel or a directional mask is constructed at each pixel location in the normalized wood-structure region. The adaptive Gaussian kernel represents a Gaussian kernel with a width parallel to the main extension direction that is greater than the width perpendicular to the main extension direction. This is described in detail below.

[0112] Sub-step a1: Gray-level normalization. First, gray-level normalization is performed on the wood structure areas to eliminate gray-level unevenness caused by differences in lighting conditions or surface conditions between different wood structure areas. Specifically, mean-variance normalization can be performed on the gray-level values ​​within each local window: subtract the mean from the pixel gray-level value within the window and divide by the standard deviation, so that the normalized gray-level distribution has zero mean and unit variance. The window size can be adaptively set according to the wood grain period.

[0113] Sub-step b1: Wood grain orientation field estimation. Estimate the wood grain orientation field of the grayscale-normalized image to obtain the main extension direction of the wood grain at each pixel location. The specific process of wood grain orientation field estimation is as follows.

[0114] 1. Construct a set of Gabor filters in different directions (e.g., take θ∈0°,15°,30°,...,165° for a total of 12 directions), and calculate the amplitude response of the Gabor filter in each direction for each pixel position.

[0115] 2. For each pixel location, the direction corresponding to the largest amplitude response is taken as the main wood grain direction at that location, forming a pixel-by-pixel orientation field map θ(x,y).

[0116] 3. Smooth the orientation field map (e.g., take the angle after vector averaging of the orientation field) to eliminate orientation estimation jitter caused by local noise and obtain a continuous and smooth wood grain orientation field.

[0117] In addition, the wood grain orientation field estimation can also be performed using a method based on the local structure tensor: calculate the covariance matrix (structure tensor) of the image gradient in the neighborhood of each pixel, and take the direction of its principal feature vector as the wood grain orientation.

[0118] Sub-step c: Directional Texture Suppression and Crack Enhancement. After obtaining the wood grain orientation field, an anisotropic filter is constructed based on the orientation field map θ(x,y). A strong smoothing response is applied along the main wood grain extension direction to suppress regular wood grain textures extending along the wood grain direction, while the original response is preserved in the direction perpendicular to the wood grain to enhance thin, elongated dark line structures that are inconsistent with the wood grain direction. Specifically, an adaptive Gaussian kernel or directional mask is constructed at each pixel position: the kernel width along the θ(x,y) direction is larger, such as a Gaussian kernel standard deviation σ_along = 5-10 pixels along the main wood grain extension direction, used to control the smoothing scale along the main wood grain extension direction; the kernel width perpendicular to the θ(x,y) direction is smaller, such as a Gaussian kernel standard deviation σ_across = 1-3 pixels perpendicular to the main wood grain extension direction, used to control the smoothing scale perpendicular to the main wood grain extension direction. Through this processing, regular strip textures consistent with the main wood grain extension direction are smoothed and suppressed, while crack edge structures with directions inconsistent with the wood grain are preserved or enhanced.

[0119] (2) Treatment of masonry areas.

[0120] For masonry areas, the processing flow is divided into two sub-steps: linear structure extraction and regularity identification and suppression. The core implementation method is as follows: extracting the linear structure features of the masonry area; performing cluster analysis on the linear structure features according to direction to identify at least one set of parallel line segments and the dominant direction of each set of parallel line segments; for each set of parallel line segments, projecting along the normal direction of the dominant direction of the parallel line segment, calculating the spacing sequence between adjacent line segments and the spacing variation coefficient of the spacing sequence; if the spacing variation coefficient is less than a preset spacing variation threshold and the number of parallel line segments is greater than or equal to a preset minimum value, the parallel line segments are identified as regular line structures; generating a suppression mask for the area where the target line segments identified as regular line structures are located; setting the gradient response value of the pixel grayscale value within the suppression mask to be lower than the gradient response value of the pixel grayscale value outside the suppression mask, and / or setting the contrast of the pixel grayscale value outside the suppression mask to be higher than the contrast of the pixel grayscale value within the suppression mask. A detailed description follows.

[0121] Sub-step a2: Linear structure extraction. First, the linear structure features in the masonry surface are extracted using the Line Segment Detector (LSD) algorithm or Hough transform to obtain the position, direction, and length parameters of each line segment. Alternatively, edge detection (such as the Canny operator) can be used, followed by morphological closing operations to connect broken edges, and then connected line segments can be extracted.

[0122] Sub-step b2: Regularity identification and suppression. Regularity analysis is performed on the extracted line segment set to identify regular line structures with periodic arrangement characteristics, such as brick joints / mortar joints. The specific process is as follows.

[0123] 1. Orientation Clustering: Perform orientation clustering analysis on all extracted line segments (e.g., using K-means or density clustering) to identify at least one group of approximately parallel line segments with the dominant orientation. Each group of approximately parallel line segments may include multiple groups of adjacent line segments.

[0124] 2. Spacing Analysis: For a set of approximately parallel line segments along the dominant direction, project along their normal direction and calculate the spacing sequence between adjacent line segments. And calculate the coefficient of variation of each spacing. ,in, This represents the sequence of spacing between the kth adjacent line segments. Represents the spacing sequence standard deviation Represents the spacing sequence The average value.

[0125] 3. Regularity Determination: When the coefficient of variation (CV) of the spacing is less than the preset spacing variation threshold (e.g., 0.2) and the number of segments in the group of approximately parallel line segments is not less than the preset minimum value (e.g., 3), the group of approximately parallel line segments is determined to be a regular line structure (i.e., brick joints or mortar joints), and its main extension direction is marked. and average spacing .

[0126] 4. Suppression Processing: For regions marked as regular line structures, a suppression mask is generated along their normal direction. The width of the suppression mask can be equal to or slightly larger than the width of the line segment itself. Within the suppression mask, the contrast of pixel grayscale values ​​or the gradient response intensity is reduced, significantly weakening the response of regular brick seams in subsequent edge and crack detection.

[0127] Simultaneously, it enhances irregular crack structures that are inconsistent with the direction, spacing, or shape of brick or mortar joints. Specifically, it strengthens crack structures whose direction is inconsistent with the edge detection results. Edge segments with differences exceeding a preset angle threshold (such as 30°) retain their original response intensity, making them more prominent in subsequent crack detection.

[0128] (3) Treatment of gray plastic / plaster / painted areas.

[0129] For decorative areas where at least one of the following is determined—plaster, plaster, or painted—the processing flow is divided into two sub-steps: non-crack edge filtering and local contrast enhancement. The core implementation method is as follows: edge detection is performed on the decorative area to obtain edge segments; if the edge segments meet the non-crack edge filtering conditions, they are identified as non-crack edges. The non-crack edge filtering conditions include at least one of the following: the mean curvature of the edge segment is less than a preset mean curvature threshold and the variance of curvature is less than a preset variance of curvature threshold; the width of the edge segment is greater than a preset width threshold; the grayscale distribution characteristics of the area where the edge segment is located are symmetrical with the grayscale distribution characteristics of the neighborhood of the edge segment along the normal direction of the edge segment; the neighborhood of the edge segment represents the building surface area adjacent to the edge segment along the normal direction of the edge segment; the edge detection response value of the pixel grayscale value within the non-crack edge area is set lower than the edge detection response value of the pixel grayscale value outside the non-crack edge area, and / or the contrast of the pixel grayscale value outside the non-crack edge area is set higher than the contrast of the pixel grayscale value within the non-crack edge area. A detailed description follows.

[0130] Sub-step a: Non-crack edge filtering. First, perform morphological analysis on the edge detection results (i.e., the detected edge segments) to identify and filter non-crack edge interference such as weathered and peeling edges and painted pattern boundaries. The judgment criteria are as follows.

[0131] 1. Curvature characteristics: Weathered and flaked edges are usually smooth arcs with slow and continuous curvature changes; crack edges have larger curvature changes and often show sharp turns. Calculate the point-by-point curvature of each edge segment. When the mean curvature is lower than a preset mean curvature threshold and the variance of curvature is less than a preset variance of curvature threshold, it is determined to be a weathered and flaked edge.

[0132] 2. Transition Zone Width: Weathered and flaked edges exhibit significant grayscale differences on both sides, but the transition zone is relatively wide (the gradual transition area between the flaked and unflaked surfaces); crack edges show abrupt grayscale changes on both sides, and the transition zone is extremely narrow. Grayscale profiles are sampled along the normal direction of the edge segments to calculate the transition zone width, i.e., the edge segment width, representing the number of pixels spanned by the grayscale change from low to high values. When the transition zone width exceeds a preset width threshold (e.g., 5 pixels), it is determined to be a weathered and flaked edge rather than a crack edge.

[0133] 3. Symmetry of grayscale distribution: The grayscale values ​​on both sides of a crack are usually asymmetrical (the grayscale value inside the crack is lower than that on the two sides of the surface), while the grayscale changes at the boundary of a painted pattern may be symmetrical. The grayscale distribution characteristics on both sides of the edge segment (i.e., the area where the edge segment is located and the neighborhood of the edge segment) are statistically analyzed along the normal direction of the edge segment to help distinguish the crack edge from the pattern boundary.

[0134] Based on the above criteria, structures determined to be non-crack edges are suppressed or weakened (by reducing their edge detection response intensity).

[0135] Sub-step b: Local contrast enhancement. After filtering out non-crack edge interference based on the above non-crack edge filtering conditions, local pixel grayscale contrast enhancement is performed to improve the detectability of low-contrast cracks or fine cracks. Specifically, the Contrast Limited Adaptive Histogram Equalization (CLAHE) method is used: the image is divided into overlapping local sub-blocks, histogram equalization is performed within each sub-block, and the histogram is cropped using a contrast limiting parameter to avoid excessive noise amplification. The equalization results of each sub-block are then stitched together using bilinear interpolation. Alternatively, adaptive gamma correction or local standard deviation normalization can also be used for contrast enhancement.

[0136] S3.3, Output.

[0137] The preprocessed images at different time points show that texture interference in different material regions has been specifically suppressed.

[0138] S3.4 Data connection with subsequent steps.

[0139] The preprocessed temporal image output by S3 can be directly input into S4 for crack detection. Since S3 has specifically suppressed texture interference in different material regions (wood grain is smoothed, brick joints are periodically suppressed, and weathered edges of gray plaster are filtered), S4's material-aware detection model can reduce interference and improve detection stability on the preprocessed temporal image. In other words, S3 and S4 form a dual anti-interference mechanism of "first suppressing interference at the pixel level, then perceiving material at the feature level."

[0140] S4. Material-sensing crack detection.

[0141] S4.1 Overview.

[0142] A target detection model is used to detect cracks in the registered and preprocessed (S3) images of each time phase to obtain crack candidate regions. Different detection threshold strategies are adopted depending on whether the current processing time phase is the reference time phase: the reference time phase uses the standard detection threshold for full-image detection; for non-reference time phases, an adaptive detection threshold mechanism based on the cumulative confidence of the historical instance library is introduced, and different detection sensitivities are applied to different known regions.

[0143] S4.2 Material-aware crack detection model.

[0144] This invention designs a material-aware crack detection model. Based on a general object detection model, this model introduces a material feature modulation mechanism, enabling it to adaptively adjust feature extraction and detection strategies according to the material type of the image region. Its core implementation method is as follows: Multi-scale feature extraction is performed on preprocessed images at various time points to obtain detection feature maps at multiple scales; feature encoding is performed on the material partition map to obtain material feature maps at multiple scales, wherein the spatial size of each scale material feature map matches the spatial size of the corresponding scale detection feature map; based on the channel-wise scaling and offset coefficients generated from the material feature maps, the detection feature maps and material feature maps at multiple scales are fused to obtain a material modulation feature map; crack detection is performed on the material modulation feature map to obtain crack candidate detection boxes and crack detection confidence scores for the crack candidate detection boxes. A detailed description follows.

[0145] S4.2.1 Overall Model Architecture.

[0146] Figure 3 An architecture diagram of a material-aware crack detection model according to an embodiment of the present invention is shown.

[0147] like Figure 3 As shown, the material-aware crack detection model consists of three core modules: backbone network 310, material feature encoding branch 320, and material modulation feature fusion module 330.

[0148] (1) Image feature extraction backbone network.

[0149] A general object detection backbone network 310 is used, such as the Cross Stage Partial Darknet (CSP Darknet, where Darknet is a native backbone network name), Residual Network (ResNet), and Swin Transformer, to perform multi-scale feature extraction on the preprocessed temporal image 311 output by S3, outputting detection feature maps 312 at multiple scales: .

[0150] (2) Material feature coding branch.

[0151] Using the material partition map 321 output by S2 as input, material features are extracted through a lightweight material feature encoding branch 320. Specifically, the material partition map 321 is a single-channel image, and each pixel value is a material category label (e.g., 0=background, 1=wood, 2=brick, 3=gray plastic). The structure of the material feature encoding branch 320 is as follows.

[0152] 1. Perform one-hot encoding on the material partition map 321 to obtain the material category probability map of channel C (C is the number of material categories).

[0153] 2. A three-layer convolutional network (hereinafter referred to as Conv-BN-ReLU) based on convolutional layers (Conv), batch normalization layers (BN), and rectified linear activation functions (ReLU) is used to downsample and extract material space features layer by layer, outputting material feature maps that match the spatial dimensions of feature maps at each scale of the backbone network. .

[0154] 3. The number of parameters in the material feature coding branch 320 is much smaller than that in the backbone network 310, and it is only used to encode the spatial distribution information of materials.

[0155] (3) Material modulation feature fusion module.

[0156] Material information is injected into image features through Feature-wise Linear Modulation (FiLM) mechanism on feature maps at various scales, as shown in formula (1).

[0157] (1).

[0158] in, This represents the detection feature map at the i-th scale of the backbone network. This is the material feature map corresponding to the i-th scale. The scaling factor is generated from the material feature map per channel. The per-channel offset coefficients are generated from the material feature map and passed through two independent 1×1 convolutional layers. from The calculation yielded: , ⊙ represents element-wise multiplication, used to perform element-wise affine transformation on the detection feature map 312 of multiple scales obtained by the backbone network 310 based on the channel-wise scaling and offset coefficients generated by the material feature map 322. This allows the detection features to adaptively adjust their response mode according to the material category, so as to output the modulated feature map 331. .

[0159] The function of the material modulation mechanism is to differentiate the enhancement or suppression of feature responses in different channels based on the material type. For example, features related to directional stripe textures in the corresponding channels of a wood structure region are suppressed, while features related to crack morphology are enhanced; features related to periodic line structures in a masonry region are suppressed. This modulation is automatically learned at the feature level, eliminating the need for manual design of filter parameters.

[0160] Modulated feature map 331 The crack detection head 340 is fed in to perform crack detection, and the crack candidate detection box + crack detection confidence 341 is output.

[0161] S4.2.2 Material-weighted training strategy.

[0162] During the model training phase, a material-aware weighted loss function is used, as shown in Equation (2).

[0163] (2).

[0164] Where M is the set of material categories, The detection loss is the sample loss within the area covered by material category m. This represents the loss weight for material category m. This is a weighted loss value using material perception.

[0165] The principles for setting weights are as follows.

[0166] Areas with strong texture interference (such as wood and brick) are given higher weights, forcing the model to learn to distinguish between cracks and crack-like textures in difficult material areas.

[0167] Material areas with weak texture interference (such as simple gray plastic surfaces) are assigned standard weights.

[0168] The background area is given a lower weight.

[0169] Meanwhile, the training data adopts a material-balanced sampling strategy: the number of samples in each material region is kept balanced in each training batch to avoid the model being biased towards the material type with the most samples.

[0170] S4.2.3, Model Advantages.

[0171] Compared with traditional general detection models, the material-aware crack detection model has the following advantages.

[0172] Through the material feature encoding branch, the model gains the ability to perceive the material type of an image region, and can distinguish between cracks and material-specific textures at the feature level.

[0173] Through the FiLM modulation mechanism, the model automatically adjusts the feature response according to the material type, eliminating the need to manually design independent filter parameters for different materials.

[0174] Through material-weighted training, the model has stronger anti-interference ability in areas with severe texture interference (wooden structures, brick and stone).

[0175] S4.3, Reference phase detection strategy (for T0).

[0176] For the baseline time-phase image T0, full-image crack detection is performed using a standard detection threshold. The core implementation method is as follows: First, candidate crack detection boxes and their first crack detection confidence scores are obtained from the preprocessed baseline time-phase image. Based on the preset standard detection threshold, target candidate crack detection boxes with target first crack detection confidence scores are determined, resulting in a set of crack detection boxes for the preprocessed baseline time-phase image, where the target first crack detection confidence score is greater than or equal to the standard detection threshold. The specific steps are as follows.

[0177] 1. Input the T0 image and corresponding material partition map after S3 preprocessing into the material-aware crack detection model.

[0178] 2. The model outputs a set of candidate detection boxes for the first crack and the corresponding confidence scores for the first crack detection.

[0179] 3. Filter the first crack candidate detection boxes using a preset standard detection threshold Th_std (such as 0.25 or 0.5), and retain the target first crack candidate detection boxes with a confidence level ≥ Th_std.

[0180] 4. Optionally, non-maximum suppression (NMS) can be performed to remove overlapping detection boxes.

[0181] 5. Output the set of crack detection frames for T0.

[0182] S4.4 Non-reference time phase adaptive detection strategy (for T1~Tn).

[0183] For non-reference time-phase images T1, T2, ..., Tn, an adaptive detection threshold mechanism based on the cumulative confidence level of a historical instance library is introduced. The core idea is to divide the image space into known regions with different confidence levels according to the cumulative confidence level of known crack instances in the historical instance library, and to apply different detection threshold strategies to different regions.

[0184] It is important to note that the adaptive detection threshold mechanism of this invention is fundamentally different from the correlation threshold adjustment in the field of Multi-Object Tracking (MOT). Threshold adjustment in MOT is a scalar operation per trajectory and frame, meaning that a scalar confidence threshold is adjusted independently for each motion trajectory, without involving regional division in the image space. The core feature of this invention is that it maps the accumulated confidence to a known region mask in the image space, allowing different known regions in the same temporal image to use different detection thresholds in parallel. This spatial partitioning and parallel multi-threshold design is driven by the unique requirement of multi-temporal re-inspection of cracks in ancient buildings at fixed locations. There is no corresponding technical solution in the field of multi-object tracking.

[0185] Figure 4 A schematic diagram of an adaptive detection threshold mechanism according to an embodiment of the present invention is shown.

[0186] like Figure 4 As shown, the adaptive detection threshold mechanism involves a detailed process of operations such as known region division and regional adaptive detection threshold setting.

[0187] S4.4.1 Known region division.

[0188] The crack detection bounding box positions and cumulative confidence scores of all detected crack instances are read from the historical instance library 410. Known region segmentation 420 is then performed, dividing the image space into three types of known regions 430. The core implementation method is as follows: based on the crack detection bounding boxes and cumulative confidence scores of all crack instances, the preprocessed non-reference time-phase image is divided into a first-confidence known region, a second-confidence known region, and a third-confidence known region. Specifically, the first-confidence known region is determined based on the crack detection bounding boxes of crack instances whose cumulative confidence scores are greater than or equal to a first confidence threshold. The second-confidence known region is determined based on the crack detection bounding boxes of crack instances whose cumulative confidence scores are less than the first confidence threshold but greater than or equal to the second confidence threshold. The third-confidence known region is determined based on the crack detection bounding boxes of crack instances whose cumulative confidence scores are less than the second confidence threshold. The third-confidence known region is also determined based on other regions not covered by the historical time-phase crack detection bounding boxes read from the historical instance library. The first confidence threshold is greater than the second confidence threshold. A detailed description of the three types of known regions 430 is as follows.

[0189] High-confidence known region (i.e., first-confidence known region): The image area (and appropriate spatial extension range) covered by the crack detection bounding box of crack instances in the historical instance database where the cumulative confidence score is ≥ the high-confidence threshold Th_high. The spatial extension range is calculated as follows: based on the shorter side length of the crack detection bounding box, extend outward by 10%~30% of this reference as a buffer zone. The purpose of the extension is to cover areas where the crack may extend slightly beyond the detection bounding box. Th_high, as the aforementioned first-confidence threshold, can take values ​​such as 0.7 or 0.8. These areas correspond to real cracks that have been confirmed multiple times by historical observations, requiring higher sensitivity to detect minor changes in the crack.

[0190] Medium-confidence known regions (i.e., second-confidence known regions): Image areas covered by crack detection boxes of crack instances in the historical instance database where the cumulative confidence satisfies: low confidence threshold Th_low ≤ cumulative confidence C_cum < high confidence threshold Th_high. Th_low, as the second confidence threshold, can take values ​​such as 0.3 or 0.4. Cracks in these regions require further confirmation; the standard detection threshold should be maintained.

[0191] Low-confidence / unknown regions (i.e., regions with known third-confidence): Image regions not covered by historical instances in the historical instance database, or regions where the cumulative confidence of historical instances covering the region is less than the low-confidence threshold Th_low. The former are entirely new exploration regions where cracks have never appeared, while the latter may be regions covered by crack instances that are being eliminated due to historical false detections. Neither type of region has reliable historical crack detection information; therefore, they are combined and a standard detection threshold is used for full-image detection to avoid introducing too many false detections by lowering the threshold.

[0192] S4.4.2, Regional adaptive detection threshold setting.

[0193] See also Figure 4 The design concept section 440 sets detection thresholds for three types of known regions. The core idea is that the region detection threshold for a known region with the first confidence level is lower than the region detection threshold for a known region with the second confidence level. The region detection threshold for a known region with the second confidence level is equal to the region detection threshold for a known region with the third confidence level. Specifically, as shown in Table 2.

[0194] Table 2:

[0195]

[0196] Among them, delta_high is the threshold reduction amount, which represents the threshold reduction magnitude in the high confidence region. It is a preset hyperparameter, and its value range is generally [0.05, 0.20].

[0197] S4.4.3 Adaptive Detection Execution Process.

[0198] According to an embodiment of the present invention, the core implementation method of this adaptive detection execution process is as follows: obtaining a second crack candidate detection box and a second crack detection confidence level of the second crack candidate detection box obtained from the detection of the preprocessed non-reference temporal image; determining a target second crack candidate detection box with a target second crack detection confidence level based on an adaptive detection threshold of the known region to which the second crack candidate detection box belongs, thereby obtaining a set of crack detection boxes for the preprocessed non-reference temporal image. The known region to which the second crack candidate detection box belongs is determined based on the known region where the center of the second crack candidate detection box is located or the known region covered by the second crack candidate detection box. The adaptive detection threshold is determined based on at least one region detection threshold corresponding to the known region to which the second crack candidate detection box belongs. The target second crack detection confidence level is greater than or equal to the adaptive detection threshold. A detailed description follows.

[0199] The adaptive detection process for non-reference time phases is as follows.

[0200] 1. Retrieves the location information and cumulative confidence information of the crack detection boxes for all crack instances from the historical instance library.

[0201] 2. Based on the cumulative confidence level, historical crack instances are divided into three categories: high, medium, and low.

[0202] 3. Based on the location information of the crack detection boxes for various instances, generate corresponding three types of known region masks on the current temporal image.

[0203] 4. Input the preprocessed image and material partition map of the current time phase into the material-aware crack detection model to obtain all candidate detection boxes for the second crack and the confidence score of the second crack detection.

[0204] 5. For each candidate detection box of the second crack, determine which type of known region its center position falls within, and filter it using the corresponding adaptive detection threshold.

[0205] 6. For detection boxes that cross the boundaries of multiple known regions, the equivalent detection threshold is calculated by weighting the area of ​​each known region within the detection box according to its proportion. ,in, Let be the proportion of the known region of type j within the detection box. The threshold value for detecting the region corresponding to the j-th known region is [value]. This is the equivalent detection threshold. When weighted calculation is inconvenient to implement, it can be simplified to taking the area detection threshold corresponding to the highest confidence level in the known area covered by the crack detection frame as a conservative strategy.

[0206] 7. Optionally, nonmaximum suppression can be performed to remove overlapping detection boxes.

[0207] 8. Output the set of crack detection frames for the current time phase.

[0208] S4.5, Output.

[0209] The current time phase contains a set of candidate crack detection boxes. Each detection box contains the following information: spatial coordinates (position in the T0 unified coordinate system), crack detection confidence, and the known region type label (only for non-baseline time phases).

[0210] S5, Multimodal Enhanced Segmentation.

[0211] S5.1 Overview.

[0212] Using the crack detection bounding boxes output from step S4 as the basis for spatial localization, an instance segmentation model is employed to perform precise pixel-level segmentation of cracks within each detection box, generating crack instance masks. Different segmentation hint strategies are used depending on whether the current processing phase is a baseline phase: the baseline phase uses standard bounding boxes and material context hints for segmentation; in non-baseline phases, cracks associated with historical instances are segmented using a three-modal fusion hint system of detection bounding boxes, historical masks, and material context, while newly discovered cracks are segmented using both detection bounding boxes and material context hints.

[0213] S5.2, Segmentation Model.

[0214] The instance partitioning model is not limited to a specific architecture, and the models that can be used include, but are not limited to, those shown below.

[0215] The Segment Anything Model (SAM) series includes models such as SAM, SAM 2, SAM-HQ (Segment Anything in High Quality), and MobileSAM (Mobile Lightweight Segment Anything Model).

[0216] Other instance segmentation models that support cue-driven segmentation.

[0217] The segmentation model supports receiving prompts (such as boxes, dots, masks, etc.) to guide the segmentation process and outputs pixel-level segmentation masks for specified regions.

[0218] S5.3 Multimodal cue feature fusion.

[0219] According to an embodiment of the present invention, the core implementation method of multimodal enhanced segmentation is as follows: feature extraction is performed on the current temporal crack detection box to obtain the detection box prompt feature vector, wherein the detection box prompt feature vector includes the spatial feature vector and visual feature vector of the current temporal crack detection box; feature extraction is performed on the historical temporal crack instance mask to obtain the historical temporal crack morphology feature vector; feature extraction is performed on the material classification information to obtain the material semantic vector; multi-head cross-attention calculation is performed with the detection box prompt feature vector as the query and the historical temporal crack morphology feature vector and the material semantic vector as the key to obtain the multimodal fusion prompt feature.

[0220] S5.3.1, The composition of multimodal prompts.

[0221] Modal 1: Detection box hints. This refers to the crack detection boxes for the current time phase output from step S4, providing approximate spatial location of the cracks. Similar to the standard method, the detection box hints define the working range of the segmentation model.

[0222] Modality 2: Historical Mask Hint. The historical instance database retrieves the segmentation mask (i.e., the historical crack instance mask) of the historical crack instance associated with the current crack detection frame in a previous phase, serving as a morphological reference hint. The historical crack instance mask provides precise morphological information about the crack's location, direction, branching structure, and boundary shape in the historical phase.

[0223] The mechanism by which the history mask hint works is as follows.

[0224] This provides a reference for the expected shape of cracks in the segmentation model, enabling the model to prioritize regions consistent with historical shapes in the current phase.

[0225] It helps the segmentation model distinguish between real cracks and background interference (such as wood grain, brick joints, etc.) that are inconsistent with the shape.

[0226] When a crack expands slightly, the crack instance mask from historical time phases can be used as an anchoring reference to help the model identify the new portion.

[0227] When a crack instance in the historical instance library has mask records in multiple previous time phases, the selection strategy depends on the crack's rate of change: for cracks that change slowly, the historical time phase crack instance mask of the most recent historical time phase is selected as the historical mask hint because it is closest to the current shape; for cracks that expand rapidly or have large shape changes, the historical time phase crack instance masks of multiple historical time phases can be selected and fused (such as taking the union to cover the area that the crack may expand to) to provide a more lenient shape reference.

[0228] Modality 3: Material Context Hints. Based on the material partition map generated by S2, material classification information (such as wood, brick, stucco, etc.) of image regions is obtained, providing material-level reference information for the segmentation model.

[0229] The mechanism by which material context hints work is as follows.

[0230] This helps the segmentation model refer to the material type of the area when determining crack boundaries, and makes more accurate pixel assignment judgments in areas where material texture remains.

[0231] The segmentation strategy should be adjusted according to the crack characteristics of different materials. For example, cracks on wood surfaces usually develop along the grain and are darker in color, while cracks on brick and stone surfaces usually cross the mortar joints and have irregular edges.

[0232] The specific representation of material context information is as follows: the pixel proportion of each material category within the crack detection box area is statistically analyzed to form a material distribution vector, which is then mapped to a material semantic vector through a learnable embedding matrix. For specific implementation details, see Part (3) of S5.3.2.

[0233] The material context hints are directly derived from the material partition map generated by S2, enabling one partition to be reused in multiple places.

[0234] S5.3.2, Cross-modal prompting fusion network.

[0235] Figure 5 An architecture diagram of a cross-modal cue fusion network according to an embodiment of the present invention is shown.

[0236] like Figure 5 As shown, the three heterogeneous modal information—the crack detection box 501 based on S4 output, the historical mask 502 obtained from the historical instance database, and the material classification information 503 obtained from the material partitioning map based on S2—are encoded into a unified high-dimensional multimodal fusion cue feature 542. This feature is used to drive the segmentation model 550 to perform accurate crack segmentation, thereby obtaining the crack instance mask 551. The output dimension of each encoder below is uniformly d, with a typical value of d=256.

[0237] (1) Detection box spatial encoder 510 to realize two-layer perceptron + region feature alignment.

[0238] Crack detection frame 501 output by S4 The encoding is a detection box prompt feature vector 511, where, This represents the minimum x-coordinate of the crack detection box in the reference time-phase unified image coordinate system. This represents the minimum ordinate of the crack detection box in the reference time-phase unified image coordinate system. This represents the maximum x-coordinate of the crack detection box in the reference time-phase unified image coordinate system. This represents the maximum ordinate of the crack detection box in the reference time-phase unified image coordinate system. Specifically, the crack detection box coordinates are first normalized to [0,1], and then a two-layer Multi-Layer Perceptron (MLP) is used to map the 4D coordinates into a d-dimensional spatial feature vector. ,in, Let represent a d-dimensional real space. Simultaneously, within the image region corresponding to the crack detection box, ROI features are extracted from the segmentation model's backbone feature map using Region of Interest Alignment (ROI Alignment), and then processed by global average pooling to obtain a d-dimensional visual feature vector. Adding the two together yields the final detection bounding box cue feature vector 511. .

[0239] (2) The historical mask morphology encoder 520 implements three-layer convolution + global pooling.

[0240] The historical mask 502 for the crack in previous time phases is obtained from the historical instance database. A lightweight convolutional encoder is used to extract the crack's morphological features. The encoder structure is: 3 convolutional layers (Conv-BN-ReLU, channel count 16→32→64) → global average pooling → fully connected layer, outputting a d-dimensional historical time-phase crack morphological feature vector 522. .

[0241] The information captured by the historical mask morphology encoder 520 includes: the general direction of the crack (through the spatial response pattern of the convolutional kernel), the branch structure (through the topological information preserved in the feature map), and the width distribution (through the activation intensity distribution of the feature map). When there are segmentation masks for multiple time phases in the historical instance library, the mask encoding result of the most recent time phase is taken as the morphological cue, or a weighted average of the mask encoding results of multiple time phases can be taken.

[0242] (3) Material context semantic encoder 530, which realizes material distribution vector × embedding matrix.

[0243] Material classification information 503 within the crack detection frame region is obtained based on the S2 material partitioning map. Specifically, the pixel proportion of each material classification information 503 within the crack detection frame region is statistically analyzed to form a material distribution vector. (The sum of all components is 1), where, This represents the percentage of pixels within the crack detection frame that are classified as wood structures. This represents the percentage of pixels within the crack detection frame that are classified as masonry. This represents the percentage of pixels within the crack detection frame that are categorized as gray plastic, plaster, or painted areas. This represents the percentage of pixels within the crack detection bounding box that are classified as other materials or background. This is achieved through a learnable material embedding matrix. (C represents the number of material categories) Representing a real space of C×d, the material distribution vector is mapped to a d-dimensional material semantic vector. .

[0244] The material semantic vector 533 encodes the material composition of the current detection box region, enabling the segmentation model 550 to know the type of interfering texture in the region (e.g., if the main material is wood, the wood grain needs to be suppressed; if the main material is brick or stone, the brick joints need to be suppressed).

[0245] (4) Cross-modal attention fusion module.

[0246] Encoding vectors for three modes Fusion is achieved through a cross-modal cross-attention mechanism. Specifically, the feature vector 511 is used as a cue box. For the query (hereinafter referred to as Q), the historical temporal crack morphology feature vector 522 is used. Material semantic vector 533 For the key / value pair (hereinafter referred to as K / V), perform multi-head cross-attention fusion 541, as shown in formula (3).

[0247] (3).

[0248] in, Indicates will and Stacking along the token dimension forms a 2×d key / value matrix. MultiHead() represents the multi-head cross-attention fusion algorithm. This indicates a multimodal fusion prompt feature.

[0249] The specific calculation process for multi-head cross-attention is as follows.

[0250] 1. Query ,key ,value Each attention head is mapped to H attention heads using a linear projection matrix, and each head has a dimension of . .

[0251] 2. For the h-th attention head, calculate the attention weight. ,in, The detection box prompts the feature vector. The query vector obtained after linear projection of the query corresponding to the h-th attention head. , Represented by the historical mask feature vector and material semantic vector The key matrix obtained after splicing and linear projection onto the key corresponding to the h-th attention head. , for transpose, express The real space, express The real number space.

[0252] 3. Attention weight It contains two components, corresponding to the importance of the history mask and material context to the current bounding box, respectively. It represents a 1×2 dimensional real number space.

[0253] 4. After calculating the weighted output for each head, the results are concatenated, and the final multimodal fusion cue feature is obtained through the output projection matrix. 542 .

[0254] The cross-attention mechanism enables the fusion process to adaptively learn which information from the historical morphology and material context is most helpful for segmentation given the location of the detection box. For example, when the detection box is located in a wooden structure region, the fusion module automatically pays more attention to the crack direction information in the historical mask (used to distinguish cracks from wood grain) and less attention to brick joint information in the material encoding. Conversely, when the detection box is located in a brick or masonry region, the fusion module will make more use of the brick joint suppression information in the material encoding.

[0255] Multimodal fusion cue features 542 The prompt input format is mapped to the segmentation model through a linear projection layer 540. The prompt drives the segmentation model 550 to perform accurate crack segmentation within the current detection box area to obtain a crack instance mask 551.

[0256] (5) Network training strategies.

[0257] The training of the cross-modal cue fusion network employs a two-stage strategy.

[0258] Phase 1: On the crack segmentation dataset, the entire segmentation pipeline (fusion network + frozen segmentation model decoder) is trained using a triplet of detection box + mask + material as input and manually annotated precise crack masks as supervision signals. The loss function is the weighted sum of the Dice loss and the Binary Cross Entropy (BCE) loss between the segmentation mask and the labeled mask, with a typical ratio of Dice:BCE = 1:1 (equal weights).

[0259] The second stage involves fine-tuning the multi-temporal crack data using time-series data. Crack masks from adjacent temporal phases are used as historical mask inputs, enabling the fusion network to learn how to effectively utilize the continuity of temporal patterns.

[0260] S5.4, Multimodal Enhanced Segmentation Execution Process.

[0261] S5.4.1, Reference Time Phase Segmentation Strategy (for T0).

[0262] For the baseline time-phase image T0, crack segmentation is performed using standard bounding boxes and material context cues. The core implementation method is as follows: based on the feature vector of the crack detection box cue and the material semantic vector of the corresponding region of the crack detection box in the current time phase, pixel-level segmentation of the crack in the current time phase is performed to generate a crack instance mask for the current time phase. The specific steps are as follows.

[0263] 1. Take the T0 crack detection frame output by S4.

[0264] 2. Treat each detection box as a standard box prompt.

[0265] 3. Obtain material classification information for the detection frame area based on the S2 material partition map as auxiliary prompts (if applicable);

[0266] 4. Input the segmentation model. The model performs segmentation within the image region specified by the detection box, combining material classification information, and outputs a pixel-level mask of the crack.

[0267] 5. Perform the above operation on each detection box to obtain the crack instance mask set of all crack instances in T0.

[0268] S5.4.2, Non-reference time phase multimodal enhancement cue segmentation strategy (for T1~Tn).

[0269] For non-reference time-phase images T1, T2, ..., Tn, different segmentation strategies are adopted based on the association between the detection results and the historical instance library: Cracks associated with historical instances are segmented using multimodal enhancement cues that fuse detection boxes, historical masks, and material context. The core implementation method is as follows: pixel-level segmentation is performed on the current time-phase crack based on the detection box cue feature vector of the current time-phase crack detection box, the material semantic vector of the corresponding region of the current time-phase crack detection box, and the morphological feature vector of the historical time-phase cracks associated with the current time-phase crack, generating a current time-phase crack instance mask; newly discovered cracks are segmented using detection boxes and material context cues. The core method is the same as the core implementation method for T0 mentioned above, and will not be repeated here. A detailed description follows.

[0270] The multimodal enhancement segmentation process for non-reference time phases is as follows.

[0271] 1. Retrieve the current phase crack detection box set output by S4.

[0272] 2. For each detection box, query the historical instance library to see if there is a matching historical crack instance (spatial overlap or proximity) based on its spatial location, in order to determine whether the detection box corresponds to a known crack.

[0273] 3. For detection boxes that have been successfully associated with historical instances.

[0274] a. Obtain the segmentation mask of the associated historical instance in the previous time phase (modal 2).

[0275] b. Obtain the material classification information (modal 3) of this region based on the S2 material partition map.

[0276] c. Merge the detection bounding box (modal 1), history mask (modal 2), and material context (modal 3) into a multimodal tooltip.

[0277] d. Input the segmentation model to obtain the current temporal crack instance mask.

[0278] 4. For detection boxes that are not associated with historical instances (newly detected crack candidates).

[0279] a. Use only the detection box as the standard box prompt.

[0280] b. Obtain the material classification information of the detection box area based on the S2 material partition map as an auxiliary prompt (if any).

[0281] c. Input the segmentation model to obtain the crack instance mask.

[0282] 5. Output the mask set of all crack instances in the current time phase.

[0283] S5.5, Output.

[0284] The current phase contains a set of crack instance masks, each mask corresponding to a crack detection box, which records the precise pixel-level outline of the crack.

[0285] S6, Confidence Management.

[0286] S6.1 Overview.

[0287] For each crack instance in the historical instance database, a cross-temporal cumulative confidence score (range 0-1) is maintained to quantify the reliability of the crack instance as a real crack. After processing at each temporal stage, the cumulative confidence score is dynamically updated based on the association matching results of the current temporal stage: the cumulative confidence score of successfully associated instances is updated incrementally, while the cumulative confidence score of unmatched instances is updated by decaying. When the cumulative confidence score of an instance is lower than the elimination threshold for multiple consecutive temporal stages, it is marked as a false detection. The historical record of this instance is retained in the database for traceability, but it will no longer participate in the detection threshold conditions and segmentation prompts for subsequent temporal stages.

[0288] S6.2 Confidence initialization.

[0289] The core implementation method is as follows: For the current phase crack instance mask of a crack instance in the preprocessed current phase image that is not associated with a historical phase crack instance mask, the cumulative confidence of the current phase crack instance is determined based on the crack detection confidence of the current phase crack detection box corresponding to the current phase crack instance mask. This is described in detail below.

[0290] For the crack instance first detected in the reference time phase T0, its initial cumulative confidence is set as the confidence score of the current detection, as shown in formula (4).

[0291] C_init = conf_detect(4).

[0292] Here, `conf_detect` represents the detection confidence score output by the object detection model (ranging from 0 to 1), and `C_init` represents the current detection confidence score. For newly discovered crack instances in subsequent time phases (detection boxes that do not match historical instances), the initial confidence score is also set to its detection confidence score.

[0293] S6.3, Association Matching.

[0294] In non-reference time phases, it is necessary to associate and match the detection results of the current time phase with existing instances in the historical instance library to determine which current detections correspond to which historical cracks.

[0295] Cost matrix construction: Suppose that N crack candidate detection boxes are detected in the current time phase, and there are Z historical instances in the historical instance library that have not been eliminated. Construct an N×Z association cost matrix Cost, where each element Cost(n,z) represents the association cost between the nth crack candidate detection box and the zth historical instance, as shown in formula (5).

[0296] (5).

[0297] The relevant parameters are explained below.

[0298] This is the intersection-union ratio (IoU) between the nth current detection box and the zth historical instance's current detection box.

[0299] The cosine similarity of the depth features between the region corresponding to the nth current detection box and the region corresponding to the current detection box of the zth historical instance is calculated after extracting ROI features from the backbone feature map of the detection model.

[0300] The normalized spatial distance penalty term is calculated by dividing the Euclidean distance of the detection box center point by the length of the image diagonal, and mapping the value range to [0,1]. When the normalized distance exceeds the preset threshold, the normalized distance value is taken; otherwise, it is 0, which is used to introduce temporal continuity constraints.

[0301] Let be the weighting coefficients of each cost term, satisfying... Typical value (IoU weight) (Feature similarity weight) (Spatial distance weighting) prioritizes spatial overlap and appearance consistency.

[0302] Matching Algorithm: Based on the cost matrix, the Hungarian Algorithm is used to solve for optimal bipartite graph matching, minimizing the total association cost. An association cost threshold is set. When associated costs The matching pair should be rejected to avoid forcibly associating detection results with significant differences in spatial location or appearance.

[0303] The matching results are processed as follows.

[0304] Successfully matched crack candidate detection box - instance pair: The current detection is considered to be a re-observation of the historical crack in the current time phase, and S6.4 incremental update is performed.

[0305] Unmatched crack candidate detection boxes: are considered newly discovered crack candidates, and S6.2 initialization is performed and written to the instance library.

[0306] Unmatched historical instance: This historical crack was not detected in the current phase. Perform S6.5 decay update.

[0307] S6.4 Incremental confidence update (successful association).

[0308] The core implementation method is as follows: For the current phase crack instance mask of the crack instance that has been associated with the historical phase crack instance mask in the preprocessed current phase image, the cumulative confidence of the current phase crack instance is determined according to the incrementing coefficient and the cumulative confidence of the previous phase crack instance associated with the current phase crack detection box. The incrementing coefficient is used to ensure that the cumulative confidence of the current phase crack instance is greater than the cumulative confidence of the previous phase crack instance but less than 1. This is described in detail below.

[0309] When a crack instance in the historical instance database is successfully associated with the detection result in the current time phase, it indicates that the crack has been observed in another time phase, and its reliability is enhanced. At this time, the cumulative confidence can be updated incrementally using the incremental formula shown in formula (6).

[0310] C1_new = min(1, C1_old + α×(1 – C1_old))(6).

[0311] The relevant parameters are explained below.

[0312] C1_old represents the cumulative confidence level before the update.

[0313] C1_new represents the updated cumulative confidence level.

[0314] α is an incrementing coefficient, with a value range of (0, 1), and a typical value of 0.1 to 0.3.

[0315] The properties of incremental updates are as follows.

[0316] C1_new is always no less than C1_old, meaning that the confidence level increases monotonically when the association is successful.

[0317] The increment α×(1 – C1_old) decreases as C1_old increases, meaning that the higher the confidence level, the smaller the increment of a single increment, so that the confidence level gradually approaches 1 but does not reach 1 quickly.

[0318] The larger α is, the faster the confidence increases, and the greater the positive incentive for the success of a single association.

[0319] S6.5, Confidence decay update (no match).

[0320] The core implementation method is as follows: For a target historical time-phase crack instance mask that is not associated with the preprocessed current time-phase image, the cumulative confidence of the target historical time-phase crack instance mask is updated based on the attenuation coefficient, the number of times the target historical time-phase crack instance mask is not associated with the preprocessed current time-phase image, and the cumulative confidence of the target historical time-phase crack detection box corresponding to the target historical time-phase crack instance mask. Here, the target historical time-phase crack instance mask represents a historical time-phase crack instance mask whose cumulative confidence is less than a preset elimination threshold for a consecutive preset number of time phases. The attenuation coefficient is used to ensure that the cumulative confidence of the updated target historical time-phase crack instance is less than the cumulative confidence of the target historical time-phase crack instance before the update. This is described in detail below.

[0321] When a crack instance in the historical instance database does not match any detection results in the current time phase, it means that the crack has not been observed in the current time phase. At this time, the decay formula shown in formula (7) can be used to decay and update the cumulative confidence.

[0322] C2_new = C2_old * β(7).

[0323] The relevant parameters are explained below.

[0324] C2_old represents the cumulative confidence level before the update.

[0325] C2_new represents the updated cumulative confidence level.

[0326] β is the attenuation coefficient, with a value range of (0, 1) and a typical value of 0.7~0.9.

[0327] The properties of decay updates are as follows.

[0328] C2_new is always less than C2_old, meaning that the confidence level decreases monotonically when there is no match.

[0329] Each decay occurs at a fixed rate, and multiple consecutive mismatches will cause the confidence level to drop exponentially.

[0330] The smaller the β, the faster the decay and the more severe the penalty for unmatched pairs.

[0331] S6.6 Parameter selection principles and sensitivity description.

[0332] The selection of the increment coefficient α and the decay coefficient β directly affects the quality of the historical instance database and the long-term stability of the detection system.

[0333] The selection principle of the increment coefficient α: α controls the response speed of confidence to a single successful association. When α is too large (e.g., >0.5), a single association can cause a significant increase in confidence, making it difficult for false detection instances to be eliminated through subsequent decay mechanisms, thus increasing the proportion of false detections in the instance library; when α is too small (e.g., <0.05), confidence accumulation is extremely slow, and even multiple consecutive successful associations of real cracks are unlikely to reach a high confidence threshold, making it difficult to trigger low-threshold, high-sensitivity detection, thus reducing the efficiency of utilizing historical information. Typical values ​​of α∈[0.1, 0.3] balance accumulation speed and fault tolerance.

[0334] The selection principle of the attenuation coefficient β: β controls the severity of the non-match penalty. When β is too small (e.g., < 0.5), a single non-match can cause a significant drop in confidence, increasing the risk of real cracks being quickly and falsely eliminated due to temporary occlusion or detection omissions, thus compromising the stability of the historical instance database. When β is too large (e.g., > 0.95), the attenuation is too slow, and false detections that have disappeared or cracks that have been repaired will remain in the historical instance database for a long time, continuously affecting the known region division and detection threshold setting in subsequent time phases. Typical values ​​of β ∈ [0.7, 0.9] strike a balance between quickly eliminating false detections and tolerating temporary occlusion.

[0335] The coordination relationship between α and β: Under steady state, a crack instance needs to be successfully associated approximately 1 / α times consecutively to bring its confidence level close to the high confidence threshold from its initial value. It requires approximately ln(Th_remove / C_init) / ln(β) consecutive unmatched phases to be eliminated. The system design should ensure that the number of consecutive unmatched phases required for elimination is greater than the maximum number of occlusion durations that may occur in practice, to ensure that real cracks are not mistakenly eliminated.

[0336] S6.7 False detection elimination mechanism.

[0337] To prevent low-quality instances from occupying the instance library for a long time and affecting subsequent detection, a false detection elimination mechanism is set up.

[0338] Elimination criteria: When the cumulative confidence of a crack instance is lower than the elimination threshold Th_remove for M consecutive time phases, it is marked as a false detection. The historical record of the instance is kept in the library for traceability, but it will no longer participate in the detection and segmentation feedback of subsequent time phases.

[0339] The relevant parameters are explained below.

[0340] Th_remove is the elimination threshold, with a value range of (0, 0.3), and a typical value of 0.1~0.2.

[0341] M is the phase number threshold for consecutive low confidence levels, typically 2 to 5, used to avoid false eliminations due to single occlusion or missed detection.

[0342] Semantic interpretation of elimination: If a crack instance is not detected in multiple consecutive time phases and its confidence continuously decays to an extremely low value, the most likely reason is that the instance itself is a false detection (such as misidentifying brick seams or wood grain as a crack), rather than a temporary obscuring of a real crack. Eliminating it can clean up the instance library and improve the quality of subsequent time phase detection and association.

[0343] S6.8, Output.

[0344] The cumulative confidence level after each crack instance is updated.

[0345] Marked instances of false positives that have been eliminated.

[0346] Confidence update records (for tracing and auditing).

[0347] S7, Historical Instance Library Update.

[0348] S7.1 Overview.

[0349] After the detection, segmentation, and confidence management of the current phase are completed, all processing results are written to the historical instance library. The historical instance library is the core data structure of this invention. It stores a complete record of all crack instances across time phases and provides historical detection information for the next phase's S4 adaptive detection and S5 (multimodal enhanced segmentation).

[0350] According to an embodiment of the present invention, at least the crack instance identifier, corresponding temporal information, crack detection box, crack instance mask, crack detection confidence, cumulative confidence, and material classification information of the region for each crack instance detected for each temporal image need to be recorded in the historical instance database.

[0351] S7.2, Data structure of the historical instance library.

[0352] Each crack instance in the historical instance library can contain the information fields in Table 3:

[0353] Table 3:

[0354]

[0355] S7.3, Update Process.

[0356] After the current phase is processed, update the historical instance library according to the following procedure.

[0357] For existing instances that have been successfully associated.

[0358] 1. Add the coordinates of the detection frame at the current time phase to the spatial information record of this instance.

[0359] 2. Append the segmentation mask of the current phase to the mask record of this instance.

[0360] 3. Update the "current detection box" and "current mask" of this instance to the latest values ​​of the current phase.

[0361] 4. Add the current time phase number to the "Observation Time Phase List".

[0362] 5. Update the "Last Detection Phase" to the current phase.

[0363] 6. Write the cumulative confidence level after the S6 update.

[0364] 7. Add the original detection confidence of the current phase to the "Detection Confidence Sequence".

[0365] 8. Reset "Consecutive Unmatched Count" to 0.

[0366] 9. Update "Association Status" to "Associated".

[0367] 10. Update structured properties (such as crack length, width, area, and other geometric features).

[0368] For newly discovered crack instances (no historical instances were matched).

[0369] 1. Create a new instance record in the historical instance database and assign a unique instance ID.

[0370] 2. Both the "first detection phase" and the "last detection phase" should be the current phase.

[0371] 3. Write the detection frame coordinates and segmentation mask for the current phase.

[0372] 4. Set the initial cumulative confidence level to the detection confidence level (S6.2).

[0373] 5. Set "Consecutive Unmatched Count" to 0 and "Eliminated" to No.

[0374] 6. Write the material category information of the area where the crack is located based on the S2 material partition map.

[0375] 7. If there are structured attributes, include them as well.

[0376] For existing instances where the current time phase does not match.

[0377] 1. Update the cumulative confidence level to the decayed value (S6.5).

[0378] 2. Increment the "consecutive unmatch count".

[0379] 3. Update "Association Status" to "Not Matched".

[0380] 4. Check if the elimination conditions are met (S6.7). If they are met, mark "whether it is eliminated" as yes. This instance will no longer participate in the known area division and detection feedback of subsequent time phases.

[0381] For instances that were eliminated.

[0382] 1. Mark "Whether it has been phased out" as yes.

[0383] 2. Record the elimination phase as the current phase.

[0384] 3. The historical records of this instance are kept in the database for future reference, but will no longer be used for known region division, detection threshold adjustment, and segmentation prompt feedback in subsequent time phases.

[0385] S7.4 Feedback output of the historical instance library.

[0386] Figure 6 A schematic diagram of a closed-loop feedback mechanism for a historical instance library according to an embodiment of the present invention is shown.

[0387] like Figure 6 As shown, the solid line describes the data interaction process based on steps S2 to S7 for updating the historical instance library. The dashed line describes how the updated historical instance library can provide two types of key historical detection information for processing in the next time phase. The specific details are as follows.

[0388] On the one hand, feedback is sent to S4 (material-aware crack detection).

[0389] Cumulative confidence score for each crack instance: used to classify the confidence level (high / medium / low) of known areas, thereby determining the detection threshold for each area.

[0390] The current bounding box position for each crack instance: used to locate the extent of a known region in the image space.

[0391] On the other hand, it feeds back to S5 (multimodal enhancement cue segmentation).

[0392] Historical mask for each crack instance: serving as a historical morphological reference to assist the segmentation model in accurately segmenting cracks in the current time phase.

[0393] Material category information for each crack instance (from the S2 material partition map): This serves as a material context hint, guiding the segmentation model to suppress interfering textures of specific materials.

[0394] S7.5, Output.

[0395] The updated historical instance library includes: complete cross-temporal records of all crack instances that have not been eliminated; updated cumulative confidence; current phase detection boxes, segmentation masks, structured attributes, and relationships; and elimination records.

[0396] Figure 7 A block diagram of a building crack timing detection device according to an embodiment of the present invention is shown.

[0397] like Figure 7 As shown, the building crack time-series detection device 700 includes an image acquisition module 710, an image registration module 720, a material recognition and partitioning module 730, a material adaptive texture suppression module 740, a material-aware crack detection module 750, and a multimodal fusion prompting segmentation module 760.

[0398] The image acquisition module 710 is used to acquire a multi-temporal image sequence of the building surface, which includes a reference temporal image and non-reference temporal images.

[0399] The image registration module 720 is used to spatially register non-reference time-phase images based on the spatial coordinates of the reference time-phase image to obtain a registered multi-time-phase image sequence.

[0400] The material recognition and partitioning module 730 is used to perform pixel-by-pixel material classification on the reference temporal image based on semantic segmentation and generate a material partitioning map.

[0401] The material adaptive texture suppression module 740 is used to perform differential texture suppression preprocessing on different material regions in the registered phase images according to the material partition map, and output the preprocessed multi-phase image sequence.

[0402] The material-aware crack detection module 750 is used to perform material-aware crack detection on preprocessed temporal images based on a material partition map. It combines the detection thresholds determined for each known region obtained from the division of the preprocessed temporal images to obtain a set of crack detection boxes for each preprocessed temporal image. Each crack detection box has a crack detection confidence level.

[0403] The multimodal fusion prompting segmentation module 760 is used to perform pixel-level segmentation of the current-phase cracks within the current-phase crack detection box of the preprocessed current-phase image based on the multimodal fusion prompting features, generate a current-phase crack instance mask, and obtain the current-phase crack detection result of the building surface. The multimodal fusion prompting features at least fuse the features of the current-phase crack detection box and the material classification information of the corresponding region of the current-phase crack detection box determined based on the material partitioning map. If the current-phase crack detection box is associated with historical-phase crack instances, the multimodal fusion prompting features also fuse the morphological features of the corresponding historical-phase crack instance mask. Historical-phase crack instances have cumulative confidence scores, and the historical-phase crack instance mask and the corresponding historical-phase crack instance cumulative confidence scores are stored in a historical instance library. The historical instance library stores cross-temporal detection information for all crack instances and is updated based on the current-phase crack detection result for subsequent temporal detection threshold adjustment and pixel-level segmentation prompt generation.

[0404] Any one or more of the modules according to embodiments of the present invention, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules according to embodiments of the present invention can be implemented by splitting them into multiple modules. Any one or more of the modules according to embodiments of the present invention can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, and firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules according to embodiments of the present invention can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.

[0405] For example, any plurality of the image acquisition module 710, image registration module 720, material recognition and partitioning module 730, material adaptive texture suppression module 740, material-aware crack detection module 750, and multimodal fusion prompting segmentation module 760 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of the present invention, at least one of the image acquisition module 710, image registration module 720, material recognition and partitioning module 730, material adaptive texture suppression module 740, material-aware crack detection module 750, and multimodal fusion prompting segmentation module 760 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the image acquisition module 710, image registration module 720, material recognition and partitioning module 730, material adaptive texture suppression module 740, material-aware crack detection module 750, and multimodal fusion prompting segmentation module 760 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0406] Figure 8 A block diagram of an electronic device suitable for implementing a time-series detection method for building cracks according to an embodiment of the present invention is shown. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0407] like Figure 8 As shown, an electronic device 800 according to an embodiment of the present invention includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage portion 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0408] RAM 803 stores various programs and data required for the operation of electronic device 800. Processor 801, ROM 802, and RAM 803 are interconnected via bus 804. Processor 801 executes various operations of the method flow according to embodiments of the present invention by executing programs in ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than ROM 802 and RAM 803. Processor 801 may also execute various operations of the method flow according to embodiments of the present invention by executing programs stored in said one or more memories.

[0409] According to an embodiment of the present invention, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to a bus 804. The system 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A driver 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the driver 810 as needed so that computer programs read from it can be installed into the storage section 808 as needed.

[0410] According to embodiments of the present invention, the method flow according to embodiments of the present invention can be implemented as a computer software program. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by processor 801, it performs the functions defined in the system of the embodiments of the present invention. According to embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0411] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the building crack timing detection method according to embodiments of the present invention.

[0412] According to embodiments of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0413] For example, according to embodiments of the present invention, a computer-readable storage medium may include ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803 as described above.

[0414] Embodiments of the present invention also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of the present invention. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the building crack timing detection method provided in the embodiments of the present invention.

[0415] When the computer program is executed by the processor 801, it performs the functions defined in the system / apparatus of this embodiment of the invention. According to embodiments of the invention, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0416] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 809, and / or installed from a removable medium 811. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0417] According to embodiments of the present invention, program code for executing the computer programs provided in the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0418] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not expressly stated in the present invention. In particular, the features described in the various embodiments and / or claims of this invention can be combined and / or combined in various ways without departing from the spirit and teachings of this invention. All such combinations and / or combinations fall within the scope of this invention.

[0419] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of the invention is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.

Claims

1. A method for time-series detection of building cracks, characterized in that, The method includes: Acquire a multi-temporal image sequence of a building surface, the multi-temporal image sequence including a reference temporal image and non-reference temporal images; Based on the spatial coordinates of the reference time phase image, the non-reference time phase image is spatially registered to obtain a registered multi-time phase image sequence. Based on semantic segmentation, pixel-by-pixel material classification is performed on the reference temporal image to generate a material partition map; Based on the material partitioning map, differential texture suppression preprocessing matching the texture interference type of the material region is performed on different material regions in the registered phase images, and the preprocessed multi-phase image sequence is output. Based on the material partitioning map, material-aware crack detection is performed on each preprocessed temporal image. Combined with the detection threshold determined for each known region obtained from the division of each preprocessed temporal image, a set of crack detection boxes for each preprocessed temporal image is obtained. Each crack detection box has a crack detection confidence level. Based on the multimodal fusion prompt features, the current-phase cracks within the current-phase crack detection box of the preprocessed current-phase image are segmented at the pixel level to generate a current-phase crack instance mask, thus obtaining the current-phase crack detection result of the building surface. The multimodal fusion prompt features fuse at least the features of the current-phase crack detection box and the material classification information of the corresponding region of the current-phase crack detection box determined based on the material partitioning map. If the current-phase crack detection box is associated with historical-phase crack instances, the multimodal fusion prompt features also fuse the morphological features of the corresponding historical-phase crack instance mask. The historical-phase crack instances have cumulative confidence scores, and the historical-phase crack instance masks and their cumulative confidence scores are stored in a historical instance database. This database stores cross-phase detection information for all crack instances and is updated based on the current-phase crack detection result for subsequent phase detection threshold adjustment and pixel-level segmentation prompt generation.

2. The method according to claim 1, characterized in that, The material partition map includes a wood structure region; the step of performing differential texture suppression preprocessing on different material regions in the registered temporal images according to the material partition map, matching the texture interference type of the material region, includes: Estimate the grain orientation field and extract the main extension direction of the grain for the grayscale normalized wood structure region; Based on the main extension direction, an adaptive Gaussian kernel or a directional mask is constructed at each pixel position in the normalized wood structure region, wherein the adaptive Gaussian kernel represents a Gaussian kernel with a kernel width parallel to the main extension direction that is greater than the kernel width perpendicular to the main extension direction.

3. The method according to claim 1, characterized in that, The material partition map includes masonry areas; the step of performing differential texture suppression preprocessing on different material areas in the registered temporal images according to the material partition map, matching the texture interference type of the material area, includes: Extract the linear structural features of the masonry area; Cluster analysis is performed on the linear structural features according to direction to identify at least one group of parallel line segments and the dominant direction of each group of parallel line segments; For each group of parallel line segments, project along the normal direction of the dominant direction of the parallel line segment, and calculate the spacing sequence between each adjacent line segment in the parallel line segment and the spacing variation coefficient of the spacing sequence; If the spacing variation coefficient is less than the preset spacing variation threshold and the number of parallel line segments is greater than or equal to the preset minimum value, the parallel line segments are determined to be regular line structures. A suppression mask is generated for the region where the target line segment, identified as having the regular line structure, is located. Set the gradient response value of the pixel grayscale value inside the suppression mask to be lower than the gradient response value of the pixel grayscale value outside the suppression mask, and / or set the contrast of the pixel grayscale value outside the suppression mask to be higher than the contrast of the pixel grayscale value inside the suppression mask.

4. The method according to claim 1, characterized in that, The material partition map includes decorative areas defined by at least one of plastering, stenciling, and painted decoration; the step of performing differential texture suppression preprocessing on different material areas in the registered temporal images according to the material partition map, matching the texture interference type of the material area, includes: Edge detection is performed on the decorative area to obtain edge line segments; If an edge segment meets the non-crack edge filtering conditions, the edge segment is determined to be a non-crack edge. The non-crack edge filtering conditions include at least one of the following: the mean curvature of the edge segment is less than a preset mean curvature threshold and the curvature variance is less than a preset curvature variance threshold; the width of the edge segment is greater than a preset width threshold; and the gray-scale distribution characteristics of the region where the edge segment is located are symmetrical with the gray-scale distribution characteristics of the neighborhood of the edge segment along the normal direction of the edge segment. The neighborhood of the edge segment represents the building surface region adjacent to the edge segment along the normal direction of the edge segment. Set the edge detection response value of the pixel grayscale value in the area where the non-crack edge is located to be lower than the edge detection response value of the pixel grayscale value outside the area where the non-crack edge is located, and / or set the contrast of the pixel grayscale value outside the area where the non-crack edge is located to be higher than the contrast of the pixel grayscale value in the area where the non-crack edge is located.

5. The method according to claim 1, characterized in that, The process of performing material-aware crack detection on the preprocessed temporal images based on the material partitioning map includes: Multi-scale feature extraction is performed on the preprocessed images at each time phase to obtain detection feature maps at multiple scales; The material partition map is feature encoded to obtain material feature maps at multiple scales, wherein the spatial size of the material feature map at each scale matches the spatial size of the detection feature map at the corresponding scale. Based on the channel-wise scaling factor and offset factor generated from the material feature map, the detection feature maps at multiple scales and the material feature maps at multiple scales are fused to obtain a material modulation feature map. Crack detection is performed on the material modulation feature map to obtain crack candidate detection boxes and crack detection confidence scores of the crack candidate detection boxes.

6. The method according to claim 5, characterized in that, The set of crack detection boxes for each preprocessed temporal image, obtained by combining the detection thresholds determined for each known region obtained from the segmentation of each preprocessed temporal image, includes: For the preprocessed reference time-phase image: Obtain the first crack candidate detection box and the first crack detection confidence of the first crack candidate detection box obtained from the preprocessed reference temporal image; Based on a preset standard detection threshold, a target first crack candidate detection box with a target first crack detection confidence level is determined, and a crack detection box set of the preprocessed reference time phase image is obtained, wherein the target first crack detection confidence level is greater than or equal to the standard detection threshold. For the preprocessed non-reference temporal images: Read the crack detection bounding boxes and cumulative confidence scores of all detected crack instances from the historical instance library; Based on the crack detection bounding boxes and cumulative confidence scores of all crack instances, the preprocessed non-reference temporal images are divided into a first confidence known region, a second confidence known region, and a third confidence known region. The first confidence known region is determined based on the crack detection bounding boxes of crack instances with a cumulative confidence score greater than or equal to a first confidence threshold. The second confidence known region is determined based on the crack detection bounding boxes of crack instances with a cumulative confidence score less than the first confidence threshold but greater than or equal to the second confidence threshold. The third confidence known region is determined based on the crack detection bounding boxes of crack instances with a cumulative confidence score less than the second confidence threshold. The third confidence known region is also determined based on other regions not covered by historical temporal crack detection bounding boxes read from the historical instance database. The first confidence threshold is greater than the second confidence threshold. Obtain the second crack candidate detection box and the second crack detection confidence of the second crack candidate detection box obtained from the preprocessed non-reference temporal image; Based on the adaptive detection threshold of the known region to which the second crack candidate detection box belongs, a target second crack candidate detection box with target second crack detection confidence is determined, resulting in a set of crack detection boxes for the preprocessed non-reference temporal image. The known region to which the second crack candidate detection box belongs is determined based on the known region where the center of the second crack candidate detection box is located or the known region covered by the second crack candidate detection box. The adaptive detection threshold is determined based on at least one region detection threshold corresponding to the known region to which the second crack candidate detection box belongs. The region detection threshold of the first confidence known region is less than the region detection threshold of the second confidence known region, the region detection threshold of the second confidence known region is equal to the region detection threshold of the third confidence known region, and the target second crack detection confidence is greater than or equal to the adaptive detection threshold.

7. The method according to claim 1, characterized in that, The method further includes: Feature extraction is performed on the current phase crack detection box to obtain the detection box prompt feature vector, wherein the detection box prompt feature vector includes the spatial feature vector and the visual feature vector of the current phase crack detection box; Feature extraction is performed on the historical temporal crack instance mask to obtain the historical temporal crack morphology feature vector; Feature extraction is performed on the material classification information to obtain the material semantic vector; Using the detection box prompt feature vector as the query and the historical time phase crack morphology feature vector and the material semantic vector as the key, multi-head cross-attention calculation is performed to obtain the multimodal fusion prompt feature.

8. The method according to claim 7, characterized in that, The step of performing pixel-level segmentation of the current-phase crack within the current-phase crack detection box of the preprocessed current-phase image based on multimodal fusion cue features, and generating a current-phase crack instance mask, includes: For the current phase crack detection bounding box in the preprocessed reference phase image, or the current phase crack detection bounding box of a crack instance in the preprocessed non-reference phase image that is not associated with a historical phase crack instance mask: Based on the detection box prompt feature vector of the current phase crack detection box and the material semantic vector of the corresponding region of the current phase crack detection box, the current phase crack is segmented at the pixel level to generate a current phase crack instance mask. For the current temporal crack detection bounding box in the preprocessed non-reference temporal image that has been associated with historical temporal crack instance masks: Based on the detection box prompt feature vector of the current phase crack detection box, the material semantic vector of the region corresponding to the current phase crack detection box, and the morphological feature vector of the historical phase crack associated with the current phase crack, the current phase crack is segmented at the pixel level to generate a current phase crack instance mask.

9. The method according to claim 1, characterized in that, The method further includes: For the current phase crack instance mask of the crack instance that is not associated with the historical phase crack instance mask in the preprocessed current phase image, the cumulative confidence of the current phase crack instance is determined based on the crack detection confidence of the current phase crack detection box corresponding to the current phase crack instance mask. For the current phase crack instance mask of the crack instance that has been associated with the historical phase crack instance mask in the preprocessed current phase image, the cumulative confidence of the current phase crack instance is determined according to the incrementing coefficient and the cumulative confidence of the previous phase crack instance associated with the current phase crack detection box. The incrementing coefficient is used to ensure that the cumulative confidence of the current phase crack instance is greater than the cumulative confidence of the previous phase crack instance and less than 1. For a target historical phase crack instance mask that is not associated with the current preprocessed temporal image, the cumulative confidence of the target historical phase crack instance mask is updated based on the attenuation coefficient, the number of times the target historical phase crack instance mask is not associated with the current preprocessed temporal image, and the cumulative confidence of the target historical phase crack detection box corresponding to the target historical phase crack instance mask. The target historical phase crack instance mask represents a historical phase crack instance mask whose cumulative confidence is less than a preset elimination threshold for a consecutive preset number of temporal phases. The attenuation coefficient is used to ensure that the cumulative confidence of the updated target historical phase crack instance is less than the cumulative confidence of the target historical phase crack instance before the update.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: The crack instance identifier, corresponding temporal information, crack detection box, crack instance mask, crack detection confidence, cumulative confidence, and material classification information of the region for each crack instance detected for each temporal image are recorded in the historical instance database.

Citation Information

Patent Citations

  • Cultural relic crack monitoring system and method based on image processing

    CN120912559A

  • Unmanned aerial vehicle building outer wall crack adaptive segmentation method based on deep learning

    CN121504960A