A method for generating underwater pier crack data generation driven by cross-scene scale controllable condition diffusion generation

CN122760718APending Publication Date: 2026-09-15DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610962443.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

[0005]有鉴于此,本发明的目的在于提出一种跨场景尺度可控条件扩散生成驱动的水下桥墩裂缝数据生成方法,以解决目前水下桥墩裂缝真实样本匮乏的技术问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122760718A_ABST
    Figure CN122760718A_ABST
Patent Text Reader

Abstract

The application provides a kind of underwater pier crack data generation method driven by cross-scene scale controllable condition diffusion generation, comprising the following steps: S1, obtaining underwater scene image, land building crack image and crack structure condition information corresponding to crack image;S2, scale modulation processing is carried out on crack structure condition information, and the crack structure condition after scale modulation is obtained;S3, a condition diffusion generation framework is constructed;S4, the crack structure condition after scale modulation is input into the condition control network for feature coding, and the structure guide feature matched with the diffusion generation process is obtained;S5, cross-scene mapping generation is executed, and underwater pier crack image is obtained;S6, underwater pier crack dataset is constructed based on underwater pier crack image of different scale conditions.The application can effectively alleviate the problem of insufficient real samples, improve the generalization ability and robustness of the crack detection model in complex underwater environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and more particularly to a method for generating underwater bridge pier crack data driven by cross-scene scale controllable conditional diffusion generation. Background Technology

[0002] As a critical component of hydraulic and transportation infrastructure, underwater bridge piers are susceptible to cracking due to factors such as water erosion, chemical corrosion, and structural aging during long-term service. Early detection and accurate assessment of cracks are crucial for ensuring bridge structural safety, conducting health status assessments, and making maintenance decisions. However, underwater bridge pier crack detection faces severe challenges from the complex underwater imaging environment, including insufficient lighting, interference from suspended objects, image blurring, and color shifts, resulting in significant differences in the visual appearance of cracks across different scenarios and scales. Furthermore, real-world underwater bridge pier crack samples are extremely scarce due to complex underwater operating conditions, high costs of manual inspection, and limitations in imaging equipment, making it difficult to construct a sufficiently large and diverse training dataset, severely restricting the performance improvement of intelligent crack detection models.

[0003] To address the problem of scarce samples, existing techniques often employ Generative Adversarial Networks (GANs) or diffusion models for image synthesis and data augmentation. These methods generate synthetic images with the visual style of the target domain by modeling the appearance distribution of source domain images. For example, GANs can be used to transfer images of land cracks to underwater scenes, or diffusion models can be used to gradually denoise and generate new samples that conform to the target distribution. Some methods also attempt to introduce semantic segmentation maps or edge detection maps as conditional inputs to guide the structural location of targets in the generated images. Overall, existing generative methods focus on fitting the overall visual style, achieving cross-domain transformation through appearance distribution alignment, and alleviating the data shortage problem to some extent.

[0004] However, existing generation methods still have significant shortcomings when applied to underwater bridge pier cracks. Compared to cracks in terrestrial buildings, underwater bridge pier cracks typically have narrower widths and shorter lengths, exhibiting finer structural features. Existing methods lack effective constraints on the geometric scale and spatial distribution of cracks, making it difficult to accurately characterize the morphological features of these fine cracks. This results in poor performance in terms of structural consistency and engineering semantic rationality. Furthermore, existing methods often rely on statistical distribution of appearance for style transfer, and the spatial location, extension morphology, and topological connectivity of cracks are easily distorted or lost during generation, failing to guarantee strict consistency of crack structure before and after cross-scene mapping. Therefore, a controllable generation method that can explicitly introduce crack structural constraints and possess cross-scene scale adaptability is urgently needed to construct a high-fidelity underwater bridge pier crack dataset to support high-precision training of intelligent crack detection models. Summary of the Invention

[0005] In view of this, the purpose of this invention is to propose a cross-scenario scale controllable conditional diffusion generation method for generating underwater bridge pier crack data, so as to solve the technical problem of the scarcity of real samples of underwater bridge pier cracks.

[0006] The technical means employed in this invention are as follows: A method for generating underwater bridge pier crack data driven by cross-scene scale controllable conditional diffusion includes the following steps: S1. Acquire underwater scene images, land building crack images, and crack structure condition information corresponding to the crack images; S2. The crack structure condition information obtained in S1 is subjected to scale modulation processing to obtain the scale-modulated crack structure condition. S3. Construct a conditional diffusion generation framework, which includes a diffusion generation model and a conditional control network. The diffusion generation model is used to characterize the visual appearance distribution of the land crack image domain and the underwater scene image domain, respectively. S4. Input the scale-modulated crack structure conditions obtained in S2 into the conditional control network in S3 for feature encoding to obtain structure-guided features that match the diffusion generation process. S5. In the diffusion generation process, the structural guidance features obtained in S4 are used as conditional constraints to input the diffusion generation model. The diffusion generation process is modulated by structural constraints. The land building crack image obtained in S1 is used as the source domain and the underwater scene image is used as the target domain appearance constraint. Cross-scene mapping generation is performed to make the generated image have the visual features of the underwater bridge pier scene while maintaining the crack structure morphology, thus obtaining the underwater bridge pier crack image. S6. Repeat S2 to S5 to obtain underwater bridge pier crack images under different scale conditions by changing the scale modulation parameters, and construct an underwater bridge pier crack dataset based on the underwater bridge pier crack images under different scale conditions.

[0007] Furthermore, in S1: The underwater scene image is a bridge pier or bridge structure image from a publicly available underwater image dataset. The underwater scene image contains underwater visual characteristics such as water scattering, light attenuation, and imaging blur, and is used to characterize the environmental appearance features of the underwater bridge pier scene. The land building crack images are sample images of building structure cracks from a publicly available building crack dataset. The land building crack images contain the distribution pattern, extension direction and connectivity features of the cracks, and are used to provide morphological and structural information of the cracks. The crack structural condition information is generated by manual annotation, crack segmentation algorithm or crack detection model. This crack structural condition information includes the spatial location, extension morphology and boundary distribution characteristics of the crack, and is used as the structural constraint condition in the subsequent diffusion generation process.

[0008] Furthermore, in S5: The cross-scene mapping generation uses a diffusion generation model as a carrier, takes land building crack images as source domain input, and underwater scene images as target domain appearance constraints, and constructs a mapping relationship between the source domain and the target domain to perform cross-domain conversion of crack images from land scenes to underwater bridge pier scenes. During the mapping process, the statistical features of the underwater scene image are used as appearance guides to ensure that the generated image is consistent with the underwater scene in terms of color distribution, contrast variation, light attenuation, and imaging blur. During the diffusion generation process, the structural guidance features obtained in S4 are fused with the noise latent variables at the current diffusion moment to obtain conditional latent variables carrying crack structure information, and the conditional latent variables are input into the diffusion generation model for reverse denoising and reconstruction. During the reverse denoising and reconstruction process, the diffusion generation model predicts and updates the denoised latent space features step by step based on the current diffusion time, the conditional latent variables, and the appearance constraints of the target domain, so that the crack structure information continuously participates in the updating of the generated features during the multi-step denoising process. By continuously constraining the latent space features through the aforementioned structural guidance features, the spatial position, extension morphology, and topological connectivity of the crack remain consistent before and after cross-scene mapping. The cross-scene mapping is accomplished through a progressive noise addition and reverse noise reduction reconstruction process of the diffusion generation model. In the multi-step generation process, crack structure information and underwater scene appearance features are gradually fused to output a crack image that conforms to the distribution of the target scene.

[0009] Furthermore, in S3, the conditional control network performs feature encoding and mapping on the crack structure condition information, converting the crack structure condition information into structural guidance features that match the feature space of the diffusion generation model. Specifically, the conditional control network performs multi-layer feature extraction on the input crack structure conditions to obtain a structural feature representation. This representation is then aligned with the intermediate features in the diffusion generation model through spatial scale alignment operations and injected into the generation network via feature fusion during the diffusion generation process, modulating the feature updates during the diffusion process. The feature fusion methods include additive fusion or residual modulation of the input features of the diffusion model, enabling the crack structure information to gradually guide the consistency of the generated cracks in terms of spatial location and morphological structure during the generation process.

[0010] Furthermore, the controllable generation of crack scale in S2 is achieved by scale modulation of crack structural condition information. This involves changing the width and length of the crack region in the crack structural condition to control the scale characteristics of the generated crack, so that the generated crack exhibits corresponding structural morphological changes under different width or thickness conditions. Scale modulation includes morphological transformation processing of crack structural condition information or adjustment of crack width and length based on manual annotation to generate cracks of different scales.

[0011] The present invention also provides a storage medium comprising a stored program, wherein, when the program is executed, it performs any of the above-described cross-scene scale controllable conditional diffusion generation driven underwater bridge pier crack data generation methods.

[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes any of the above-described cross-scenario scale controllable conditional diffusion generation driven underwater bridge pier crack data generation methods through the computer program.

[0013] Compared with the prior art, the present invention has the following advantages: This invention introduces crack structure condition information during the diffusion generation process and explicitly models the crack geometry, achieving controllable generation of crack width, length, and other scale features. This overcomes the problem of existing generation methods failing to effectively represent fine cracks, enabling the generated results to cover a fine-scale distribution that better matches the early crack characteristics of underwater bridge piers. Furthermore, a structural condition constraint mechanism ensures the consistency of generated cracks in spatial location and morphological structure, avoiding structural distortion caused by relying solely on visual appearance modeling. Simultaneously, by constructing a cross-scene mapping mechanism, the invention achieves mapping generation of land cracks to underwater bridge pier scenes, allowing the generated images to retain crack structural features while possessing underwater visual appearance characteristics, thus effectively alleviating the difficulty of obtaining real underwater crack samples.

[0014] The underwater bridge pier crack images generated by the above method have a high degree of consistency with real samples in terms of crack scale distribution, structural features and visual appearance. They can be used to improve the crack detection model's ability to identify small cracks and enhance its adaptability in complex underwater environments, showing good prospects for engineering applications. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the overall framework of the method of the present invention. Detailed Implementation

[0017] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0018] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0019] like Figure 1 As shown, this invention provides a method for generating underwater bridge pier crack data driven by cross-scene scale controllable conditional diffusion generation, including the following steps: S1. Acquire underwater scene images, land building crack images, and crack structure condition information corresponding to the crack images; The underwater scene images are derived from publicly available underwater image datasets and are used to characterize the visual appearance features of the underwater scene in the target domain. The land building crack images are derived from publicly available building crack datasets and are used to provide original crack image samples. The crack structure condition information is generated through manual annotation, crack segmentation algorithms, or existing crack detection models and is used to describe the spatial location, extension morphology, and boundary distribution characteristics of the cracks.

[0020] To ensure the consistency of the input data, the underwater scene images and the land building crack images are preprocessed, including image size unification and normalization.

[0021] S2. The crack structure condition information obtained in S1 is subjected to scale modulation processing to obtain the scale-modulated crack structure condition. In this invention, based on the crack structure condition information obtained in S1, the corresponding crack segmentation map is subjected to geometric shape adjustment processing to obtain crack structure segmentation maps with different scale parameters.

[0022] The scale modulation process includes reducing the width of the crack region and shortening the length of the crack region to simulate the morphological characteristics of early underwater bridge pier cracks, which are thinner and shorter than those of land-based building cracks.

[0023] The scale modulation processing can be achieved by performing morphological transformation on the crack region through image processing methods, or by correcting the crack segmentation map through manual annotation, thereby forming structural constraint information that meets the conditions of different crack scales, providing a basis for the subsequent controllable generation of crack scales.

[0024] S3. Construct a conditional diffusion generation framework, which includes a diffusion generation model and a conditional control network. The diffusion generation model is used to characterize the visual appearance distribution of the land crack image domain and the underwater scene image domain, respectively. S4. Input the scale-modulated crack structure conditions obtained in S2 into the conditional control network in S3 for feature encoding to obtain structure-guided features that match the diffusion generation process. In this invention, a conditional generation framework is constructed based on a diffusion generation model, which includes the diffusion generation model and a conditional control network. The diffusion generation model is used to learn the visual appearance distribution and generation representation of crack images, while the conditional control network processes the crack structure condition information obtained in step S2 and converts it into structure-guided features that match the diffusion generation process.

[0025] Specifically, images of land-based building cracks are input into a diffusion generation model for training, enabling the model to learn the texture features, structural morphology, and background distribution characteristics of the crack images. Scale-modulated crack structural condition information is input into a conditional control network, and structural guidance features are obtained through multi-layer feature extraction and spatial scale alignment. These structural guidance features are injected into the input or intermediate features of the diffusion model during the diffusion generation process using feature superposition, additive fusion, and residual modulation, thereby modulating the reverse denoising generation process with structural constraints.

[0026] By inputting crack structure condition information modulated with different width and length parameters, the condition generation framework can generate crack images corresponding to the input structure conditions, so that the generated cracks are consistent with the input crack structure conditions in terms of spatial location, extension direction, boundary distribution, width, length and connectivity, thereby achieving controllable crack scale generation.

[0027] S5. In the diffusion generation process, the structural guidance features obtained in S4 are introduced to modulate the diffusion generation process with structural constraints. The land building crack image obtained in S1 is used as the source domain and the underwater scene image is used as the target domain appearance constraint. Cross-scene mapping generation is performed to make the generated image have the visual features of the underwater bridge pier scene while maintaining the crack structure morphology, thus obtaining the underwater bridge pier crack image. Specifically, during the diffusion generation process, the structural guidance features obtained in S4 are fused with the noise latent variables at the current diffusion moment to obtain conditional latent variables carrying crack structure information, and the conditional latent variables are input into the diffusion generation model for reverse denoising and reconstruction. During the reverse denoising and reconstruction process, the diffusion generation model predicts and updates the denoised latent space features step by step based on the current diffusion time, the conditional latent variables, and the appearance constraints of the target domain, so that the crack structure information continuously participates in the updating of the generated features during the multi-step denoising process. By continuously constraining the latent space features through the aforementioned structural guidance features, the spatial position, extension shape, and topological connectivity of the crack remain consistent before and after cross-scene mapping.

[0028] After generating scale-controlled cracks, a cross-scene mapper is further constructed to establish the mapping relationship between the land crack domain and the underwater bridge pier scene domain. The land crack domain is used to provide crack images with clear structural features, while the underwater bridge pier scene domain is used to provide the visual appearance features of the underwater environment of the target domain.

[0029] The cross-scene mapper takes a scale-controllable crack image or its latent space representation as input and the visual appearance distribution of an underwater bridge pier scene image as a target domain constraint. By learning the appearance transformation relationship between the two domains, the input crack image generates a crack image with the visual appearance features of an underwater bridge pier scene while maintaining the crack structure characteristics.

[0030] The cross-scene mapping process mainly operates at the level of image appearance expression, making the generated results conform to the characteristics of underwater scenes in terms of color distribution, contrast changes, light attenuation, water scattering, texture representation, and imaging blur. At the same time, through structural preservation constraints, the spatial position, extension shape, boundary distribution, width, length, and connectivity of cracks are kept consistent before and after mapping, avoiding crack structure shift or morphological distortion caused by the cross-scene mapping process.

[0031] Thus, scale-controllable crack images are converted into crack images with visual features of underwater bridge pier scenes, enabling cross-scene generation from land building crack samples to underwater bridge pier crack samples.

[0032] S6. Repeat S2 to S5 to obtain underwater bridge pier crack images under different scale conditions by changing the scale modulation parameters, and construct an underwater bridge pier crack dataset based on the underwater bridge pier crack images under different scale conditions.

[0033] The underwater bridge pier crack images are generated and their corresponding crack structural condition information is compiled to construct an underwater bridge pier crack dataset. The underwater bridge pier crack dataset includes crack image samples with different crack widths, different crack lengths and widths, different spatial distribution patterns, and different underwater visual appearance conditions.

[0034] Furthermore, the crack structure condition information can be used as the structural annotation information corresponding to the generated image to characterize the spatial location, extension morphology and boundary distribution of the crack; the generated underwater bridge pier crack image is used to expand the number of underwater bridge pier crack samples, alleviating the difficulty in obtaining real underwater bridge pier crack samples and the problem of insufficient early small crack samples.

[0035] Therefore, based on the underwater bridge pier crack images and their corresponding structural condition information, a crack sample set with scale-controllable features and underwater scene appearance features can be formed, providing a data foundation for subsequent underwater bridge pier crack detection, identification, or data augmentation tasks.

[0036] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating underwater pier crack data generation driven by cross-scene scale controllable condition diffusion generation, characterized in that, Includes the following steps: S1. Acquire underwater scene images, land building crack images, and crack structure condition information corresponding to the crack images; S2. The crack structure condition information obtained in S1 is subjected to scale modulation processing to obtain the scale-modulated crack structure condition. S3. Construct a conditional diffusion generation framework, which includes a diffusion generation model and a conditional control network. The diffusion generation model is used to characterize the visual appearance distribution of the land crack image domain and the underwater scene image domain, respectively. S4. Input the scale-modulated crack structure conditions obtained in S2 into the conditional control network in S3 for feature encoding to obtain structure-guided features that match the diffusion generation process. S5. In the diffusion generation process, the structural guidance features obtained in S4 are used as conditional constraints to input the diffusion generation model. The diffusion generation process is modulated by structural constraints. The land building crack image obtained in S1 is used as the source domain and the underwater scene image is used as the target domain appearance constraint. Cross-scene mapping generation is performed to make the generated image have the visual features of the underwater bridge pier scene while maintaining the crack structure morphology, thus obtaining the underwater bridge pier crack image. S6. Repeat S2 to S5 to obtain underwater bridge pier crack images under different scale conditions by changing the scale modulation parameters, and construct an underwater bridge pier crack dataset based on the underwater bridge pier crack images under different scale conditions.

2. The underwater bridge pier crack data generation method driven by cross-scene scale controllable conditional diffusion generation according to claim 1, characterized in that, In S1: The underwater scene image is a bridge pier or bridge structure image from a publicly available underwater image dataset. The underwater scene image contains underwater visual characteristics such as water scattering, light attenuation, and imaging blur, and is used to characterize the environmental appearance features of the underwater bridge pier scene. The land building crack images are sample images of building structure cracks from a publicly available building crack dataset. The land building crack images contain the distribution pattern, extension direction and connectivity features of the cracks, and are used to provide morphological and structural information of the cracks. The crack structural condition information is generated by manual annotation, crack segmentation algorithm or crack detection model. This crack structural condition information includes the spatial location, extension morphology and boundary distribution characteristics of the crack, and is used as the structural constraint condition in the subsequent diffusion generation process.

3. The underwater bridge pier crack data generation method driven by cross-scene scale controllable conditional diffusion generation according to claim 1, characterized in that, In S5: The cross-scene mapping generation uses a diffusion generation model as a carrier, takes land building crack images as source domain input, and underwater scene images as target domain appearance constraints, and constructs a mapping relationship between the source domain and the target domain to perform cross-domain conversion of crack images from land scenes to underwater bridge pier scenes. During the mapping process, the statistical features of the underwater scene image are used as appearance guides to ensure that the generated image is consistent with the underwater scene in terms of color distribution, contrast variation, light attenuation, and imaging blur. During the diffusion generation process, the structural guidance features obtained in S4 are fused with the noise latent variables at the current diffusion moment to obtain conditional latent variables carrying crack structure information, and the conditional latent variables are input into the diffusion generation model for reverse denoising and reconstruction. During the reverse denoising and reconstruction process, the diffusion generation model predicts and updates the denoised latent space features step by step based on the current diffusion time, the conditional latent variables, and the appearance constraints of the target domain, so that the crack structure information continuously participates in the updating of the generated features during the multi-step denoising process. By continuously constraining the latent space features through the aforementioned structural guidance features, the spatial position, extension morphology, and topological connectivity of the crack remain consistent before and after cross-scene mapping. The cross-scene mapping is accomplished through a progressive noise addition and reverse noise reduction reconstruction process of the diffusion generation model. In the multi-step generation process, crack structure information and underwater scene appearance features are gradually fused to output a crack image that conforms to the distribution of the target scene.

4. The underwater bridge pier crack data generation method driven by cross-scene scale controllable conditional diffusion generation according to claim 1, characterized in that, In S3, the conditional control network performs feature encoding and mapping on the crack structure condition information, converting the crack structure condition information into structural guidance features that match the feature space of the diffusion generation model. Specifically, the conditional control network performs multi-layer feature extraction on the input crack structure conditions to obtain structural feature representations. These representations are then aligned with intermediate features in the diffusion generation model through spatial scale alignment operations and injected into the generation network via feature fusion during the diffusion generation process, modulating feature updates during the diffusion process. Feature fusion methods include additive fusion or residual modulation of the input features of the diffusion model, enabling the crack structure information to gradually guide the consistency of the generated cracks in terms of spatial location and morphological structure during the generation process.

5. The underwater bridge pier crack data generation method driven by cross-scene scale controllable conditional diffusion generation according to claim 1, characterized in that, The controllable generation of crack scale in S2 is achieved by scale modulation of crack structure condition information. Specifically, by changing the width and length of the crack region in the crack structure condition, the scale characteristics of the generated crack are controlled, so that the generated crack exhibits corresponding structural morphological changes under different width or thickness conditions. Scale modulation involves morphological transformation of crack structure information or adjustment of crack width and length based on manual annotation to generate cracks of different scales.

6. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it performs the underwater bridge pier crack data generation method driven by cross-scene scale controllable conditional diffusion generation as described in any one of claims 1 to 5.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the underwater bridge pier crack data generation method driven by cross-scene scale controllable conditional diffusion generation as described in any one of claims 1 to 5 through the computer program.