Frequency decoupling multi-weather scene crack identification method and system
Through the frequency decoupling of multi-weather scene crack recognition method, the denoising and frequency decoupling modules are used to solve the problem of image quality degradation caused by multiple degradation factors in bad weather, and high-precision crack detection is achieved.
Patent Information
- Application Number
- CN202510586720.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
When the existing crack detection methods face the random superposition of multiple deterioration factors, the recognition effect is poor, especially in severe weather conditions, which reduces the image quality, which affects the detection accuracy.
A frequency decoupling multi-weather scene crack recognition method is proposed. Through the combination of denoising, SharedSTEM processing and frequency decoupling modules, high-frequency and low-frequency semantic information are decoupled to improve the edge accuracy of crack binary segmentation.
It effectively improves the accuracy of crack segmentation under complex weather conditions, and ensures accurate barriers to infrastructure maintenance and public transportation safety.
Smart Images

Figure CN120107260A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a frequency-decoupled multi-weather scenario crack recognition method and system. Background Art
[0002] Surface health monitoring of engineering structures is crucial to infrastructure management. Accurate crack detection and timely repair are directly related to the safety and quality improvement of engineering concrete structures. However, crack topology is complex, has narrow and long directions, and sub-cracks are multiplied linearly with low edge contrast, which poses significant challenges to detection.
[0003] Natural images can be decomposed into low spatial frequency components describing smooth changes and high spatial frequency components representing rapid changes during feature extraction. In the semantic segmentation task, it is considered to directly decouple these two types of information through frequency filters, extract detail semantics, and combine prototype learning to retain the semantic information of high-resolution images. At present, traditional crack detection methods rely on high-quality image datasets, but in practical applications, crack detection often faces many challenges, such as insufficient lighting, low contrast between cracks and background, and image quality degradation caused by bad weather. In addition, noise factors such as engineering surface pollution, light and shadow occlusion, and complex background also significantly affect the detection performance of drones, especially in low-light scenes such as tunnels, subways, and night inspections, where the performance is far inferior to detection under conventional lighting conditions.
[0004] To address the above problems, image enhancement is usually used to improve brightness and contrast before inputting into the segmentation model. However, the imaging quality in bad weather still seriously restricts crack detection, especially in drone inspections, where improving the detection effect under low-quality images is still a problem.
[0005] In drone road health inspections or daily traffic scenarios, images are often affected by multiple weather conditions at the same time. Common situations include sleet, rain in fog, and snow in heavy fog. The random superposition of these multiple degradation factors further deteriorates the image quality, greatly challenging the effectiveness of image restoration and processing. Current research focuses on solving single degradation factors, such as haze, rain and snow. Although some networks are designed to handle combinations of severe weather conditions, such as rain and snow or rain and fog, they are usually limited to fixed combinations. In contrast, real-world scenarios usually involve the random superposition of multiple degradation factors, where two or even three severe weather conditions often lead to severe image degradation. Current methods are relatively poor in solving randomly superimposed crack identification. Summary of the invention
[0006] The present invention proposes a frequency-decoupled multi-weather scenario crack identification method and system to solve the technical problem that the existing identification method has poor identification effect when facing the random superposition of multiple degradation factors.
[0007] In order to solve the above technical problems, the present invention provides a frequency decoupled multi-weather scenario crack identification method, comprising the following steps: Step S1: De-noising the concrete engineering structure image to be identified to obtain a de-noised image; Step S2: down-sampling the denoised image and processing it through SharedSTEM to obtain a first feature; Step S3: inputting the first feature into a plurality of frequency decoupling modules for processing and then classifying to identify whether there are cracks in the concrete engineering structure image; The first frequency decoupling module divides the first feature into grids, clusters the grids, and performs position encoding; extracts high-frequency features and low-frequency features of the clustered grids with position encoding information respectively, and fuses them according to the correlation between the low-frequency features and the high-frequency features; and fuses the fused features with the first feature and inputs them into the subsequent frequency decoupling module as a new first feature.
[0008] Preferably, in step S1, the concrete engineering structure image to be identified is denoised by a multi-weather integrated restoration module; When the multi-weather integrated restoration module is trained, the noise-free crack image is converted into a low-quality crack image through the inverse generative adversarial network AdverseGAN, and the low-quality crack image is denoised through several AIIR modules, and the denoising effect is judged for training optimization; The AIIR module includes several multi-scale spatial attention modules MSA and a feature pyramid FFN. Several MSAs are used to integrate the results of different heads, and after 1×1 convolution superposition, the results are input into FFN to integrate local and global information to obtain the restored crack image.
[0009] Preferably, the expression of the AIIR module is: ; ; In the formula, It represents the preliminary denoising result obtained by concatenating the results of several MSA modules; Cat Represents a splicing operation; Indicates i Results of the MSA modules; Indicates i Characteristics of the group; Represents feature pyramid operation; Represents a 1×1 convolution operation.
[0010] Preferably, every two AIIR modules are processed through a feature mapping network to decouple the image style and the crack style, thereby increasing the distance between the two features.
[0011] Preferably, the low-frequency features are extracted by the following method: the clustering grid is divided into multiple groups, adaptive pooling of different sizes is performed on each group, and up-sampled to the original size, and the up-sampled results are spliced and activated by ReLU to extract the low-frequency features.
[0012] Preferably, the high-frequency features are extracted by the following method: the clustering grid is divided into multiple groups, and for each group, a convolution layer with different kernels is used to simulate the cutoff frequencies of different high-pass filters to form a channel frequency domain information list, and the channel frequency domain information list is connected in the channel dimension through a tensor splicing operation to obtain the high-frequency features.
[0013] Preferably, the expression for fusion is based on the correlation between low-frequency features and high-frequency features. for: ; In the formula, q represents the query vector; Represents high-frequency features; Represents low-frequency features; Frequency Component i and frequency components j The correlation coefficient between .
[0014] Preferably, the features obtained by fusing the high-frequency features and the low-frequency features are input into a recurrent neural network DA-FFN based on a dual-stage attention mechanism for processing before subsequent operations are performed.
[0015] Preferably, the method uses a loss function L To train: ; In the formula, is the i-th element of the actual label vector; is the i-th element of the probability distribution predicted by the model; N represents the number of samples.
[0016] The present invention also provides a frequency-decoupled multi-weather scenario crack identification system, comprising: one or more processors and memories, and one or more programs, wherein the one or more programs are stored in the memories and are configured to be executed by the one or more processors, and the one or more programs include methods for executing the above-mentioned method.
[0017] The beneficial effects of the present invention include at least: the present invention integrates diversified weather restoration and spatial domain frequency guidance, firstly performs denoising through diversified weather, and then uses a stacked frequency decoupling module for segmentation processing. The design effectively decouples high-frequency and low-frequency semantic information, improves the edge accuracy of crack binary segmentation, and solves the problem of poor recognition accuracy of traditional models in inspection scenarios based on low-resource drones due to limited computing resources, ensures high-precision crack segmentation under complex weather conditions, and provides accurate protection for infrastructure maintenance and public transportation safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram of a model structure of an embodiment of the present invention; Figure 2 A schematic diagram of a frequency decoupling module according to an embodiment of the present invention; Figure 3 A schematic diagram of a comparison of the method of an embodiment of the present invention with other model results on multiple data sets; Figure 4 The figure is a schematic diagram for comparing the result data of an embodiment of the present invention with the result data of other models on multiple data sets. DETAILED DESCRIPTION
[0019] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.
[0020] Example 1 The application of lightweight equipment such as drones in road inspection is crucial for traffic engineering structural health monitoring. In crack detection, adverse weather conditions often lead to image quality degradation.
[0021] Therefore, if Figure 1 As shown, the embodiment of the present invention provides a frequency decoupled multi-weather scene crack recognition method, which first denoises the collected concrete engineering structure image to be identified, and then constructs a segmented part by stacking frequency decoupling modules. The spatial domain frequency is used to directly separate and extract the crack edge and main body information from the original image to ensure the consistency of the internal crack object and enhance the edge area supervision. Specifically, the following steps are included: Step S1: De-noising the concrete engineering structure image to be identified to obtain a de-noised image.
[0022] Step S2: After downsampling the denoised image, the denoised image is processed by SharedSTEM to obtain the first feature.
[0023] Specifically, the denoised concrete engineering structure image is first downsampled multiple times to make the feature map 1 / 8 of the original, and SharedSTEM uses two convolutional layers to process the downsampled features. After that, each denoised image block is converted into an embedded feature, and the model initializes a grid point T as the prototype of the crack image, as follows: ; n represents the 3×3 neighborhood of the cluster center, Represents the local feature vector or data point considered when each cluster center is initialized, for The weights of each point in the grid act as local cluster centers. Each cluster center contains weighted initialized semantic information in its corresponding region but without specific semantic description. At this time, the model's understanding of the semantic information in the image is relatively simple and vague. Therefore, we perform multi-layer frequency decoupling and clustering updates on these cluster centers.
[0024] The relative position of the grid T is encoded through the convolutional layer, and the position information is added to the features to retain the spatial information of the image. Our goal is to update each cluster center s in the grid T instead of directly updating the initial features. This reduces the number of learnable parameters compared to the parameterized method of updating the entire feature map.
[0025] In positional encoding, the input grid T is mapped to query Q, key K, and value V respectively through three independent linear transformations, usually fully connected layers or 1x1 convolutions, and the query Q, key K, and value V are taken as the first features and passed to subsequent steps.
[0026] Step S3: After the first feature is input into a plurality of frequency decoupling modules for processing, classification is performed to identify whether there are cracks in the concrete engineering structure image.
[0027] The first frequency decoupling module divides the first feature into grids, clusters the grids, and performs position encoding; extracts high-frequency features and low-frequency features of the clustered grids with position encoding information respectively, and fuses them according to the correlation between the low-frequency features and the high-frequency features; and fuses the fused features with the first feature and inputs them into the subsequent frequency decoupling module as a new first feature.
[0028] In this embodiment, the correlation between low-frequency features and high-frequency features is extracted through a frequency domain correlation module; the low-frequency features are extracted through a variable low-frequency domain tuning filter; and the high-frequency features are extracted through a variable high-frequency domain tuning filter.
[0029] Specifically, Figure 2As shown in the figure, a variable low-frequency domain tuning filter strategy is proposed based on the characteristics that the low-frequency components in the image occupy most of the energy and carry the main semantic information. The expression is: ; I represents the interpolation operation, AP () represents adaptive average pooling, Represents the P groups after V is grouped, and low-pass filtering is achieved through adaptive average pooling and grouping: the input grid 𝑇 is divided into multiple groups, and different kernels and strides are controlled to generate dynamic low-pass filters, and adaptive pooling of different sizes is performed respectively, and upsampled to the original size. Finally, the up-sampled results are spliced and activated by ReLU to extract low-frequency information The cutoff frequency is controlled by adjusting the kernel size and stride.
[0030] The variable high-frequency domain tuning filter module uses the query Q and high-frequency characteristics to perform a matrix element-by-element product operation to modulate all frequency components. Its expression is: ; in DW () is a deep convolutional layer. First, the input value V is grouped into Q groups, and then For each group, different convolution layers with different kernels are used to simulate the cutoff frequencies of different high-pass filters to form a list of channel frequency domain information. Subsequently, these frequency domain information are connected into a larger tensor in the channel dimension through tensor concatenation operation , thus globally considering the frequency domain structure of the input value V. The shape rearrangement operation combines The height and width dimensions are adjusted so that Becomes a tensor suitable for weighted addition operations.
[0031] Since different frequency components are distributed in the grid T, this embodiment proposes a frequency domain correlation module to calculate the correlation degree between two different frequency components.
[0032] This module first introduces a fixed-size frequency domain correlation module , used to represent the strength of different frequency components. The original input is decomposed into multi-core signal components through a linear layer, thereby refining the understanding of the spatiotemporal changes of image features.
[0033] The module calculates the key K and value V of the frequency component through a linear layer, and normalizes the key K on the frequency component through a Softmax operation. Querying this module can represent the correspondence between different frequency components, and then selecting important frequency components by querying the frequency domain association module.
[0034] Finally, the outputs of the two variable frequency domain tunable filters are combined to obtain the variable frequency filter output, as shown in the following formula: ; In this embodiment, the output of the frequency domain correlation module It is regarded as a function transfer so that each component of the grid integrates a similarity coefficient, which is used for adaptive fusion of variable low-frequency domain tuning filters and variable high-frequency domain tuning filters.
[0035] Through element-by-element product operation, query Q and high-frequency information The features interact with each other to fine-tune the feature strength of each position and channel. At the same time, the low-cost addition ADD is used to replace the high-computation concatenation Concat to fuse the information of low-frequency and high-frequency tuning under the same number of channels. The obtained features are input into DA-FFN. This mechanism optimizes the correlation between the query and high-frequency features and suppresses the high-frequency noise inside the object.
[0036] Finally, the first feature processed by SharedSTEM is restored to the original spatial feature F through bilinear interpolation, and finally the high-frequency and low-frequency fused features are fused with the original spatial features through a set of linear projection layers. The updated feature Fi′ contains rich pixel semantic information from the original feature F, as well as the prototype semantics collected by the cluster center T(s). This method is conducive to training convergence and gradient flow. By stacking four such frequency decoupling modules and aggregating semantics around cluster centers, the crack segmentation network achieves accurate crack segmentation. The pixel features and updated grid features are input to the classification layer, and the output is used for pixel-level classification.
[0037] Example 2 Based on Example 1, this embodiment provides a multi-weather integrated restoration module for denoising the concrete engineering structure image to be identified.
[0038] Specifically, given a clean image, the corresponding degraded image is first generated through AdverseGAN to construct a comparison dataset of degraded and clean images. Subsequently, the generated degraded image is input into the restoration network to generate a consistent restored image that is independent of the degradation. Finally, the restoration result is input into the discriminator to determine the type of degradation experienced by the restored image before restoration. The entire process design is concise and clear. The test phase only relies on the restoration network, without the need for additional reasoning steps or model modifications, thus ensuring efficiency and practicality.
[0039] In this embodiment, the inverse generative adversarial network AdverseGAN is used to effectively generate mixed adverse scenarios to obtain the training data for the first stage. Subsequently, a diversified multi-weather integrated restoration module and a discriminative learning method are used to train the integrated restoration network. Specifically, given a degraded image, a 3 × 3 convolution is first used to generate shallow feature embeddings. Next, these shallow features are then processed by the four-layer attention integrated iterative restoration module AIIR provided in this embodiment, and after four layers of AIIR processing, deep features are obtained.
[0040] In this embodiment, each AIIR consists of several multi-scale spatial attention modules MSA and a feature pyramid FFN, where MSA extracts long- and short-range context information by introducing a self-attention branch in parallel with the convolution branch. The attention branch mainly enhances the model's attention to weather style, and uses 3×3 convolution and GeLU activation functions to extract the main features of the image. AIIR uses multiple MSAs to integrate the results of different heads and perform 1×1 convolution superposition. Local related information is the key to image restoration. FFN is used to integrate local and global information to obtain the restored crack image. The formula is as follows: ; ; In the formula, It represents the preliminary denoising result obtained by concatenating the results of several MSA modules; then the number of channels is restored through FFN and 1×1 convolution to obtain the final denoised image; Cat Represents a splicing operation; Indicates i Results of the MSA modules; Indicates i Characteristics of the group; Represents feature pyramid operation; Represents a 1×1 convolution operation.
[0041] In this embodiment, the features extracted by each AIIR module are processed by a feature mapping network. The feature mapping network locates the reconstruction vector in the Codebook by learning the projection of the encoded features to the corresponding clean embedding, provides additional auxiliary visual atoms for the restoration process, decouples the image style and the crack style, and increases the distance between the two features.
[0042] Example 3 This example effectively verifies the method of Example 1 on the Crack500, Crack200 and DeepCrack datasets, where CRACK500 contains 3,363 crack images with a resolution of 640×360; CRACK200 contains 4,507 crack images with a resolution of 800×600; and DEEPCRACK contains 1,611 crack images with a resolution of 544×384. These datasets show diverse crack morphologies and multiple challenging factors.
[0043] like Figure 3 As shown, GhostNetV2 is the ghost segmentation network V2, STDC represents the short-term densely connected network, SeaFormer is the squeeze-enhanced axial transformer, BiSeNetV2 is the bilateral segmentation network V2, SegFormer represents a simple and efficient unified Transformers semantic segmentation framework, AFFormer is the adaptive frequency transformer, TopFormer is the top forming transformer, and PP-LITESEG represents a novel lightweight model based on the Baidu PaddlePaddle platform for real-time semantic segmentation tasks. It can be seen that the method of this embodiment, combined with the discriminative learning strategy, can effectively distinguish various degradation types and significantly improve the model's ability to extract features from low-quality images. In addition, the present invention strengthens the correspondence between frequency domain components and spatial domain pixels, and applies different processing to different frequency components so that the model not only distinguishes between cracks and non-crack elements in the edge area, but also further improves the segmentation effect of the overall crack.
[0044] like Figure 4As shown, HardNet represents the harmonic densely connected network, BiSeNetV2 is the bilateral segmentation network V2, STDC represents the short-term densely connected network, SegFormer represents a simple and efficient unified Transformers semantic segmentation framework, GhostNetV2 is the ghost segmentation network V2, PP-LITESEG represents a novel lightweight model based on the Baidu PaddlePaddle platform for real-time semantic segmentation tasks, DDRNet represents the deep dual-resolution network for real-time accurate semantic segmentation, TopFormer is the top forming transformer, SeaFormer is the squeeze-enhanced axial transformer, and AFFormer is the adaptive frequency transformer. It can be seen that our model outperforms the suboptimal SegForemr on the Crack500 dataset, reducing the computational cost by 68.0%, reducing the parameter count by 28.9%, and increasing the mIoU by 0.16%. On the DeepCrack dataset, our model achieves the highest mIoU and the second-best precision and recall. Compared with the precision-excellent TopFormer, our model reduces 37.1% of parameters. Although our model is not optimal in terms of parameters and FLOPs, results show that, for example, GhostNetV2 and RHACrackNet, which aggressively prioritize low computational cost and number of parameters but suffer significant penalties as a result, our model achieves the best segmentation performance at minimal computational cost.
[0045] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is more specific and detailed, but it cannot be understood as limiting the scope of the present invention. As long as there is no contradiction in the combination of these technical features, they should be considered as within the scope of this specification.
[0046] It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these modifications and improvements all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.
Claims
1. A frequency decoupled multi-weather scenario crack identification method, characterized by: The following steps are involved: Step S1: De-noising the concrete engineering structure image to be identified to obtain a de-noised image; Step S2: downsampling the denoised image and processing it through SharedSTEM to obtain a first feature; Step S3: inputting the first feature into a plurality of frequency decoupling modules for processing and then classifying to identify whether there are cracks in the concrete engineering structure image; The first frequency decoupling module divides the first feature into grids, clusters the grids, and performs position encoding; extracts high-frequency features and low-frequency features of the clustered grids with position encoding information respectively, and fuses them according to the correlation between the low-frequency features and the high-frequency features; and fuses the fused features with the first feature and inputs them into the subsequent frequency decoupling module as a new first feature.
2. The frequency-decoupled multi-weather scenario crack identification method according to claim 1 is characterized in that: In step S1, the concrete engineering structure image to be identified is denoised by using a multi-weather integrated restoration module; When the multi-weather integrated restoration module is trained, the noise-free crack image is converted into a low-quality crack image through the inverse generative adversarial network AdverseGAN, and the low-quality crack image is denoised through several AIIR modules, and the denoising effect is judged for training optimization; The AIIR module includes several multi-scale spatial attention modules MSA and a feature pyramid FFN. Several MSAs are used to integrate the results of different heads, and after 1×1 convolution superposition, the results are input into FFN to integrate local and global information to obtain the restored crack image.
3. The frequency-decoupled multi-weather scenario crack identification method according to claim 2 is characterized by: The expression of the AIIR module is: ; ; In the formula, It represents the preliminary denoising result obtained by concatenating the results of several MSA modules; Cat Represents a splicing operation; Indicates i Results of the MSA modules; Indicates i Characteristics of the group; Represents feature pyramid operation; Represents a 1×1 convolution operation.
4. The frequency-decoupled multi-weather scenario crack identification method according to claim 3 is characterized by: Each two AIIR modules are processed through a feature mapping network to decouple the image style and the crack style and increase the distance between the two features.
5. The frequency decoupled multi-weather scenario crack identification method according to claim 1 is characterized by: The low-frequency features are extracted by the following method: the cluster grid is divided into multiple groups, adaptive pooling of different sizes is performed on each group, and up-sampled to the original size, and the up-sampled results are spliced and activated by ReLU to extract the low-frequency features.
6. The frequency decoupled multi-weather scenario crack identification method according to claim 5 is characterized by: The high-frequency features are extracted by the following method: the clustering grid is divided into multiple groups, and for each group, a convolution layer with different kernels is used to simulate the cutoff frequencies of different high-pass filters to form a channel frequency domain information list, and the channel frequency domain information list is connected in the channel dimension through a tensor splicing operation to obtain the high-frequency features.
7. The frequency-decoupled multi-weather scenario crack identification method according to claim 6 is characterized by: The expression for fusion according to the correlation between low-frequency features and high-frequency features for: ; In the formula, q represents the query vector; Represents high-frequency features; Represents low-frequency features; Frequency Component i and frequency components j The correlation coefficient between .
8. The frequency decoupled multi-weather scenario crack identification method according to claim 1 is characterized by: The features after the fusion of high-frequency features and low-frequency features are input into the recurrent neural network DA-FFN based on the dual-stage attention mechanism for processing before subsequent operations.
9. The frequency decoupled multi-weather scenario crack identification method according to claim 1, characterized in that: The method uses the loss function L To train: ; In the formula, is the i-th element of the actual label vector; is the i-th element of the probability distribution predicted by the model; N represents the number of samples.
10. A frequency decoupled multi-weather scenario crack identification system, characterized in that: include: One or more processors and memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 9.
Citation Information
Patent Citations
Remote sensing image fusion processing method based on non-subsampled Laplacian pyramid and bi-dimensional empirical mode decomposition (BEMD)
CN102622730A
Power communication network fault diagnosis method and system based on machine learning algorithm
CN114757379A
Quartz glass detection method, device and equipment based on frequency spectrum and medium
CN116645365A
Bridge concrete crack detection method under complex background based on deep learning
CN116823800A
Power distribution network net load prediction method and system considering multiple time-space correlation
CN118411062A
Cited By
Semi-supervised segmentation method combining double-segmentation-head frequency decoupling learning and entropy change pseudo-label screening
CN120635457A