Multi-mode crack detection method, medium and equipment

By using gradient resolution conflict method in multimodal crack detection to deal with the conflict problem between strength and distance crack characteristics, the gradient conflict problem is solved, and the stability and detection accuracy of the model are improved.

CN120219342AActive Publication Date: 2025-06-27ANHUI UNIV

Patent Information

Application Number
CN202510306474.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

In multimodal crack detection, the introduction of multiple single-modal loss functions may lead to gradient conflicts, affecting the optimization process and performance of the model.

Method used

The gradient resolution conflict method is used to deal with the conflict problem between intensity and distance crack characteristics, and ensure the consistency of gradient direction by updating the parameters of the encoder and decoder.

Benefits of technology

It alleviates the problem of inconsistent gradient direction, improves the training stability and generalization ability of the model, and improves the accuracy and robustness of crack detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219342A_ABST
    Figure CN120219342A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode crack detection method, a medium and equipment. In the method, multiple modes comprise an intellectual mode and a range mode. The method comprises the following steps: extracting an intensity crack feature, simultaneously extracting a distance crack feature range, and introducing two groups of decoupling adapter modules to finely adjust an SAM2 encoder; inputting the strength crack feature intensity, the distance crack feature range and fusion features of the strength crack feature and the distance crack feature range into an SAM2 decoder; and a gradient conflict solving method is adopted to relieve gradient conflicts, so that three updated gradients are obtained. Finally, all updated gradients are added to generate a final gradient; the updated gradient is used for updating parameters of an intra-range encoder and a range encoder and a decoder, and finally a crack detection image is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a multi-modal crack detection method, medium and device that balance the introduction of single-modal losses by using a gradient conflict resolution method. Background Art

[0002] In recent years, the wide application of deep learning techniques in the field of computer vision, especially convolutional neural networks (CNNs), has greatly improved the performance of image processing tasks. In crack detection, deep learning methods can achieve more accurate crack localization by automatically extracting image features. However, single-modal image feature extraction methods have certain limitations. Especially in crack detection, different-modal image information often provides complementary information. Therefore, it is of great significance to combine multi-modal information for crack detection.

[0003] Traditional multi-modal fusion methods mainly rely on simple feature concatenation or weighted fusion, ignoring the differences in features between different modalities. Especially when dealing with features at different activation levels, it may lead to inaccurate feature learning, thereby affecting the performance of the model.

[0004] To solve these problems, decoupled adapter modules based on deep learning have emerged in recent years. This technology can automatically select effective features in multi-modal inputs and perform different optimizations on low-activation and high-activation features. By adopting different optimization strategies for channels with higher and lower activation values, the decoupled adapter can effectively reduce the risk of overfitting and improve the generalization ability of the model.

[0005] However, in multi-modal crack detection, introducing multiple single-modal loss functions may lead to gradient conflicts. Especially when the directions of the gradients are inconsistent, the optimization process of the model may be disturbed. Therefore, the gradient conflict problem has become a key factor affecting the performance of multi-modal crack detection. Existing methods usually adopt simple gradient accumulation or merging strategies, ignoring the mutual influence between different loss functions, resulting in unstable training processes and even preventing the full learning of features in a certain modality.

[0006] Therefore, solving the gradient conflict problem and improving the stability and accuracy of multi-modal crack detection methods have become an important challenge in current research. By introducing a gradient conflict resolution strategy, the problem of inconsistent gradient directions can be effectively alleviated, making the model more stable during the training process. At the same time, the utilization efficiency of features in each modality can be improved, thereby enhancing the accuracy and robustness of crack detection. Summary of the Invention

[0007] A multi-modal crack detection method proposed by the present invention can solve at least one of the technical problems in the background art.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A multi-modal crack detection method, comprising the following steps:

[0010] S1. Extract the intensity crack feature intensity and the distance crack feature range from the image to be measured, and introduce two groups of decoupled adapter modules to fine-tune the SAM2 encoder;

[0011] S2. Input the intensity crack feature intensity, the distance crack feature range and the fusion feature into the SAM2 decoder to obtain crack detection information;

[0012] S3. Adopt the gradient conflict resolution method to handle the conflict problem between the intensity crack feature intensity and the distance crack feature range, and obtain a new gradient;

[0013] S4. Use the new gradient to update the parameters of the intensity encoder, the range encoder and the decoder, determine the crack detection information, and finally generate a crack detection map.

[0014] Further, the method for extracting the intensity crack feature intensity and the distance crack feature range in step S1 of the present invention includes:

[0015] Obtain the feature representations of the intensity crack feature intensity and the distance crack feature range through a bimodal encoder:

[0016] F R =E R (I R ,θ R )

[0017] F T =E T (I T ,θ T )

[0018] In the formula, F R is the intensity crack feature intensity, E R is the intensty intensity encoder, F T is the distance crack feature range, E T is the range distance encoder;

[0019] First, freeze and share the SAM2 encoder, and at the same time fine-tune the two groups of adapter parameters θ R and θ T, to adapt the SAM model to intensity - range crack detection, an adapter module is usually adopted in practice. The standard adapter adopts a bottleneck structure design:

[0020] Given the input features which contains the parameters of the lower projection layer the GeLU activation function σ(·) and the parameters of the upper projection layer where the intermediate dimension of the bottleneck The calculation formula of the bottleneck structure is:

[0021]

[0022] In the highly activated region after GeLU activation, the contour features of significant targets are highlighted. The foreground adapter selects the top P for proportion of highly activated channels through the TopK operation, and the background adapter selects the bottom P back proportion of low - activated channels:

[0023]

[0024] where TopK(input,k,largest) retains the largest k channels when largest = True and the smallest k channels when False. P for and P back control the retention ratios of foreground / background activations respectively.

[0025] Furthermore, obtaining crack detection information in step S2 of the present invention includes:

[0026] The SAM2 decoder first deeply fuses the input intensity crack features and range crack features, which is achieved through a multi - layer convolutional neural network CNN. Assuming the input single - modality features are intensity crack feature I and distance crack feature R, the deep fusion is expressed as:

[0027] F = CNN(I,R)

[0028] where F is the fused feature;

[0029] To capture global context information through a multi - layer convolutional neural network, the convolutional operation in the neural network is expressed as:

[0030]

[0031] where F in is the input feature map, F out is the output feature map, and K is the convolution kernel;

[0032] Meanwhile, using the attention mechanism, automatically focus on the local area according to the significance of the crack, enhancing the detailed information of the crack edge. Assume that F is the fused feature, and the attention weight A is expressed as:

[0033]

[0034] where Q and K are the query and key matrices, and d k is the dimension of the key vector. The output of the attention mechanism is expressed as:

[0035] F att = A·V

[0036] where V is the value matrix;

[0037] Through this joint processing of multi-modal features, the decoder can further optimize the recognition of the local area based on the global features, accurately locate the position and shape of the crack, and achieve it through weighted fusion:

[0038] F final = α·F global +(1-α)·F local

[0039] where F global is the global feature, F local is the local feature, and α is the weight coefficient;

[0040] To accurately locate the position and shape of the crack, cross-entropy loss or Dice loss will be used. For example, the cross-entropy loss is expressed as:

[0041]

[0042] where y i is the true label, is the predicted probability.

[0043] Furthermore, the method for handling two single-modal conflict problems by using the gradient conflict resolution method in step S3 of the present invention includes:

[0044] According to the chain rule, the gradients of the two encoders are expressed as:

[0045]

[0046] Observing the gradient relationship, the ratio of G R and G T in the vast majority of regions is greater than 1. According to the following parameter update formula:

[0047]

[0048] Conclusion: The parameter update speed of the intensity encoder is significantly faster than that of the range encoder, and the difference is due to G R The gradient value is continuously greater than G T the gradient value, resulting in a larger parameter update amplitude for the intensity encoder in each iteration;

[0049] where η represents the learning rate and t represents the optimization step;

[0050] Use a shared decoder to decode the intensity feature and the range feature to obtain the unimodal predictions P R and P T ;

[0051] P R = D(F R , θ D ), P T = D(F T , θ D )

[0052] Calculate two unimodal losses:

[0053]

[0054] Introduce the auxiliary constraints of the two unimodal losses, and the gradient of the intensity encoder tends to be equal to the gradient of the range encoder;

[0055] In the encoder E R part, the gradient R from the unimodal loss L and the gradient F from the fusion loss L jointly affect the parameter update; in the encoder E T part, the gradient T from the unimodal loss L and the gradient F from the fusion loss L jointly affect the parameter update; in the decoder D part, the gradient from the intensity stream loss the gradient from the range stream loss and the gradient from the fusion stream jointly affect the parameter update.

[0056]

[0057] The cosine angle of different gradients is not always greater than 0, indicating that there are some opposite gradient directions, which are harmful to model training;

[0058]

[0059] Among them, GradDeConflict is a gradient conflict resolution method that adopts the PCGrad strategy. For the given gradient G i ∈G, the current gradient after conflict resolution is initialized to G j ; then, estimate the current gradient and the cosine similarity between another gradient G i ∈G (j≠i);

[0060]

[0061] If the cosine similarity is less than 0, it means that the optimization directions of the two gradients conflict. At this time it will subtract its projection on the normal plane of G j , otherwise it remains unchanged;

[0062]

[0063] Then, eliminate the potential conflicts between and all other gradients in the same way;

[0064] This process will be repeated among all gradients;

[0065] Finally, all the updated gradients are added together to generate the final gradient G:

[0066]

[0067] After eliminating the gradient conflicts in all the encoders and decoders, the parameters of the two encoders and decoders are updated as follows:

[0068]

[0069] Furthermore, the crack detection map generation method in step S4 of the present invention includes:

[0070] Fuse the two features F R and F T and input them into the decoder to generate the prediction result P F , and its calculation process can be expressed as:

[0071] P F = D(P R + F T , θ D )

[0072] Among them, D represents the decoder with parameters θ D ;

[0073] Calculate the prediction result Pf The loss function between the predicted label and the ground truth GT:

[0074]

[0075] wherein, and respectively represent the weighted intersection over union loss function IoU and the weighted binary cross entropy loss function BCE.

[0076] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above method.

[0077] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to execute the steps of the above method.

[0078] As can be seen from the above technical solutions, the present invention proposes a significant crack detection method based on Intensity-Range multi-modal fusion, which has higher detection accuracy and robustness than traditional methods. Through multi-modal feature fusion, Intensity and Range information are effectively combined to enhance the recognition ability of crack regions; the introduction of the decoupled adapter module enables the model to distinguish high and low activation features of different modalities, adaptively optimize feature expressions, and improve detection performance; in addition, the present invention adopts a gradient conflict resolution strategy to alleviate the gradient conflict problem in the multi-modal learning process, and improve the training stability and generalization ability of the model. The present invention can effectively improve the accuracy, stability and adaptability of crack detection in complex scenarios, and is applicable to various engineering detection and intelligent detection systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 is a flowchart of a multi-modal crack detection method of the present invention;

[0080] Figure 2 is a neural network structure diagram of the multi-modal crack detection method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.

[0082] As Figure 1 shown, a multi-modal crack detection method described in this embodiment is executed by a computer device through the following steps,

[0083] S1. Extract the intensity crack feature intensity and the distance crack feature range from the image to be measured;

[0084] S2. Input the intensity crack feature intensity, the distance crack feature range, and the fusion feature into the SAM2 decoder to obtain crack detection information;

[0085] S3. Use the gradient conflict resolution method to handle the conflict between the intensity crack feature intensity and the distance crack feature range to obtain a new gradient;

[0086] S4. Use the new gradient to update the parameters of the intensity and range distance encoders and the decoder, and finally generate a crack detection map.

[0087] The following is a detailed description of each step:

[0088] S1. Extract the intensity crack feature intensity and the distance crack feature range from the image to be measured, and introduce two groups of decoupled adapter modules to fine-tune the SAM2 encoder;

[0089] When extracting the intensity crack feature intensity and the distance crack feature range, there is a problem of indistinguishable high-activation and low-activation learning. Two groups of decoupled adapter modules are introduced to fine-tune the SAM2 encoder. The decoupled adapter modules decouple the learning processes of the intensity modality and the range modality, enabling the model to more effectively process features at different activation levels.

[0090] The methods for extracting the intensity crack feature intensity and the distance crack feature range include:

[0091] Obtain the feature representation through a dual-modal encoder:

[0092] F R = E R (I R , θ R )

[0093] F T = E T (I T , θ T )

[0094] In the formula, F R is the intensity crack feature intensity, E R is the intensty intensity encoder, F T is the distance crack feature range, and E T is the range distance encoder;

[0095] According to conventional practice, we freeze and share the SAM2 encoder while fine-tuning the two sets of adapter parameters θ R (intensity branch) and θ T (range branch). To adapt the SAM model to intensity-range crack detection, adapter modules are usually adopted in practice. The standard adapter is designed with a bottleneck structure:

[0096] Given the input feature which contains the parameters of the lower projection layer the GeLU activation function σ(·) and the parameters of the upper projection layer where the bottleneck intermediate dimension The calculation formula of this structure is:

[0097]

[0098] However, traditional adapters use the same gradient optimization for both highly activated neurons and lowly activated neurons. The highly activated region after GeLU activation can highlight the contour features of significant targets, while the lowly activated region mainly corresponds to task-irrelevant background information. Over-optimization of the lowly activated region is likely to introduce noise and increase the risk of overfitting. Therefore, we propose to decouple the adapter into a foreground adapter and a background adapter.

[0099] Specifically, the foreground adapter selects the top P for proportion of highly activated channels through the TopK operation, and the background adapter selects the bottom P back proportion of lowly activated channels:

[0100]

[0101] where TopK(input,k,largest) retains the largest k channels when largest = True and the smallest k channels when largest = False. P for and P back control the retention ratios of foreground / background activations respectively.

[0102] S2. Input the intensity crack feature intensity, the distance crack feature range, and the fusion feature into the SAM2 decoder to obtain crack detection information;

[0103] The SAM2 decoder first performs deep fusion on the input single-modal features (including the intensity crack feature and the range crack feature), usually implemented through a multi-layer convolutional neural network (CNN). Assuming the input single-modal features are I (intensity crack feature) and R (distance crack feature), the deep fusion can be expressed as:

[0104] F = CNN(I, R)

[0105] Where F is the fused feature.

[0106] The global context information is captured by a multi-layer convolutional neural network. The convolution operation in the neural network can be expressed as:

[0107]

[0108] Where F in is the input feature map, F out is the output feature map, and K is the convolution kernel.

[0109] Meanwhile, using the attention mechanism, it automatically focuses on the local area according to the saliency of the crack, enhancing the detailed information of the crack edge. Assuming F is the fused feature, the attention weight A can be expressed as:

[0110]

[0111] Where, Q and K are the query and key matrices, and d k is the dimension of the key vector. The output of the attention mechanism can be expressed as:

[0112] F att = A · V

[0113] Where, V is the value matrix.

[0114] Through this joint processing of multi-modal features, the decoder can further optimize the recognition of the local area based on the global features, accurately locating the position and morphology of the crack. It can be achieved through weighted fusion:

[0115] F final = α · F global + (1 - α) · F local

[0116] Where, F global is the global feature, F local is the local feature, and α is the weight coefficient.

[0117] To accurately locate the position and morphology of the crack, cross-entropy loss or Dice loss etc. may be used. For example, the cross-entropy loss can be expressed as:

[0118]

[0119] Where, y i is the true label, is the predicted probability.

[0120] S3. Use the gradient conflict resolution method to handle the conflict between the intensity crack feature and the range crack feature, and obtain a new gradient.

[0121] Due to the introduction of the losses of two single modalities (intensity crack feature and range crack feature), gradient conflict occurs. The gradient conflict resolution method is adopted to alleviate the gradient conflict, thereby obtaining three updated gradients. Finally, all the updated gradients are added up to generate the final gradient.

[0122] According to the chain rule, the gradients of the two encoders can be expressed as:

[0123]

[0124] Observing the gradient relationship, the ratio of G in the vast majority of regions R and G T is greater than 1. According to the following parameter update formula:

[0125]

[0126] It can be concluded that the parameter update speed of the intensity encoder is significantly faster than that of the range encoder. This difference stems from the fact that the gradient value of G R continually exceeds the gradient value of G T , resulting in a larger parameter update amplitude for the intensity encoder in each iteration.

[0127] Here, η represents the learning rate and t represents the optimization step. The comparison of the parameter update speeds indicates that the intensity modality dominates the training of the entire model, while the range modality is not fully optimized. Therefore, the two modalities are not well utilized, thereby reducing the performance.

[0128] To address this challenge, we introduce single-modal supervision to enhance the optimization of the range modality. Specifically, a shared decoder is used to decode the intensity feature and the range feature to obtain the single-modal predictions P R and P T .

[0129] P R = D(F R , θ D ), P T = D(F T , θ D )

[0130] Therefore, two single-modal losses are calculated:

[0131]

[0132]

[0133] By introducing the auxiliary constraints of two unimodal losses, the gradient of the intensity encoder tends to be equal to the gradient of the range encoder.

[0134] However, the introduction of two unimodal losses leads to gradient conflicts. In the encoder E R section, the gradient R from the unimodal loss L and the gradient F from the fusion loss L jointly affect the parameter update; in the encoder E T section, the gradient T from the unimodal loss L and the gradient F from the fusion loss L jointly affect the parameter update; in the decoder D section, the gradient from the intensity flow loss the gradient from the range flow loss and the gradient from the fusion flow jointly affect the parameter update.

[0135]

[0136] The cosine angles of different gradients are not always greater than 0, indicating that there are some opposite gradient directions, which are harmful to model training. In this case, gradient conflict resolution is applied to alleviate these gradient conflicts, resulting in three updated gradients. By this method, the model can avoid conflicts between gradient directions during training, thus improving the stability and efficiency of optimization.

[0137]

[0138] Among them, GradDeConflict is a gradient conflict resolution method adopting the PCGrad strategy. Specifically, for a given gradient G i ∈G, the current gradient after conflict resolution is initialized as G j . Then, the cosine similarity between the current gradient and another gradient G i ∈G (j≠i) is estimated.

[0139]

[0140] If the cosine similarity is less than 0, it means that there is a conflict in the optimization directions of the two gradients. At this time it will subtract its projection on Gj Projection on the normal plane, otherwise remain unchanged.

[0141]

[0142] Then, eliminate in the same way Potential conflicts with all other gradients.

[0143] This process is repeated among all gradients. Finally, all the updated gradients are added together to generate the final gradient G:

[0144]

[0145] After eliminating the gradient conflicts in all the encoders and decoders, the parameters of the two encoders and decoders are updated as follows:

[0146]

[0147] S4. Use the new gradient to update the parameters of the intensity encoder, range encoder, and decoder, determine the crack detection information, and finally generate the crack detection map;

[0148] Fuse the two features F R and F T and input them into the decoder to generate the prediction result P F . Its calculation process can be expressed as:

[0149] P F = D(P R + F T , θ D )

[0150] where D represents the decoder with parameter θ D .

[0151] Furthermore, calculate the loss function between the prediction result P F and the true label GT:

[0152]

[0153] In the formula, and represent the weighted intersection over union (IoU) loss function and the weighted binary cross - entropy (BCE) loss function respectively. Through the weight adjustment mechanism, these two loss functions can effectively improve the model's learning ability for difficult example samples. Among them, the weighted intersection over union loss strengthens the boundary localization accuracy, and the weighted cross - entropy loss enhances the robustness of pixel - level classification.

[0154] The following is illustrated by a specific example:

[0155] As Figure 2 shown, the multi-modal crack detection method described in this embodiment is verified on the FIND dataset. For the FIND dataset, the same settings as those in the paper "A Dual-Stream End-to-end Pavement Crack Segmentation Network based on Transformer and Multi-modal Fusion" are adopted, including 2000 pairs of intensity-range images in the training set and 500 pairs of intensity-range images in the test set.

[0156] In the training and testing phases, the input intensity-range images are resized to 512*512. The AdamW optimizer is selected for model training, with an initial learning rate of 1e-3, a batch size of 4, SAM2 pre-training parameters and the default settings of PyTorch are used, and the GPU used is NVIDIA GTX 3090. The number of training epochs in the warm-up phase and the joint training phase is 45 epochs each.

[0157] The proposed method is compared with 7 RGBD detection methods, namely Crack Transformer, SegFormer, Swin-UperNet, VGG19-FCN, UNet-FCN, DSEN-uno, and CrackFusionNet. The results are shown in Table 1.

[0158]

[0159] The experimental results in Table 1 are as follows:

[0160] [1]W. Wang, Q. Lai, H. Fu, J. Shen, H. Ling, and R. Yang, “Salient object detection in the deep learning era: An in-depth survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 6, pp. 3239 - 3259, 2021.

[0161] [2]Z. Tu, Y. Ma, Z. Li, C. Li, J. Xu, and Y. Liu, “RGBT Salient Object Detection: A Large-Scale Dataset and Benchmark,” IEEE Transactions on Multimedia, vol. 25, pp. 4163 - 4176, 2022.

[0162] [3]A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al., “Segment Anything,” in Proceedings of the IEEE / CVF International Conference on Computer Vision, 2023, pp. 4015 - 4026.

[0163] [4]N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rä dle, C. Rolland, L. Gustafson et al., “SAM 2: Segment Anything in Images and Videos,” arXiv preprint arXiv:2408.00714, 2024.

[0164] [5]S. Gao, P. Zhang, T. Yan, and H. Lu, “Multi-scale and Detail-enhanced Segment Anything Model for Salient Object Detection,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 9894 - 9903.

[0165] [6] P. Zhang, T. Yan, Y. Liu, and H. Lu, “Fantastic animals and where to find them: Segment any marine animal with dual SAM,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 2578-2587.

[0166] [7] S. Lian, Z. Zhang, H. Li, W. Li, L. T. Yang, S. Kwong, and R. Cong, “Diving into underwater: Segment anything model guided underwater salient instance segmentation and a large-scale dataset,” in Forty-first International Conference on Machine Learning, 2024, pp. 29545-29559.

[0167] [8] K. Wang, D. Lin, C. Li, Z. Tu, and B. Luo, “Adapting segment anything model to multi-modal salient object detection with semantic feature fusion guidance,” arXiv preprint arXiv:2408.15063, 2024.

[0168] [9] Y. Liu, P. Wu, M. Wang, and J. Liu, “CPAL: Cross-prompting adapter with LoRAs for RGB+X semantic segmentation,” IEEE Transactions on Circuits and Systems for Video Technology, pp. 1-14, 2025.

[0169]

[10] S. Chen, C. Ge, Z. Tong, J. Wang, Y. Song, J. Wang, and P. Luo, “AdaptFormer: Adapting vision transformers for scalable visual recognition,” Advances in Neural Information Processing Systems, vol. 35, pp. 16664 - 16678, 2022.

[0170]

[11] T. Chen, L. Zhu, C. Deng, R. Cao, Y. Wang, S. Zhang, Z. Li, L. Sun, Y. Zang, and P. Mao, “SAM-Adapter: Adapting segment anything in underperformed scenes,” in Proceedings of the IEEE / CVF International Conference on Computer Vision, 2023, pp. 3367 - 3375.

[0171]

[12] T. Chen, A. Lu, L. Zhu, C. Ding, C. Yu, D. Ji, Z. Li, L. Sun, P. Mao, and Y. Zang, “SAM2-Adapter: Evaluating&adapting segment anything 2 in downstream tasks: Camouflage, shadow, medical image segmentation, and more,” arXiv preprint arXiv:2408.04579, 2024.

[0172]

[13] X. Peng, Y. Wei, A. Deng, D. Wang, and D. Hu, “Balanced multimodal learning via on-the-fly gradient modulation,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 8238 - 8247.

[0173]

[14] Y. Wei and D. Hu, “MMPareto: Boosting multimodal learning with innocent unimodal assistance,” in Forty-first International Conference on Machine Learning, 2024, pp. 52559 - 52572.

[0174]

[15] T. Li, Z. Wen, Y. Li, and T. S. Lee, “Emergence of shape bias in convolutional neural networks through activation sparsity,” Advances in Neural Information Processing Systems, vol. 36, pp. 71755 - 71766, 2023.

[0175]

[16] Y. Pang, X. Zhao, L. Zhang, and H. Lu, “CAVER: Cross-modal view-mixed transformer for bi-modal salient object detection,” IEEE Transactions on Image Processing, vol. 32, pp. 892 - 904, 2023.

[0176] As can be seen from Table 1, the method of the present invention achieves the best results in the evaluation indexes of IoU, F1, Precision, and Recall, proving the effectiveness of the method.

[0177] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.

[0178] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above method.

[0179] In another embodiment provided by the present application, a computer program product containing instructions is also provided, which when running on a computer causes the computer to execute any one of the multimodal crack detection methods in the above embodiments.

[0180] It is understandable that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention. For the explanations, examples, and beneficial effects of related content, reference can be made to the corresponding parts in the above methods.

[0181] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state disk (SSD)).

[0182] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.

[0183] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0184] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-modal crack detection method, characterized in that: The following steps are involved: S1, extracting intensity crack feature intensity and distance crack feature range from the image to be tested; S2, input the intensity crack feature intensity, the distance crack feature range and the fusion feature into the SAM2 decoder to obtain crack detection information; S3, using the gradient conflict resolution method to deal with the conflict between the intensity crack feature intensity and the distance crack feature range, and obtain a new gradient; S4. Use the new gradient to update the parameters of the intensity encoder, range encoder, and decoder, determine the crack detection information, and finally generate a crack detection map.

2. A multi-modal crack detection method according to claim 1, characterized in that: The method for extracting the intensity crack feature intensity and the distance crack feature range in step S1 includes: The intensity crack feature and the range crack feature are obtained through the dual-mode encoder: F R =E R (I R ,i R ) F T =E T (I T ,i T ) In the formula, F R is the intensity crack characteristic, E R F is the intensity encoder, T is the crack characteristic range, E T It is the range distance encoder; First, freeze and share the SAM2 encoder and fine-tune the two sets of adapter parameters θ R and θ T In order to adapt the SAM model to intensity-range crack detection, an adapter module is usually used in practice. The standard adapter adopts a bottleneck structure design: Given input features It contains the lower projection layer parameters GeLU activation function σ(·) and up-projection layer parameters The bottleneck intermediate dimension The calculation formula of the bottleneck structure is: The highly activated area after GeLU activation highlights the contour features of the salient target, and the foreground adapter selects the top P through the TopK operation. for Ratio of highly activated channels to background adapter selection after P back Proportional low activation channel: TopK(input,k,largest) retains the largest k channels when largest=True, and the smallest k channels when False. for and P back Controls the proportion of foreground / background activations retained respectively.

3. A multi-modal crack detection method according to claim 1, characterized in that: Acquiring crack detection information in step S2 includes: The SAM2 decoder first performs deep fusion of the input intensity crack features and range crack features through a multi-layer convolutional neural network (CNN). Assuming that the input unimodal features are the intensity crack feature I and the range crack feature R, the deep fusion is expressed as: F=CNN(I,R) Where F is the fused feature; The global context information is captured by a multi-layer convolutional neural network. The convolution operation in the neural network is expressed as: where F in is the input feature map, F out is the output feature map, K is the convolution kernel; At the same time, the attention mechanism is used to automatically focus on the local area according to the significance of the crack and enhance the detail information of the crack edge. Assuming that F is the fused feature, the attention weight A is expressed as: where Q and K are the query and key matrices, d k is the dimension of the key vector, and the output of the attention mechanism is expressed as: F att =A·V Where V is the value matrix; Through the joint processing of this multimodal feature, the decoder further optimizes the recognition of local areas based on the global features, accurately locates the position and shape of the crack, and achieves the following through weighted fusion: F final =α·F global +(1-α)·F local Among them, F global is the global feature, F local is a local feature, α is a weight coefficient; Use cross entropy loss or Dice loss to accurately locate the location and shape of the crack, where the cross entropy loss is expressed as: In the formula, y i is the true label, is the predicted probability.

4. A multi-modal crack detection method according to claim 1, characterized in that: The method for using the gradient conflict resolution method to handle the two single-mode conflict problems in step S3 includes: According to the chain rule, the gradients of the two encoders are expressed as: Observe the gradient relationship, G in most areas R and G T If the ratio is greater than 1, update the formula according to the following parameters: Conclusion: The parameter update speed of the intensity encoder is significantly faster than that of the range encoder. The difference comes from G R The gradient value is continuously greater than G T The gradient value causes the intensity encoder to obtain a larger parameter update amplitude in each iteration; Among them, η represents the learning rate and t represents the optimization step; A shared decoder is used to decode the intensity feature and range feature to obtain a unimodal prediction P R and P T ; P R =D(F R ,θ D ),P T =D(F T ,θ D ) Compute two unimodal losses: Two auxiliary constraints of unimodal loss are introduced, and the gradient of the intensity encoder tends to be equal to the gradient of the range encoder; In encoder E R Part, comes from the unimodal loss L R Gradient and the fusion loss L F Gradient Jointly affect parameter update; in encoder E T Part, comes from the unimodal loss L T Gradient and the fusion loss L F Gradient Jointly affect parameter updates; in the decoder D part, the gradient from the intensity flow loss Gradients from range flow loss and the gradient from the fused stream Jointly influence parameter updates; The cosine angles of different gradients are not always greater than 0, which indicates that there are some opposite gradient directions, which is harmful to model training; Among them, GradDeConflict is a gradient conflict resolution method using the PCGrad strategy. For a given gradient G i ∈G, the current gradient after conflict resolution Initialize to G j ; Then, estimate the current gradient With another gradient G i ∈G(j≠i) cosine similarity; If the cosine similarity is less than 0, it means that there is a conflict in the optimization directions of the two gradients. will subtract it from G j The projection onto the normal plane of , otherwise it remains unchanged; Then, eliminate the Potential conflicts with all other gradients; This process is repeated between all gradients; Finally, all updated gradients are added together to generate the final gradient G: After eliminating all gradient conflicts in the encoder and decoder, the parameters of the two encoders and decoders are updated as follows:

5. A multi-modal crack detection method according to claim 1, characterized in that: The method for generating a crack detection map in step S4 includes: The two features F R and F T The fusion is input into the decoder to generate the prediction result P F , and its calculation process can be expressed as: P F =D(P R +F T ,θ D ) Where D represents the parameter θ D Decoder; Calculate the prediction result P F The loss function between the real label GT: In the formula, and They represent the weighted intersection-over-union loss function IoU and the weighted binary cross entropy loss function BCE respectively.

6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 5.

7. A computer device comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Strip steel surface defect detection method based on improved Yolov5s network

    CN117350948A

  • New energy battery shell surface tiny defect detection method

    CN118314085A

  • Pavement crack detection method for different environments

    CN118967638A

  • Building outer wall crack detection method and system, medium and program product

    CN119006454A

  • Crack detector, crack detection method, and computer program

    JP2019066264A

Cited By

  • Controllable image edge detection method and system

    CN121391912A

  • A controllable image edge detection method and system

    CN121391912B