Long-distance crack detection method and system based on sparse query mechanism

Through the multimodal information fusion method based on the sparse query mechanism, combined with the advantages of two-dimensional image mode and three-dimensional point cloud mode, the problem of insufficient single mode information in long-distance crack detection is solved, and high-precision and robust detection effects are achieved.

CN119919637APending Publication Date: 2025-05-02哈尔滨工业大学人工智能研究院有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411983929.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing long-distance crack detection model relies on single mode information and cannot fully utilize the complementarity of image mode and point cloud mode, resulting in insufficient detection accuracy and robustness.

Method used

The multimodal information fusion method based on the sparse query mechanism is adopted, and the two-dimensional image mode and three-dimensional point cloud mode are processed separately through parallel modal independent detection branches, and the unified encoded query data is generated, and the proxy attention mechanism and cross attention mechanism are optimized to finally generate long-distance crack detection results.

Benefits of technology

It significantly improves the accuracy and robustness of long-distance crack detection, solves the problem of insufficient single mode information, enhances feature characterization capabilities and computing efficiency, and meets the needs of real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919637A_ABST
    Figure CN119919637A_ABST
Patent Text Reader

Abstract

The invention provides a long-distance crack detection method and system based on a sparse query mechanism, and belongs to the technical field of computer vision. The method is used for solving many technical problems existing in an existing crack detection method. The method comprises the steps of obtaining a two-dimensional image mode and a three-dimensional point cloud mode of a detection target; performing long-distance crack target preliminary detection on the two-dimensional image mode and the three-dimensional point cloud mode by adopting parallel mode independent detection branches to generate a two-dimensional long-distance crack detection result and a three-dimensional long-distance crack detection result; respectively carrying out sparse query on the two-dimensional and three-dimensional long-distance crack detection results to generate image query data and point cloud query data, and carrying out unified coding on the image query data and the point cloud query data; and performing long-distance crack target accurate detection on the uniformly coded image query data and point cloud query data to generate a long-distance crack detection result. The method is suitable for target detection research in the field of computer vision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of computer vision, and in particular relates to a long-distance crack detection method based on a sparse query mechanism. Background Art

[0002] Crack target detection technology is an engineering method based on computer vision. It automatically identifies and accurately locates the location and related features of cracks by analyzing structural surface data. This technology extracts the geometric morphology and texture features of cracks and combines traditional image processing algorithms or machine learning models to achieve accurate detection and identification of cracks in complex scenes. It is widely used in health assessment of infrastructure such as dams, bridges, and tunnels, providing support for preventive maintenance, post-disaster repair, and quality control. Compared with traditional manual detection, this technology has significant advantages such as high efficiency, accuracy, non-contact, and automation, meeting the needs of large-scale monitoring and complex environments.

[0003] Traditional crack detection methods mainly rely on edge detection, histogram statistics and other technologies to achieve recognition by extracting local features of cracks from images. For example, Sobel or Canny edge detection operators extract crack edges by analyzing grayscale value changes, histogram statistics extract crack regions through threshold segmentation, region growing algorithms use seed points to expand connected regions to generate crack contours, and morphological processing enhances features through operations such as corrosion and dilation. In addition, Gabor filters, wavelet transforms and local binary patterns (LBP) are often used to extract texture information of cracks. However, these methods are highly dependent on manually set thresholds or prior rules, lack adaptability in complex backgrounds, and are difficult to capture subtle features and irregular shapes of cracks. Especially in long-distance crack detection tasks, traditional methods show significant performance degradation and are difficult to meet the requirements of modern monitoring tasks for high accuracy, robustness and real-time performance.

[0004] In recent years, traditional machine learning algorithms have been initially applied in the field of crack detection, such as crack classification and strength prediction models based on support vector machines (SVM), pattern recognition algorithms based on wavelet transforms, and soft edge detection methods. These methods have improved the adaptability of detection to a certain extent by combining manually designed features with machine learning models. However, such methods still have the problems of high computational complexity and difficulty in parameter tuning. At the same time, they show significant lack of robustness when dealing with complex crack morphology, background interference, and long-distance detection tasks. In addition, these algorithms have limited generalization capabilities in cross-scenario applications and are difficult to meet the needs of crack detection in diverse and dynamic environments.

[0005] With the improvement of computing performance and the development of deep learning technology, crack detection methods based on deep learning have gradually become a research focus. Models such as Convolutional Neural Network (CNN) and Vision Transformer (ViT) can efficiently learn multi-level features of cracks on large-scale data sets, including low-level texture details, high-level semantic information and contextual relationships, significantly improving detection accuracy and robustness. Compared with traditional methods, deep learning automatically extracts features from raw data through an end-to-end training model, reducing manual intervention, showing strong feature expression capabilities and domain adaptability, and providing a new direction for the development of crack detection technology.

[0006] The long-distance crack detection model is a type of target detection model based on deep learning. It achieves accurate classification and three-dimensional spatial position regression of long-distance cracks by extracting modal features. Compared with conventional crack target detection, long-distance crack detection faces challenges such as small pixel ratio of long-distance crack targets, insufficient feature expression, susceptibility to image noise interference, and missing edge and texture information, which significantly increases the complexity of detection. Existing methods usually adopt a multi-level feature pyramid network (FPN) structure to effectively capture the fine-grained features of small targets such as long-distance cracks with a multi-scale fusion strategy, and use contextual information to enhance the detection accuracy and robustness in complex backgrounds. In addition, recent research has introduced the attention mechanism, which further improves the model's attention and representation ability to the target area through global context modeling and dynamic allocation of attention weights, thereby significantly improving the accuracy and semantic understanding of long-distance crack detection.

[0007] Limitations of existing methods: Current long-distance crack detection models mostly rely on single-modal information and fail to fully utilize the complementarity between different modalities. Although the image modality provides rich texture and semantic information, it lacks depth data and cannot accurately locate the three-dimensional spatial position of the target; the point cloud modality provides accurate spatial geometric information, but due to data sparsity, it lacks sufficient semantic expression capabilities, especially when detecting long-distance targets. Therefore, single-modal detection methods are difficult to achieve high-precision and robust detection effects in complex scenarios.

[0008] Modern hierarchical models usually adopt a feature fusion strategy to combine deep coarse-grained features with shallow high-resolution features through upsampling to achieve comprehensive utilization of multi-scale features. However, the fused features often show intra-class inconsistency, and the feature values ​​inside the object vary greatly. In addition, the boundaries of the fused features are often blurred due to the lack of high-frequency information, resulting in boundary offset problems, which in turn affects the accuracy and stability of detection. This problem is particularly prominent in the long-distance crack detection task, especially for small-target long-distance cracks, whose feature expression is limited. These defects further reduce the detection performance and robustness of the model.

[0009] The model based on the attention mechanism shows excellent global modeling ability in long-distance crack target detection and can effectively capture long-distance dependencies. However, the traditional Softmax attention mechanism is difficult to be efficiently applied in practical scenarios due to its quadratic complexity computational overhead. Although the linear attention mechanism significantly reduces the computational complexity by introducing linear decomposition, this improvement usually comes at the expense of weakening the global information interaction ability, thereby limiting the model's feature expression and semantic modeling capabilities in complex scenarios, making it difficult to meet the requirements of long-distance crack detection for efficiency and robustness. Summary of the invention

[0010] The present invention provides a long-distance crack detection method based on a sparse query mechanism, which is used to solve the technical problems mentioned in the above background technology.

[0011] To achieve the above object, the present invention provides the following solutions:

[0012] The present invention provides a long-distance crack detection method based on a sparse query mechanism, the method comprising the following steps:

[0013] Step S1: Acquire the two-dimensional image modality and three-dimensional point cloud modality of the detection target;

[0014] Step S2: using parallel modality-independent detection branches to perform preliminary detection of long-distance crack targets on the two-dimensional image modality and the three-dimensional point cloud modality, respectively, to generate two-dimensional long-distance crack detection results and three-dimensional long-distance crack detection results;

[0015] Step S3: performing sparse query on the two-dimensional and three-dimensional long-distance crack detection results to generate image query data and point cloud query data, and uniformly encoding the image query data and the point cloud query data;

[0016] Step S4: Perform precise detection of long-distance crack targets on the uniformly encoded image query data and point cloud query data to generate long-distance crack detection results.

[0017] Furthermore, there is a preferred embodiment in which the above-mentioned parallel modality-independent detection branch includes a YOLOv11 image detection module with enhanced multi-level feature fusion and a fully sparse voxel FSDv2 point cloud detection module.

[0018] Furthermore, there is another preferred embodiment, the above step S2 is specifically as follows:

[0019] Step S21: adding an enhanced feature fusion module to the neck network part of YOLOv11 to obtain a YOLOv11 image detection module with enhanced multi-level feature fusion;

[0020] Step S22: using the YOLOv11 image detection module with enhanced multi-level feature fusion to process the two-dimensional image modality and generate a two-dimensional long-distance crack detection result;

[0021] Step S23: Use the fully sparse voxel FSDv2 point cloud detection module to process the three-dimensional point cloud modality to generate a three-dimensional long-distance crack detection result.

[0022] Furthermore, in a preferred embodiment, the enhanced feature fusion module includes an adaptive low-pass filter generator, an offset generator and an adaptive high-pass filter generator.

[0023] Furthermore, there is another preferred embodiment, the above step S3 is specifically as follows:

[0024] Step S31: Use an uncertainty-aware image query generator to perform sparse query on the two-dimensional long-distance crack detection results.

[0025] Generate image query data including content features and location features;

[0026] Step S32: taking the center point of the three-dimensional long-distance crack detection result as the position feature part of the point cloud query;

[0027] Step S33: combining the appearance features and the geometric features to obtain the content feature part of the point cloud query;

[0028] Step S34: retain the content features in the above image query and the content features in the point cloud query, and perform encoding conversion on the position features in the image query and the position features in the point cloud query to complete unified encoding.

[0029] Furthermore, there is another preferred embodiment, the above step S4 is specifically as follows:

[0030] Step S41: using a proxy attention mechanism to process the uniformly encoded query data, capturing dependencies between queries, and obtaining optimized queries;

[0031] Step S42: using a cross attention mechanism to process the above optimized query data and external modal features, aggregating useful features, and obtaining an optimized query again;

[0032] Step S43: using a query calibration mechanism to calibrate the image query in the query data;

[0033] Step S44: After calibration, the query content features and anchor point positions are processed by the classification head and the regression head to generate long-distance crack detection results.

[0034] Furthermore, there is a preferred embodiment, in which the method for constructing the above-mentioned proxy attention mechanism is: introducing a set of additional proxy tags A into the traditional attention module to obtain the proxy attention mechanism.

[0035] The long-distance crack detection method based on a sparse query mechanism described in the present invention can be fully implemented using computer software. Therefore, correspondingly, the present invention also provides a long-distance crack detection system based on a sparse query mechanism, and the system includes a storage device, which is used to execute the above-mentioned long-distance crack detection method based on a sparse query mechanism.

[0036] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the long-distance crack detection method based on a sparse query mechanism described in any one of the above preferred embodiments is executed.

[0037] The present invention also provides a computer device, which includes a memory and a processor, wherein the memory stores a computer program. When the processor runs the computer program stored in the memory, the processor implements a long-distance crack detection method based on a sparse query mechanism as described in any one of the above preferred embodiments.

[0038] The beneficial effects of the present invention are:

[0039] 1. Aiming at the problem that the accuracy of long-distance crack detection is limited due to insufficient single modal information, the present invention proposes a detection method of multimodal information fusion. This method makes full use of the complementary characteristics of the rich semantic information of the image modality and the precise spatial position information of the point cloud modality, effectively improving the detection accuracy. By designing a parallel independent branch architecture, the independent prior information of each modality is extracted and retained respectively, ensuring the integrity and complementarity of the modal features, providing a more comprehensive feature representation and enhanced robustness for long-distance crack detection.

[0040] 2. To address the intra-category inconsistency and boundary offset problems generated in the standard multi-level feature fusion process, the present invention introduces a spatially variable low-pass filter in the feature fusion stage to adaptively smooth high-level features; replaces inconsistent features by resampling features with consistent adjacent categories; and introduces an adaptive high-pass filter to enhance the high-frequency information of low-level features, thereby significantly improving the category consistency of the fused features and refining the boundary representation of long-distance crack targets.

[0041] 3. In view of the high computational overhead of the existing attention mechanism or the shortcomings of computational optimization at the expense of limited representation capabilities, the present invention introduces additional proxy tags into the traditional attention mechanism. The proxy tags aggregate key-value information and broadcast the results back to the query to achieve efficient information interaction. While retaining the global context modeling capability, the computational efficiency is greatly improved, thereby significantly enhancing the performance and robustness of long-distance crack detection.

[0042] 4. In view of the limitations of existing methods, such as slow reasoning speed and difficulty in meeting the needs of actual scenarios, the present invention is committed to lightweight network design, which significantly improves the reasoning efficiency of long-distance crack target detection, while achieving a good balance between detection accuracy and reasoning speed, meeting the application requirements of real-time detection.

[0043] The present invention is applicable to the research direction of target detection in the field of computer vision. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0045] Figure 1 It is an overall block diagram of a long-distance crack target detection method based on a sparse query mechanism described in an embodiment of the present invention;

[0046] Figure 2 It is the improved YOLO v11 network described in the implementation mode of the present invention;

[0047] Figure 3 It is the FSDv2 network described in the implementation manner of the present invention. DETAILED DESCRIPTION

[0048] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0049] The specific embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.

[0050] Implementation method 1, see Figure 1 This embodiment describes a long-distance crack detection method based on a sparse query mechanism. By integrating the complementary advantages of the two modalities of image and point cloud, an efficient multimodal feature fusion framework is designed to significantly improve the detection accuracy and robustness. The method includes the following three stages: the preliminary detection stage of long-distance crack targets, the sparse query generation and unified modality encoding stage, and the precise detection stage of long-distance crack targets. The method framework is as follows Figure 1 As shown, it can effectively adapt to the crack target detection tasks in long distances and complex scenes.

[0051] The detection method specifically includes the following steps:

[0052] Step S1: Acquire the two-dimensional image modality and three-dimensional point cloud modality of the detection target;

[0053] Step S2: using parallel modality-independent detection branches to perform preliminary detection of long-distance crack targets on the two-dimensional image modality and the three-dimensional point cloud modality, respectively, to generate two-dimensional long-distance crack detection results and three-dimensional long-distance crack detection results;

[0054] Step S3: performing sparse query on the two-dimensional and three-dimensional long-distance crack detection results to generate image query data and point cloud query data, and uniformly encoding the image query data and the point cloud query data;

[0055] Step S4: Perform precise detection of long-distance crack targets on the uniformly encoded image query data and point cloud query data to generate long-distance crack detection results.

[0056] In practical application, if Figure 1As shown, firstly, a data set based on high-resolution two-dimensional image and point cloud data is constructed to obtain the two-dimensional image modality and three-dimensional point cloud modality of the detection target; secondly, in order to fully retain the independent prior information of each modality, this implementation method designs a parallel modality-independent detection branch, namely, the YOLOv11 image detection module and the fully sparse voxel FSDv2 point cloud detection module that enhance the multi-level feature fusion; the parallel modality-independent detection branch is used to perform preliminary detection of long-distance crack targets on the two-dimensional image modality and the three-dimensional point cloud modality, respectively, to generate two-dimensional long-distance crack detection results and three-dimensional long-distance crack detection results; since the detection results of the two-dimensional image and three-dimensional point cloud modalities provide key information for the recognition of long-distance crack targets, but their representation forms are essentially different. Therefore, by extracting target-level semantics from the detection results, the semantic information of long-distance crack targets is encoded in the form of sparse target query. The target query consists of two parts: content features and position features, which respectively describe the semantic attributes and spatial distribution of the target, and provide a consistent representation for the fusion of multimodal information; finally, an additional proxy tag is introduced to improve the attention mechanism, which significantly improves the computational efficiency while retaining the global context modeling capability. The uniformly encoded image query data and point cloud query data are used to accurately detect long-distance crack targets and generate long-distance crack detection results.

[0057] Implementation Method 2: See Figure 2 and Figure 3 This embodiment is described. This embodiment is to specifically describe step S2 in the long-distance crack detection method based on the sparse query mechanism described in the first embodiment above;

[0058] Step S2: using parallel modality independent detection branches to perform preliminary detection of long-distance crack targets on the two-dimensional image modality and the three-dimensional point cloud modality to generate two-dimensional long-distance crack detection results and three-dimensional long-distance crack detection results;

[0059] Specifically:

[0060] Step S21: adding an enhanced feature fusion module to the neck network part of YOLOv11 to obtain a YOLOv11 image detection module with enhanced multi-level feature fusion;

[0061] Step S22: using the YOLOv11 image detection module with enhanced multi-level feature fusion to process the two-dimensional image modality and generate a two-dimensional long-distance crack detection result;

[0062] Step S23: Use the fully sparse voxel FSDv2 point cloud detection module to process the three-dimensional point cloud modality to generate a three-dimensional long-distance crack detection result.

[0063] In practical applications, the parallel modality-independent detection branches of this implementation include a YOLOv11 image detection module with enhanced multi-level feature fusion and a fully sparse voxel FSDv2 point cloud detection module;

[0064] Among them, the enhanced feature fusion module (Enhance Fusion) is introduced in the neck network part of YOLOv11 to form a YOLOv11 image detection module with enhanced multi-level feature fusion, which can significantly enhance the feature representation ability. The enhanced feature fusion module includes an adaptive low-pass filter generator, an offset generator, and an adaptive high-pass filter generator. The improved YOLOv11 framework is as follows Figure 2 shown.

[0065] The processing flow of YOLO v11 is as follows:

[0066] (1) The input image F is passed through the backbone network to extract multi-scale features. The backbone network consists of a stacked convolutional layer (Conv) and a C3K2 module. The C3K2 block optimizes the information flow in the network by segmenting the feature map and applying a small kernel convolution to each branch:

[0067] (X3,X4,X5)=Backbone(F)

[0068] Among them, X3, X4, and X5 represent multi-scale feature maps, whose sizes decrease successively but the semantic information is gradually enhanced.

[0069] (2) Deep feature X 5 Perform the spatial pyramid pooling module (SPFF) to aggregate contextual information and enhance the network receptive field:

[0070] F SPFF =SPFF(X5)

[0071] (3) Through the cross-stage spatial attention module (C2PSA) SPFF The features are:

[0072] Y5=C2PSA(F SPFF )

[0073] (4) Through the pyramid network (FPN) composed of a strong feature fusion module (Enhance Fusion), deep fusion and extraction of multi-scale features are achieved:

[0074] (P3,Y4)=FPN(Y5,X4,X3)

[0075] (5) Further integration of features through path aggregation network (PAN):

[0076] (P4,P5)=PAN(P3,Y4,Y5)

[0077] (6) Each scale feature is used for target classification and bounding box regression by the detection head (Detect), and the final detection result is obtained through post-processing:

[0078] b img =Detect(P3,P4,P5)

[0079] Through the above process, the YOLOv11 detector generates image features F img And 2D detection results M img Indicates the number of detection boxes.

[0080] The processing flow of the enhanced feature fusion module is as follows:

[0081] The enhanced feature fusion module (Enhance Fusion) can be formally expressed as:

[0082]

[0083] Among them, F LP represents the low-pass filter predicted by the low-pass filter generator, (u, v) represents the offset value predicted by the offset generator for the feature coordinate (i, j), and F HP represents the high-pass filter predicted by the high-pass filter generator, F UP represents the upsampling operation, X l represents the shallow features of the lth layer of YOLOv11, Y l+1 represents the deep features of the l+1th layer, Y l Represents the fused features.

[0084] To efficiently generate a low-pass filter F LP , offset (u,v) and high-pass filter F HP , feature X l , Y l+1 After compression and fusion, they are input into the three generators. This process is called initial fusion. The process can be formally expressed as:

[0085] Z l =Concat(F UP (F LP (Conv 1×1 (Y l+1 )))),F HP (Conv 1×1 (X l ))+X l )

[0086] in, is the initial fusion feature, r is the channel compression rate, which is used to reduce the computational overhead. The 1×1 convolution kernel is used for channel compression. l+1 ∈R C×H×W is the deep feature, X l ∈R C×2H×2W It is a shallow feature.

[0087] The adaptive low-pass filter generator consists of a 3×3 convolution layer and a softmax layer, and its form is expressed as:

[0088]

[0089] in, represents the spatially variable filter weights, is the kernel size of the low-pass filter. After reshaping, it contains the Filter weight, Ω represents the window range of the filter, and p, q represent the relative position offset of the filtering operation. After the softmax operation of each kernel, the filter weight is constrained to be non-negative and the sum is 1, generating a low-pass filter with smooth characteristics. The result is expressed as

[0090] Upsampling Technology F UP The specific implementation can be expressed as: Pixel Unshuffle the weight matrix The reshaped weight matrix is ​​divided into 4 groups, each corresponding to a spatially variable low-pass filter, expressed as Where g∈{1,2,3,4} is the group index. Through this process, 4 sets of low-pass filtered features can be obtained, expressed as These features are then rearranged to form a 2x upsampled feature The calculation formula is as follows:

[0091]

[0092] The calculation formula of PixelShuffle() is specifically expressed as:

[0093]

[0094] Where c represents the channel index; i and j are the spatial positions in the low-resolution feature map; a and b∈{0,1} are the offsets when upsampling by a factor of 2; and g=a·2+b+1 represents the grouping index.

[0095] The adaptive high-pass filter generator consists of a 3×3 convolution layer, a softmax layer, and a filter inversion operation, and its form is expressed as:

[0096]

[0097] in, represents the initial kernel weight at each position (i, j), represents the kernel size of the high-pass filter, and E represents the unit filter weight matrix. The final result is calculated through the residual connection as:

[0098]

[0099] The offset generator first guides the offset direction by calculating the local cosine similarity, which is mathematically expressed as:

[0100]

[0101] Among them, S∈R 8×H×W Represents the cosine similarity between each pixel and its eight neighboring pixels.

[0102] The offset generator consists of two 3×3 convolutional layers, with feature Z l And cosine similarity S as input, predict the offset direction and offset amplitude, which is formally expressed as:

[0103] O l =D l ·A l

[0104] D l =Conv 3×3 (Concat(Z l ,S l ))

[0105] A l =Sigmoid(Conv 3×3 (Concat(Zl,S l )))

[0106] Among them, D l ∈R 2D×H×W Indicates the offset direction, A l ∈R 2G×H×W Indicates the magnitude of the control deviation, O l ∈R 2G×H×W The final offset prediction for each pixel in the deep feature. G represents the number of offset groups, which is achieved by dividing the features into different groups and assigning a unique spatial offset to each group to achieve finer resampling.

[0107] In this embodiment, the YOLOv11 detector generates enhanced image features F by embedding the Enhance Fusion module. img And 2D detection results M img Indicates the number of detection boxes.

[0108] The fully sparse voxel FSDv2 point cloud detection module is built on the FSDv2 detector, such as Figure 3 As shown in the figure, the fully sparse voxel mechanism is used to realize efficient 3D target detection of point clouds. The specific process is as follows:

[0109] (1) Input point cloud P∈R N×3 The initial voxel feature F is extracted by the Sparse Voxel Feature Extractor voxel :

[0110] F voxel =SparseEncoder(P)

[0111] (2) Extracting point features F from voxel features point , predict the foreground points through point classification and perform center voting to generate the voting center C voted :

[0112]

[0113] Among them, ΔP i Indicates the center shift of the prediction.

[0114] (3) The voting center C voted And the original point cloud P is voxelized to form virtual voxels and real voxels:

[0115] V virtual =Voxelize(C voted )

[0116] Virtual voxels contain at least one voting center, while real voxels are only generated from the original point cloud.

[0117] (4) Aggregate the point features inside each virtual voxel and generate enhanced virtual voxel features F through the virtual voxel encoder (VVE) virtual :

[0118] F virtual =VVE(F point ,C voted )

[0119] (5) The features of virtual voxels and real voxels are input into a lightweight virtual voxel mixer (VVM) for feature fusion to generate a unified feature F mixed :

[0120] F mixed =VVM(F virtual ,F real )

[0121] (6) Directly regress the 3D detection box parameters from the mixed features, including position (x, y, z), size (w, l, h) and direction θ:

[0122] b pc =MLP(F mixed )

[0123] Through the above process, the FSDv2 detector generates sparse voxel features F pc And 3D detection results M pc Indicates the number of detection boxes.

[0124] In summary, this implementation addresses the intra-category inconsistency and boundary offset problems generated in the standard multi-level feature fusion process by introducing a spatially variable low-pass filter in the feature fusion stage to adaptively smooth high-level features; replacing inconsistent features by resampling features with consistent adjacent categories; and introducing an adaptive high-pass filter to enhance the high-frequency information of low-level features, thereby significantly improving the category consistency of the fused features and refining the boundary representation of long-distance crack targets.

[0125] Furthermore, in response to the need for detecting fine textures of long-distance cracks, this implementation significantly improves the initial detection performance by enhancing the multi-level fusion features of the modalities while maintaining computational efficiency.

[0126] Implementation method 3: This implementation method specifically describes step S3 in the long-distance crack detection method based on the sparse query mechanism described in the above implementation method;

[0127] Step S3: performing sparse query on the two-dimensional and three-dimensional long-distance crack detection results to generate image query data and point cloud query data, and uniformly encoding the image query data and the point cloud query data;

[0128] Specifically:

[0129] Step S31: Use an uncertainty-aware image query generator to perform sparse query on the two-dimensional long-distance crack detection results.

[0130] Generate image query data including content features and location features;

[0131] Step S32: taking the center point of the three-dimensional long-distance crack detection result as the position feature part of the point cloud query;

[0132] Step S33: combining the appearance features and the geometric features to obtain the content feature part of the point cloud query;

[0133] Step S34: retain the content features in the above image query and the content features in the point cloud query, and perform encoding conversion on the position features in the image query and the position features in the point cloud query to complete unified encoding.

[0134] In practical applications, the detection results of the image and point cloud modalities generated by the above-mentioned second embodiment provide key information for long-distance crack target recognition, but their representation forms are essentially different. Therefore, this embodiment extracts target-level semantics from the detection results and encodes the semantic information of long-distance crack targets in the form of sparse target queries. The target query consists of two parts: content features and location features, which respectively describe the semantic attributes and spatial distribution of the target, providing a consistent representation form for the fusion of multimodal information.

[0135] Among them, long-distance crack target query generation is performed for two-dimensional image modality, specifically: by introducing an uncertainty-aware image query generator to generate image queries containing two parts: content and location.

[0136] The content part is obtained from the feature map F by the ROI-Align method. img and the two-dimensional detection box b img Extract appearance features Among them, H t ×W t Indicates the spatial size, expressed in the form of:

[0137] O img =RoI-Align(F img ,b img )

[0138] In order to make up for the lost geometric information in the RoI-Align process, content features Enhancement is performed by stitching the camera's equivalent internal parameter matrix K, which can be expressed as:

[0139] c img =MLP(Concat(Pool(Conv(o img )),Flat(K)))

[0140] Among them, Flat() means flattening the tail dimension of the tensor. Represents the equivalent internal parameter matrix, and the matrix corresponding to the i-th detection target is:

[0141]

[0142] in, f x 、f y Respectively represent the horizontal and vertical focal lengths of the camera, o x , o y Respectively represent the horizontal and vertical coordinates of the camera's principal point.

[0143] The position part is within the predefined depth range [d min ,d max ] Uniform sampling n d depth values, forming a depth set Then, a set of 2D sampling locations is predicted and the corresponding probability The form is:

[0144] [s 2d ;u logit ]=MLP(c img )

[0145] u img =softmax(u logit )

[0146] Where (";") represents the concatenation operation along the channel dimension. 2d And the depth value d, the corresponding three-dimensional sampling position s can be obtained img Combining the content part and the location part, the image query q is finally generated img =(c img ,s img ,u img ).

[0147] The long-distance crack target query generation is performed for the three-dimensional point cloud mode. Specifically, the FSDv2 detector directly outputs the long-distance crack target detection result in the three-dimensional space. In this implementation, the center point of the target is As the location part of the point cloud query, for the content part Combined with appearance features pc and geometric features b pc , expressed in the form of:

[0148] c pc =MLP(o pc +MLP(SinPos(b pc )))

[0149] Among them, pcis the sparse voxel feature extracted by the FSDv2 detector, and SinPos() represents the sinusoidal position encoding. By combining the content part and the position part, the point cloud query q is finally generated. pc =(c pc ,r pc ).

[0150] The specific unified coding is as follows: Since the query forms of different modalities are different, in order to achieve a unified representation, this implementation method retains the content part of the query and converts the position part to achieve a unified coding representation. The specific method is as follows:

[0151] p pc =PE(r pc )

[0152] p img =U-PE(s img ,u img )

[0153] Among them, PE represents position encoding, U-PE represents uncertainty-aware position encoding, and its calculation formula is:

[0154] PE(r pc )=MLP(SinPos(r pc ))

[0155] U-PE(s img ,u img )=MLP(MLP(Flat(s img ))⊙σ(MLP(u img )))

[0156] Among them, Flat() represents the flattening operation, ⊙ represents element-by-element multiplication, and σ represents the sigmoid function.

[0157] Implementation mode 4: This implementation mode specifically describes step S4 in the long-distance crack detection method based on the sparse query mechanism described in the above implementation mode;

[0158] Step S4: Perform precise detection of long-distance crack targets on the uniformly encoded image query data and point cloud query data to generate long-distance crack detection results.

[0159] Specifically:

[0160] S41: Use the proxy attention mechanism to process the uniformly encoded query data and capture the dependencies between queries to obtain optimized queries;

[0161] S42: using a cross attention mechanism to process the above optimized query data and external modal features, aggregating useful features, and obtaining an optimized query;

[0162] S43: calibrating the image query in the query data using a query calibration mechanism;

[0163] S44: After calibration, the query content features and anchor point positions are processed by the classification head and the regression head to generate long-distance crack detection results.

[0164] In practical applications, this implementation improves the attention mechanism by introducing additional proxy tags, significantly improving computational efficiency while retaining the global context modeling capability. The stacked decoder consists of a proxy attention layer, a cross attention layer, a layer normalization module, a feedforward network, and a query calibration module to achieve efficient information interaction and feature optimization.

[0165] Among them, the proxy attention mechanism is specifically:

[0166] Proxy attention introduces an additional set of proxy tags A into the traditional attention module, first aggregating information from K and V, and then broadcasting the aggregated information back to Q. In order to reduce the impact of redundant calculations in the traditional attention mechanism, the present invention significantly reduces the computational overhead while retaining the global context modeling capability by designing the number of proxy tags to be much smaller than the number of Q. The process can be expressed as:

[0167] O=σ(QA T +B2)σ(AK T +B1)V+DWC(V)

[0168] The calculation formulas for query, key, value, and proxy tag are:

[0169] Q=WQ(c sa +p sa ),K=WK(c sa +p sa ),V=WVc sa

[0170] A=AdaptivePooling(Q)

[0171] Among them, p sa Indicates that p pc and p img For splicing, c sa Similarly, σ represents the softmax function, AdaptivePooling represents the global average pooling operation, B1 and B2 are learnable proxy biases used to enhance the proxy markers’ modeling capabilities for spatial information, and DWC represents the deep convolution module used to supplement the lack of feature diversity.

[0172] The cross attention mechanism is as follows:

[0173] For image features, this implementation adopts a projection-based deformable attention mechanism. First, the anchor point position of each query is determined, where the anchor point position a of the point cloud query is pc Its three-dimensional position r pc , the anchor point position of the image query is distributed through the probability distribution u img and sampling position s img The weighted average of is calculated to obtain that for the i-th image query, the calculation formula for its anchor point position is:

[0174]

[0175] The calculation formula of the projection-based variability attention mechanism is expressed as:

[0176]

[0177] Among them, the attention weight matrix A k and offset Δa k From the content part c m In the prediction, K represents the number of sampling points, Proj() represents the projection transformation from world coordinates to camera coordinates, and m represents the query modality (image or point cloud).

[0178] For point cloud features, a proxy attention mechanism is used to achieve feature aggregation. Point cloud feature F pc Generate the content part c by average pooling in the height direction pillar , the position code of the column feature p pillar Then the BEV position r pillar generate:

[0179] p pillar =MLP(SinPos(r pillar ))

[0180] Furthermore, this embodiment uses the updated features to calibrate the image query after each decoder layer to optimize the query position and reduce uncertainty. In this process, only the probability distribution u img Calibration is performed, and the sampling position s img Keeping the same, the calculation formula of this process is expressed as:

[0181] u logit =log(u img )

[0182] u img ←softmax(ulogit+MLP(c img ))

[0183] Querying the calibration layer will affect the position encoding p accordingly imgand anchor point position a img .

[0184] The final long distance crack detection result output:

[0185] After obtaining the query Q of the last layer decoder, the query is processed by the classification head and regression head to generate the result output. The process is expressed as:

[0186] z cls =MLP(c)

[0187] z reg =MLP(c)+[a;0]

[0188] Among them, z cls represents the classification score, z reg represents the regression target, formally expressed as (x, y, z, w, l, h, rot), c and a represent the content features and anchor positions of the last layer, respectively.

[0189] Implementation mode 5: This implementation mode is to provide a detailed description of the long-distance crack detection method based on the sparse query mechanism described in the above implementation mode as a whole:

[0190] Step 1: Dataset construction:

[0191] In order to meet the needs of long-distance crack target detection, the dataset is constructed based on high-resolution two-dimensional images and point cloud data, aiming to comprehensively utilize multimodal information to improve detection performance.

[0192] Step 1.1 Image data acquisition and annotation:

[0193] Using high-resolution industrial cameras, we can collect 2D images of cracks on the surface of target structures at long distances. The collected objects cover key infrastructure such as dams, bridges, and tunnels, and cover a variety of complex scenes, such as changing lighting, complex backgrounds, and noise interference environments, to ensure the diversity and representativeness of the data.

[0194] The collected original images are digitally processed. First, the target area of ​​1280×1280 is cropped, followed by denoising, resampling and data enhancement operations. Data enhancement includes random flipping, brightness adjustment, scaling and cropping to improve the adaptability of the model to diverse data.

[0195] A manual annotation strategy is adopted to select the crack target area using professional annotation tools to generate two-dimensional bounding box annotation data. Each annotation box contains the coordinates of the target (x min ,y min ,x max ,y max ), and at the same time, classify and distinguish the target from the background.

[0196] Step 1.2 Point cloud data collection and annotation:

[0197] A high-precision LiDAR sensor is used to synchronously collect point cloud data to generate a 3D point cloud representation of the target area. The specific parameters of the LiDAR are set to: horizontal resolution 0.2°, vertical resolution 0.1°, ranging accuracy ±2mm, and acquisition frequency 10Hz. The point cloud acquisition range is consistent with image acquisition to achieve alignment of data between modalities.

[0198] The collected point cloud data is subjected to noise reduction, and invalid points are removed using spatial filtering and isolated point elimination algorithms.

[0199] A manual annotation strategy is adopted to annotate the crack target area with a 3D bounding box using a 3D point cloud annotation tool to generate point cloud annotation data. Each bounding box contains the target's position (x, y, z), size (w, l, h) and orientation rot.

[0200] Step 2 Network training:

[0201] In the network training part, the training process of this embodiment includes independent training of the two-dimensional detector and the three-dimensional detector and joint training of the overall network to gradually improve the detection performance.

[0202] Step 2.1 YOLOv11 detector training:

[0203] The training of the YOLOv11 detector adopts a three-stage strategy, using the COCO dataset, the CrackForest dataset, and a self-constructed long-distance crack dataset. The specific process is as follows:

[0204] (1) Pre-training stage: large-scale pre-training is performed on the COCO dataset to enable the network to learn general object detection features;

[0205] (2) Specific domain adaptation stage: training on the CrackForest dataset to adapt the model to the characteristic distribution of crack targets;

[0206] (3) Fine-tuning stage: Use a self-constructed long-distance crack dataset for fine-tuning to further improve the model's detection capabilities for specific scenarios.

[0207] During the training process, the loss function and training process follow the original design of YOLOv11, including classification loss, regression loss, and confidence loss.

[0208] Step 2.2 FSDv2 detector training:

[0209] The training of the FSDv2 detector is divided into two stages, using the Waymo Open Dataset (WOD) and a self-constructed long-distance crack point cloud dataset:

[0210] (1) Pre-training stage: large-scale pre-training is performed on the WOD dataset to enable the model to learn the feature expressions of various 3D objects;

[0211] (2) Fine-tuning stage: Fine-tuning is performed on the self-constructed dataset to optimize the model performance based on the specific characteristics of long-distance fracture targets.

[0212] The training process and loss function design of the FSDv2 detector completely follow the definition of its original model, including center point regression loss, voxel feature learning loss, etc.

[0213] Step 2.3 Overall network training:

[0214] In the joint training of the whole network, the decoder and the final detection head are trained using the self-constructed image dataset and point cloud dataset to achieve multimodal fusion and 3D object detection. The specific process is as follows:

[0215] (1) Target allocation and loss function:

[0216] 1) The output target allocation and loss function of the decoder follow the design of DETR, and the Hungarian algorithm is used for label allocation;

[0217] 2) Using Focal Loss for target classification and L1 Loss for bounding box regression, the 3D target detection loss is summarized as:

[0218] L out =λ cls ·L cls +L reg

[0219] (2) Auxiliary depth estimation supervision:

[0220] Auxiliary supervision is added to the image query generator to facilitate depth estimation. For the image predicted 2D bounding box b img and the 2D bounding box projected from the true 3D bounding box First, calculate the pairwise intersection over union (IoU) matrix U between the two:

[0221]

[0222] Satisfy the row and column maximum value constraint and exceed the preset IoU threshold Assign to target Depth distribution of image query generator output Using auxiliary loss L aux Supervised by target depth

[0223]

[0224] Among them, CELoss() represents the cross entropy loss.

[0225] (3) Overall loss function:

[0226] The loss function of the overall network combines the losses of two-dimensional detection, three-dimensional detection, auxiliary supervision and final detection, and the formula is:

[0227] L=λ det2D ·L det2D +λ det3D ·L det3D +λ aux ·L aux +λ out ·L out

[0228] Among them, L det2D and L det3D They represent the loss functions of YOLOv11 and FSDv2 respectively, and the hyperparameter {λ} is used to balance the scales of different loss terms.

[0229] The above steps are used to implement a long-distance crack detection method based on a sparse query mechanism described in the present invention, thereby solving many technical problems existing in existing crack detection methods, such as: 1. The problem of limited accuracy of long-distance crack detection due to insufficient single-modal information, the intra-category inconsistency and boundary offset problems generated in the standard multi-level feature fusion process, the high computational overhead of the existing attention mechanism or the insufficiency of computational optimization at the cost of limited representation ability, and the limitation of the existing method that the reasoning speed is slow and it is difficult to meet the needs of actual scenarios.

[0230] The above description is only the implementation mode of the present invention and is not limited to the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the scope of the claims of the present invention.

Claims

1. A long-distance crack detection method based on sparse query mechanism, characterized in that: The method is: S1: Acquire the 2D image modality and 3D point cloud modality of the detection target; S2: Use parallel modality-independent detection branches to perform preliminary detection of long-distance crack targets on the two-dimensional image modality and the three-dimensional point cloud modality, respectively, to generate two-dimensional long-distance crack detection results and three-dimensional long-distance crack detection results; S3: Perform sparse query on the two-dimensional and three-dimensional long-distance crack detection results to generate image query data and point cloud query data, and uniformly encode the image query data and point cloud query data; S4: Perform accurate detection of long-distance crack targets on the uniformly encoded image query data and point cloud query data to generate long-distance crack detection results.

2. A long distance crack detection method based on sparse query mechanism according to claim 1, characterized in that: The parallel modality-independent detection branch includes the YOLOv11 image detection module with enhanced multi-level feature fusion and the fully sparse voxel FSDv2 point cloud detection module.

3. A long distance crack detection method based on sparse query mechanism according to claim 2, characterized in that: S2 is specifically: S21: Add an enhanced feature fusion module to the neck network part of YOLOv11 to obtain a YOLOv11 image detection module with enhanced multi-level feature fusion; S22: The YOLOv11 image detection module with enhanced multi-level feature fusion is used to process the two-dimensional image modality to generate two-dimensional long-distance crack detection results; S23: The fully sparse voxel FSDv2 point cloud detection module is used to process the three-dimensional point cloud modality to generate three-dimensional long-distance crack detection results.

4. A long distance crack detection method based on sparse query mechanism according to claim 3, characterized in that: The enhanced feature fusion module includes an adaptive low-pass filter generator, an offset generator and an adaptive high-pass filter generator.

5. The long-distance crack detection method based on sparse query mechanism according to claim 1 is characterized in that: S3 is specifically: S31: Using an uncertainty-aware image query generator to perform sparse queries on two-dimensional long-distance crack detection results, generating image query data containing content features and location features; S32: taking the center point of the three-dimensional long-distance crack detection result as the position feature part of the point cloud query; S33: Combining appearance features and geometric features to obtain content feature part of point cloud query; S34: retaining the content features in the above-mentioned image query and the content features in the point cloud query, and performing encoding conversion on the position features in the image query and the position features in the point cloud query to complete unified encoding.

6. The long-distance crack detection method based on sparse query mechanism according to claim 1 is characterized in that: S4 is specifically: S41: Use the proxy attention mechanism to process the uniformly encoded query data and capture the dependencies between queries to obtain optimized queries; S42: using a cross attention mechanism to process the above optimized query data and external modal features, aggregating useful features, and obtaining an optimized query again; S43: calibrating the image query in the query data using a query calibration mechanism; S44: After calibration, the query content features and anchor point positions are processed by the classification head and the regression head to generate long-distance crack detection results.

7. The long-distance crack detection method based on sparse query mechanism according to claim 5 is characterized in that: The proxy attention mechanism is constructed by introducing an additional set of proxy tags A into the traditional attention module to obtain the proxy attention mechanism.

8. A long distance crack detection system based on sparse query mechanism, characterized in that: The system includes a storage device, which is used to execute the long-distance crack detection method based on a sparse query mechanism described in claim 1.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the long-distance crack detection method based on a sparse query mechanism described in any one of claims 1 to 7 is executed.

10. A computer device, characterized in that: The device includes a memory and a processor, wherein a computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes a long-distance crack detection method based on a sparse query mechanism as described in any one of claims 1 to 7.

Citation Information

Cited By

  • BEV dense and sparse hybrid multi-task sensing method and device, electronic equipment and storage medium

    CN121582759A