Corrosion reinforced concrete crack detection method and device based on YOLOV11

By constructing the YOLOV11 cross-view comparison model and DVT_FVS module, the problem of low accuracy and efficiency in crack detection of rusted reinforced concrete is solved, efficient and accurate crack recognition is achieved, and the stability and accuracy of detection are enhanced.

CN120259319AActive Publication Date: 2025-07-04CHINA RAILWAY FIRST GROUP CO LTD +3

Patent Information

Application Number
CN202510757549.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-04
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In the detection of cracks of rusted reinforced concrete, the detection accuracy and low efficiency are poor, making it difficult to achieve efficient and accurate crack identification under complex concrete surface conditions.

Method used

The crack detection method of rust reinforced concrete based on YOLOV11 is adopted. By constructing a cross-view comparison model of YOLOV11, including backbone network, neck network and head network, combined with the DVT_FVS module for video stabilization processing, real-time crack detection is achieved.

Benefits of technology

It improves the accuracy and efficiency of rusted reinforced concrete crack detection, enhances feature learning ability, improves the generalization ability of the network, and provides reliable technical support for building structure safety assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259319A_ABST
    Figure CN120259319A_ABST
Patent Text Reader

Abstract

The invention provides a rusted reinforced concrete crack detection method and device based on YOLOV11, and belongs to the technical field of image processing. The method comprises the following steps: acquiring crack image data of a corroded reinforced concrete member, and marking the image data; combining the labeled image data with labels to form a data set; constructing a YOLOV11 cross-view comparison model, wherein the model comprises a backbone network, a neck network and a head network; the YOLOV11 cross-view comparison model is trained, and an optimal model is obtained; and acquiring crack video data of a to-be-detected rusted reinforced concrete member, inputting the crack video data into the DVTFVS module for image stabilization processing, transmitting a video into the YOLOV11 cross-view comparison model for crack detection, and outputting position, size and category information of the crack. According to the invention, the efficiency and the accuracy of corrosion reinforced concrete crack detection are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and device for detecting rusted reinforced concrete cracks based on YOLOV11, belonging to the technical field of image processing. Background Art

[0002] The detection of rusted reinforced concrete cracks is a core task for ensuring structural safety in the field of construction engineering, which is widely related to the durability and reliability of building facilities, covering various infrastructure such as bridges, building buildings, and hydraulic dams. In this field, deep learning technology is gradually becoming a key means to improve the accuracy and efficiency of crack detection, used to achieve precise positioning, quantitative analysis of concrete cracks, and prediction of the development trend of cracks. With the continuous improvement of the requirements for structural health monitoring in the construction industry, more cutting-edge technologies have been introduced into the modern field of rusted reinforced concrete crack detection, aiming to break through the limitations of traditional detection methods and achieve more efficient and intelligent detection.

[0003] Compared with the cracking of concrete caused by stress, the rust expansion cracking is usually accompanied by characteristics such as reddish-brown rust marks, cracking along the reinforcement, and relatively large crack widths, which are easy to distinguish and identify. At present, there are many traditional methods for detecting rusted reinforced concrete cracks. The method based on threshold segmentation divides the pixels in the image into crack and non-crack regions by setting a fixed or adaptive threshold. Its calculation is simple and it can quickly achieve segmentation for images with a relatively simple background and obvious crack features. However, the surface conditions of concrete in actual engineering are complex and there are various interference factors. This method is extremely sensitive to the selection of the threshold, and improper threshold setting will lead to misjudgment or missed judgment of cracks. The edge detection algorithm determines the crack edge by detecting the mutation of the gray value in the image and has a certain effect on cracks with clear edges. However, in the scenario of rusted reinforced concrete, the crack edge is often blurred and affected by rust products of steel bars, etc., and it is easy to produce discontinuous edge detection results, affecting the complete identification of cracks. The semantic segmentation algorithm in the detection of rusted reinforced concrete cracks attempts to classify each pixel in the image to accurately divide the crack region and the background region. Although theoretically it can provide relatively detailed crack contour information, in actual applications, due to the diversity of the surface materials of concrete, the change of lighting conditions, and the irregularity of crack shapes, the semantic segmentation model is prone to category confusion, misjudging non-crack regions as cracks, or vice versa, resulting in limited detection accuracy. Object detection algorithms based on region proposals such as FasterR-CNN (Fast R-CNN) generate candidate regions that may contain cracks, and then classify and regress these regions to determine the location and size of the cracks. However, in the task of detecting rusted reinforced concrete cracks, the crack shapes are variable and often intertwined, and the candidate regions generated by FasterR-CNN are difficult to accurately cover all cracks, and the calculation cost is relatively high, and the efficiency is low when processing a large number of images.

[0004] Therefore, how to provide a rusty reinforced concrete crack detection method based on YOLOV11 to improve the accuracy and stability of detection is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] The object of the present invention is to provide a rusty reinforced concrete crack detection method and device based on YOLOV11, which solves the problems of poor accuracy and low efficiency in detecting rusty reinforced concrete cracks under complex concrete surface conditions.

[0006] To achieve the above object, the present invention is realized through the following technical solutions: A rusty reinforced concrete crack detection method based on YOLOV11 includes: Obtain the crack image data of the rusty reinforced concrete component, and annotate the crack image data to generate corresponding labels; Combine the annotated image data and labels into a data set, and divide it into a training set, a validation set and a test set; Construct a YOLOV11 cross-view contrast model, which includes a backbone network, a neck network and a head network; the backbone network sequentially includes a VMTV module (cross-view attention network), a C3K2 module, a CSWin_tiny module (cross-window sliding Transformer module), an SPPF (spatial pyramid pooling fast layer) module and a C2PSA module; wherein the VMTV module includes a view branch module, a cross-view interaction module, a two-stage dynamic fusion module and an auxiliary module; Use the training set to train the YOLOV11 cross-view contrast model, and tune the hyperparameters through the validation set to obtain the best model; During the movement of the camera, the crack video data of the rusty reinforced concrete component to be detected is collected in real time, input into the DVT_FVS module (cross-view dynamic video stabilization module) for video stabilization processing, and then the video is input into the YOLOV11 cross-view contrast model for crack detection, and the position, size and category information of the cracks are output.

[0007] Preferably, the C2PSA module output feature of the backbone network of the YOLOV11 cross-view contrast model Is input to the neck network for upsampling and then concatenated with the feature output by the SPPF module of the backbone network, and the concatenated feature is processed by the C3K2 module to obtain a feature , the feature After upsampling, it is concatenated with the feature output by the C3K2 module of the backbone network to obtain a feature , the feature Is output to the DVT_FVS module of the head network, and the feature Is also processed through a convolutional layer and concatenated with the feature The input is spliced and fed into the C3K2 module to obtain features , and the features are output to the DVT_FVS module of the head network, and the features are also processed through a convolutional layer and spliced with the features to be input into the C3K2 module to obtain features , and the features are output to the head network.

[0008] Preferably, the VMTV module includes: View branch module: It adopts 3 parallel multi-view Transformer branches, and each branch includes an image cell region embedding layer, a dynamic position encoding layer, 4 layers of VIT hierarchical Transformer layers (visual Transformer hierarchical Transformer layers), an SE channel attention module, a feed-forward network, a normalization residual block, and a cross-scale feature pyramid connection layer; Cross-view interaction module: It includes an intra-view multi-scale self-attention layer, a cross-scale cross-attention layer, and a global semantic alignment layer; Two-stage dynamic fusion module: It includes a scale dimension fusion module and a view dimension fusion module; Auxiliary module: It includes hierarchical dynamic routing, a cross-view contrast learning module, and a Figure 1 consistency supervision module.

[0009] Preferably, in the view branch module: The input image is divided into image cell regions (Patches) of corresponding sizes through the image cell region embedding layer and embedding vectors are generated; The learnable sine encoding of the embedding vector is spliced with the local position deviation matrix through the dynamic position encoding layer and linearly projected to the feature dimension; Then it is successively processed through 4 layers of VIT hierarchical Transformer layers, an SE channel attention module, a feed-forward network, and a normalization residual block; the VIT hierarchical Transformer layer includes an 8-head self-attention layer; The cross-scale feature pyramid connection layer splices the three-size features output by the normalization residual block to generate a single-view multi-scale feature, specifically by upsampling the low-scale feature and splicing it with the current-scale feature to generate a single-view multi-scale feature.

[0010] Preferably, in the cross-view interaction module: The intra-view multi-scale self-attention layer includes 4 layers of VIT hierarchical Transformer layers, an SE channel attention module, a feed-forward network, and a normalization residual block, and each layer of the VIT hierarchical Transformer layer includes an 8-head self-attention layer; The cross-scale cross-attention layer includes a bilinear interpolation layer, a convolutional layer, a 12-head cross-attention layer, and a feed-forward network; The global semantic alignment layer includes a feature concatenation layer and a 3-layer global self-attention layer. The feature concatenation layer flattens the multi-scale features of all views into a unified sequence, and then adds view identity embeddings and scale identity embeddings to mark the feature sources. The global self-attention layer includes an 8-head self-attention layer, a feed-forward network, and a dynamic position encoding layer.

[0011] Preferably, in the two-stage dynamic fusion module: The scale dimension fusion module includes a feature concatenation layer, a double-hidden layer FC weight generator (double-hidden layer fully connected weight generator), and a weighted fusion layer; The view dimension fusion module serializes the scale fusion feature of each view into tokens, and then uses a 3-layer VIT-style self-attention layer to complete cross-view feature fusion. The VIT-style self-attention layer includes an 8-head self-attention layer, a feed-forward network, layer normalization, and residual connections.

[0012] Preferably, in the auxiliary module: The hierarchical dynamic routing generates weights for local tokens through a 1-layer fully connected layer and generates weights for global views through a 1-layer fully connected layer; The cross-view contrast learning module uses a 2-layer fully connected projection head to calculate the NT-Xent contrast loss (normalized temperature-scaled cross-entropy loss); The view Figure 1 The consistency supervision module uses a 3-layer MLP prediction head (multi-layer perceptron prediction head), combines the MSE loss (mean squared error loss) and the contrast loss to generate prediction values.

[0013] Preferably, the DVT_FVS module includes: The preprocessing module: uses blind deconvolution motion blur kernel estimation and adaptive non-local mean denoising, combines scale normalization and data augmentation to complete the preprocessing of the input frame; The feature fusion module: embeds the optical flow into a latent representation through a ResNet34_CBAM encoder (Residual Neural Network 34 Convolutional Block Attention Module Encoder), and connects it with the poses of real and virtual cameras; infers a new virtual camera pose through a BiLSTM+MobileViT block (Bidirectional Long Short-Term Memory Network + Mobile Vision Transformer Block), and uses this virtual camera pose to generate a warping grid; then realizes spatio-temporal feature fusion through a 3D convolutional module and an Hourglass network to predict the new virtual camera pose as a quaternion; warps the input frame based on optical image stabilization and the virtual camera pose to generate a stable frame.

[0014] Preferably, the training process of the DVT_FVS module adopts: the dynamic weight GradNorm strategy (dynamic weight gradient norm strategy), and cooperatively optimizes the smoothing loss in the BiLSTM stage (bidirectional long short-term memory network stage), the tracking loss in the Transformer stage, the ArcFace improved angle loss in the Hourglass stage, and the reconstruction loss of the entire process.

[0015] The advantages of the present invention are as follows: The present invention not only enhances the feature learning ability, but also improves the generalization ability of the network through reasonable model structure and hyperparameter tuning, ensures the efficiency and accuracy of corrosion reinforced concrete crack detection, and provides more reliable technical support for the safety assessment of building engineering structures. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.

[0017] Figure 1 It is a schematic flowchart of the method provided by the present invention; Figure 2 It is a schematic structural composition diagram of the YOLOV11 cross-view contrast model provided by the present invention; Figure 3 It is a schematic diagram of the view branch module provided by the present invention; Figure 4 It is a schematic diagram of the cross-view interaction module provided by the present invention; Figure 5 It is a schematic diagram of the two-stage dynamic fusion module provided by the present invention; Figure 6 It is a schematic diagram of the auxiliary module provided by the present invention; Figure 7 It is an architecture diagram of the DVT_FVS module provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment 1 As Figure 1 shown, a method for detecting corrosion reinforced concrete cracks based on YOLOV11 includes: S1: Obtain the crack image data of the corroded reinforced concrete member, and annotate the image data to generate corresponding labels.

[0020] Specifically, obtain the crack image data of the corroded reinforced concrete, and assign corresponding labels to the crack image data of the corroded reinforced concrete. The method for assigning corresponding labels is: use a high-definition camera to collect data of corroded reinforced concrete under different background environmental conditions, and use the professional annotation tool Labelimg to annotate the positions and categories of the cracks.

[0021] S2: Combine the annotated image data and labels into a data set, and divide it into a training set, a validation set, and a test set.

[0022] Specifically, combine the crack image data of the corroded reinforced concrete and the labels to form a data set, and divide the data set into training set samples, test set samples, and validation set samples. Divide the data set according to the ratio of 70%, 15%, and 15%, and perform preprocessing operations such as normalization, rotation, and flipping on the data to improve the data quality.

[0023] S3: Construct a YOLOV11 cross-view contrast model, which includes a backbone network, a neck network, and a head network. The backbone network sequentially includes a VMTV module, a C3K2 module, a CSWin_tiny module, an SPPF module, and a C2PSA module. Among them, the VMTV module includes a view branch module, a cross-view interaction module, a two-stage dynamic fusion module, and an auxiliary module.

[0024] S4: Use the training set to train the YOLOV11 cross-view contrast model, and obtain the best model by tuning the hyperparameters through the validation set.

[0025] Specifically, input the training set samples into the YOLOV11 cross-view contrast model for training to obtain an initial model, and use the validation set samples to tune the hyperparameters of the initial model to obtain the best model. Adopt the Adam optimization algorithm to tune hyperparameters such as the learning rate, batch size, and number of training epochs.

[0026] S5: The camera collects the crack video data of the corroded reinforced concrete member to be detected in real time during the movement process, inputs it into the DVT_FVS module for frame stabilization processing, and then transmits the video into the YOLOV11 cross-view contrast model for crack detection, and outputs the position, size, and category information of the cracks.

[0027] Specifically, input the test set samples into the best model to obtain the crack detection results of the corroded reinforced concrete. The model outputs the position, size, and category information of the cracks, and calculates evaluation indicators such as accuracy and recall rate at the same time.

[0028] As a refinement of the above embodiment, as Figure 2 shown, the YOLOV11 cross-view contrast model includes a backbone network, a neck network, and a head network.

[0029] The C2PSA module of the backbone network of the YOLOV11 cross-view contrast model outputs features which are input to the neck network. After upsampling, they are concatenated with the features output by the SPPF module of the backbone network. The concatenated features are processed by the C3K2 module to obtain features , and the features after upsampling are concatenated with the features output by the C3K2 module of the backbone network to obtain features , and the features are output to the DVT_FVS module of the head network. The features are also processed by a convolutional layer and concatenated with the features and then input to the C3K2 module to obtain features , and the features are output to the DVT_FVS module of the head network. The features are also processed by a convolutional layer and concatenated with the features and then input to the C3K2 module to obtain features , and the features are output to the DVT_FVS module of the head network.

[0030] As a refinement of the above embodiment, the VMTV module includes: View branch module: It adopts 3 parallel multi-view Transformer branches. Each branch includes an image cell region embedding layer, a dynamic position encoding layer, 4 layers of VIT hierarchical Transformer layers, an SE channel attention module, a feed-forward network, a normalization residual block, and a cross-scale feature pyramid connection layer; Cross-view interaction module: It includes an intra-view multi-scale self-attention layer, a cross-scale cross-attention layer, and a global semantic alignment layer; Two-stage dynamic fusion module: It includes a scale dimension fusion module and a view dimension fusion module; Auxiliary module: It includes hierarchical dynamic routing, a cross-view contrast learning module, and a Figure 1 consistency supervision module.

[0031] Specifically, as Figure 3 shown, in the view branch module: Within the view branch, three parallel multi-view Transformer branches are adopted (with image cell region sizes of 16 / 8 / 4 respectively). Each branch consists of an image cell region embedding layer, a dynamic position encoding layer, four layers of VIT hierarchical Transformer layers, an SE channel attention module, a feed-forward network, a normalization residual block, and a cross-scale feature pyramid connection layer.

[0032] Specifically, the input image is first divided into image cell regions of corresponding sizes by the image cell region embedding layer and embedding vectors are generated; subsequently, through the dynamic position encoding layer, the learnable sine encoding of the embedding vectors is concatenated with the local position deviation matrix and linearly projected onto the feature dimension; then it enters the four layers of VIT hierarchical Transformer layers, where each layer sequentially contains an 8-head self-attention layer, and then through the SE channel attention module (realizing dimensionality reduction and dimensionality increase through two fully connected layers), a feed-forward network (FFN, 64→1024→64), and a normalization residual block; finally, through the cross-scale feature pyramid connection layer, the low-scale features are upsampled and concatenated with the current-scale features to generate multi-scale features of a single view.

[0033] Specifically, as Figure 4 shown, in the cross-view interaction module: The multi-scale self-attention layer within the view includes four layers of VIT hierarchical Transformer layers, an SE channel attention module, a feed-forward network, and a normalization residual block. Each layer of the VIT hierarchical Transformer layer contains an 8-head self-attention layer; the SE channel attention module is used to perform global average pooling, dimensionality reduction and dimensionality increase fully connected layers, and an activation function to perform channel weighting on the self-attention output; the feed-forward network (FFN) consists of a fully connected layer and a GELU activation function (Gaussian error linear unit activation function); in addition, each layer also contains two layers of LayerNorm (layer normalization) and two layers of residual connections.

[0034] The cross-scale cross-attention layer includes a bilinear interpolation layer, a convolutional layer, a 12-head cross-attention layer, and a feed-forward network; the cross-scale cross-attention layer unifies the number of channels through the bilinear interpolation layer and a 1×1 convolutional layer; the 12-head cross-attention layer is used to split the input features into 12 heads to calculate the cross-view attention interaction, combines a lightweight FFN network to enhance the non-linear expression, and uses layer normalization and residual connections to fuse the cross-attention output with the input features.

[0035] The global semantic alignment layer includes a feature splicing layer and three layers of global self-attention layers. The multi-scale features of all views are flattened into a unified sequence through the feature splicing layer, and then the view identity embedding Eview and the scale identity embedding Escale are added to mark the feature sources. Global semantic integration is performed through three layers of global self-attention layers, and each layer contains eight-head self-attention layers (calculating the interactions between all tokens), a feed-forward network, and a dynamic position encoding layer.

[0036] Specifically, as Figure 5 shown, in the two-stage dynamic fusion module: Scale dimension fusion includes a feature splicing layer, a double-hidden-layer FC weight generator, and a weighted fusion layer. First, the three-scale features are spliced, and channel weights are generated through a double-hidden-layer FC network. After layer normalization, weighted summation with channel attention is performed on the multi-scale features. View dimension fusion serializes the scale fusion features of each view into tokens and completes the final cross-view feature fusion through three layers of VIT-style self-attention layers (each layer contains eight-head self-attention, FFN, layer normalization, and residual connections).

[0037] Specifically, as Figure 6 shown, in the auxiliary module: Hierarchical dynamic routing generates weights for local tokens through one layer of fully connected layer and generates weights for global views through one layer of fully connected layer, guiding the model to focus on key regions. The cross-view contrast learning module calculates the NT-Xent contrast loss using two layers of fully connected projection heads to enhance the view invariance of features; Figure 1 The view consistency supervision module generates prediction values through three layers of MLP prediction heads (with Dropout), and combines the MSE loss and the contrast loss to ensure the semantic consistency of different views. The outputs of each component are finally fused with the main process features to optimize the model detection performance.

[0038] As a refinement of the above embodiment, as Figure 7 shown, the DVT_FVS module includes: Preprocessing module: Blind deconvolution motion blur kernel estimation and adaptive non-local mean denoising are adopted, and input frame preprocessing is completed by combining scale normalization and data augmentation; Feature fusion module: The optical flow is embedded into a latent representation through a ResNet34_CBAM encoder, and it is connected with the poses of real and virtual cameras. A new virtual camera pose is inferred through a BiLSTM + MobileViT block, and a warping grid is generated using this virtual pose. Then, spatio-temporal feature fusion is achieved through a 3D convolution module and an Hourglass network to predict the new virtual camera pose as a quaternion. The input frame is warped based on optical image stabilization (OIS) and the virtual camera pose to generate a stable frame.

[0039] Specifically, the DVT_FVS module stabilizes videos using sensor data and optical flow through unsupervised learning. In the preprocessing stage, a motion blur kernel estimation module based on blind deconvolution and an adaptive non-local mean denoising technique are adopted, and input frame preprocessing is completed by combining scale normalization and data augmentation. The core network includes a lightweight ResNet34-CBAM encoder (including 4 residual blocks, a CBAM attention module, and 2 fully connected layers) to extract multi-scale features, a 4-layer bidirectional BiLSTM to capture local temporal dependencies, and 3 layers of MobileViT (including 1 fully connected layer) to replace the Transformer to handle long-distance spatio-temporal correlations. In the feature fusion stage, optical flow features and a 3D convolution module are introduced, and multi-scale spatio-temporal feature fusion is achieved by combining the Hourglass network. During the training process, the dynamic weight GradNorm strategy is adopted to jointly optimize the smoothing loss in the BiLSTM stage, the tracking loss in the Transformer stage, the improved ArcFace angle loss in the Hourglass stage, and the reconstruction loss of the entire process, and finally a stable video sequence with global motion compensation is generated.

[0040] It should be noted that: the above VMTV module has multi-view complementarity: each token fuses the observation perspectives of different views (the occluded area in the left view is visible in the right view), reducing the lack of single-view information. Multi-scale hierarchical: low-resolution features capture global semantics ("rust crack" category), and high-resolution features retain local details ("crack" texture), forming a coarse-to-fine representation. Dynamic alignment ability: cross-view cross-attention and a dynamic weight generator enable the model to adaptively focus on key regions (ROI in object detection, edges in segmentation), which is more efficient than fixed fusion rules (such as averaging, splicing).

[0041] Embodiment 2 This embodiment takes the input of image data with a size of 640*640 as an example to detail the implementation process of a rusty reinforced concrete crack detection method based on YOLOV11: Patch division and embedding (3 parallel branches); patch size (image cell area size) = 16 branches; Number of Patches: =(640 / 16)^2 = 40*40 = 1600; Embedding dimension: = 64 (initial channels); Input image: ; Image cell area embedding (Patch Embed): ; Patch size = 8 branches: ; ; Patch size = 4 branches: ; ; Learnable sine encoding: ; Among them, represents the sine positional encoding function, and represent the positional encoding value of the position calculated by this function in dimension on, indicates that the embedding dimension of the model is 512, represents a constant.

[0042] Generate and add it to : ; Taking the patch size = 16 branch as an example (feature dimension C = 64, sequence length L = 1600), the step-by-step calculation process of 4 layers of Transformer layers is detailed (with the first layer as the core, and subsequent layers follow the same structure, only the input and output are different): Input , (batch dimension = 1, sequence length = 1600, channel dimension = 64), multi-head self-attention (8 heads, head dimension = 64), total dimension D = 8 * 64 = 512) linear transformation to generate Q / K / V: ; Among them, represents the dynamic positional encoding, represents the weight matrix of the query (Query), represents the weight matrix of the key (Key), represents the weight matrix of the value (Value), represents the bias vector of the query, represents the bias vector of the key represents the bias vector of the value.

[0043] ; Split into 8 heads (batch dimension remains unchanged, new head dimension): ; ; ; Among them, The multi-head tensor representing the query (Query). The multi-head tensor representing the key. The multi-head tensor representing the value. Represents the query. Represents the value. Represents the key. Represents dimension reshaping. Represents dimension transposition.

[0044] In the Transformer architecture, to implement the multi-head attention mechanism, the previously computed Q (query tensor), K (key tensor), and V (value tensor) need to be re-partitioned and transformed. Here, the tensor is sliced along a certain dimension to form multiple "heads" so that different heads can capture the relationships between input features from different perspectives.

[0045] Calculate attention similarity (including dynamic position bias Δh ∈ R1600*1600): Among them, and represent the query and key matrices after multi-head splitting. Attn represents the attention weight of the h-th head (Softmax normalization, the dependency distribution between tokens). Represents the attention score matrix of the h-th head (dot product similarity, including scaling and position bias). Represents the scaling factor (the square root of the feature dimension per head, to stabilize the gradient).

[0046] Weighted aggregation of value vectors and concatenation of heads: ; ; Among them, Represents the attention output tensor of the h-th head. Represents the final output tensor of multi-head attention. Represents the attention weight matrix of the h-th head. Represents the value matrix of the h-th head.

[0047] Residual connection and LayerNorm (after self-attention): ; Among them, represents the final output tensor of the self-attention layer (Self-Attention, SA), and LayerNorm represents the layer normalization function that normalizes the last dimension of the tensor.

[0048] Feed-forward network (FFN, hidden dimension = 1024), first fully connected layer (expanding dimensions): ; where represents the weight matrix of the first layer of the FFN, represents the bias vector of the first layer of the FFN, represents the output of the self-attention layer.

[0049] Second fully connected layer (shrinking dimensions): ; ; where represents the final output of the feed-forward network represents the output of the first layer of the feed-forward network.

[0050] Residual connection and LayerNorm (after FFN): ; where represents the final output feature tensor of the Transformer encoder layer.

[0051] Table 1 Complete calculation process table of 4-layer Transformer Cross-scale cross-attention (view Figure 1 at 16 scales → view Figure 2 at 8 scales, taking view N = 2 as an example here) Scale alignment: view Figure 2 features at 8 scales ; Bilinear interpolation to adjust the resolution to 40×40 (consistent with 16 scales): ; 1x1 convolution to reduce the number of channels to 64: ; Cross-attention calculation (12 heads, head dimension = 64): Query ; Key-value ; Split into 12 heads ; Attention output ; where Represents the normalized exponential function.

[0052] After concatenation, CrossAttn (Cross Attention) ∈ R1*1600*768, and after passing through the FFN, the output is FCross (Feature Cross Module) ∈ R1*1600*64; Global semantic alignment feature concatenation: ; Dimension: 1*(1600 + 6400 + 25600)*256 = 1*33600*256; 3-layer global self-attention (8 heads, head dimension = 64): Calculation of self-attention for each layer: ; After the output passes through the residual connection and LayerNorm, the final Vglobal (global semantic alignment feature) ∈ R1*33600*256 Scale dimension fusion dynamic weight generator input concatenation: ; Two-hidden-layer FC: ; Among them, Represents the activation function of the middle layer in the two-hidden-layer fully connected network (FC), Represents the Softmax output; Weighted fusion: ; View dimension fusion (N = 2 views) view feature sequence: ; 3-layer self-attention: Calculation for each layer ; Output ; Cross-view contrastive learning projection head ; Hierarchical dynamic routing local token weights ; Global view weights ; Each Patch of each view contains cross-view and cross-scale global semantic information. The low-resolution Patch (patchsize = 4) encodes the global context, and the high-resolution Patch (patchsize = 16) retains local details (texture, edges); cross-view interaction enables each token to fuse complementary information from other views (occluded regions from different perspectives).

[0053] Contrastive learning features Calculate the cross-view contrastive loss (NT-Xent) to enhance the view invariance of the features.

[0054] View Figure 1 Consistency prediction Generated by 3-layer prediction heads for view Figure 1 Consistency supervision loss (MSE + contrast loss) to ensure semantic consistency of different views.

[0055] Object detection (multi-view fusion): Preserve the spatial location information of features and restore them to multi-scale feature maps (1600 → 40×40, 6400 → 80×80, 25600 → 160×160).

[0056] Generate bounding box coordinates and class probabilities through a detection head (such as YOLOV11): .

[0057] Output result: The position (x, y, w, h) of the object and the class confidence.

[0058] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting cracks in rusty reinforced concrete based on YOLOV11, characterized in that, Including: Obtain the crack image data of the corroded reinforced concrete member, and annotate the crack image data to generate corresponding labels; Combine the annotated image data with the labels into a data set, and divide it into a training set, a validation set and a test set; Construct a YOLOV11 cross-view contrast model, which includes a backbone network, a neck network and a head network; the backbone network successively includes a VMTV module, a C3K2 module, a CSWin_tiny module, an SPPF module and a C2PSA module; wherein the VMTV module includes a view branch module, a cross-view interaction module, a two-stage dynamic fusion module and an auxiliary module; Use the training set to train the YOLOV11 cross-view contrast model, and tune the hyperparameters through the validation set to obtain the best model; The camera collects the crack video data of the corroded reinforced concrete member to be detected in real time during the movement, inputs it into the DVT_FVS module for frame stabilization processing, and then transmits the video into the YOLOV11 cross-view contrast model for crack detection, and outputs the position, size and category information of the cracks.

2. The method for detecting rusty reinforced concrete cracks based on YOLOV11 according to claim 1, characterized in that, The output features of the C2PSA module of the YOLOV11 cross-view contrast model backbone network The features output by the SPPF module of the backbone network are concatenated with the upsampled features of the input neck network, and the concatenated features are processed by the C3K2 module to obtain features , the features After upsampling, the features are concatenated with the features output by the C3K2 module of the backbone network to obtain features , the features are output to the DVT_FVS module of the head network, and the features are also processed by a convolutional layer and concatenated with the features and then input into the C3K2 module to obtain features , the features are output to the DVT_FVS module of the head network, and the features are also processed by a convolutional layer and concatenated with the features and then input into the C3K2 module to obtain features , the features are output to the head network.

3. The method for detecting rusted reinforced concrete cracks based on YOLOV11 according to claim 2, wherein The VMTV module contains: View branch module: Adopt 3 parallel multi-view Transformer branches, each branch contains an image cell area embedding layer, a dynamic position encoding layer, 4 layers of VIT hierarchical Transformer layers, an SE channel attention module, a feed-forward network and a normalized residual block and a cross-scale feature pyramid connection layer; Cross-view interaction module: Contains an intra-view multi-scale self-attention layer, a cross-scale cross-attention layer and a global semantic alignment layer; Two-stage dynamic fusion module: Contains a scale dimension fusion module and a view dimension fusion module; Auxiliary module: Contains hierarchical dynamic routing, cross-view contrast learning module and view consistency supervision module.

4. The method for detecting cracks in corroded reinforced concrete based on YOLOV11 according to claim 3, wherein, In the view branch module: The input image is divided into corresponding-sized image cell areas through the image cell area embedding layer and embedding vectors are generated; The learnable sine encoding of the embedding vector is concatenated with the local position deviation matrix through the dynamic position encoding layer and linearly projected to the feature dimension; Then it is processed successively through 4 layers of VIT hierarchical Transformer layers, an SE channel attention module, a feed-forward network and a normalized residual block; the VIT hierarchical Transformer layer contains an 8-head self-attention layer; The cross-scale feature pyramid connection layer concatenates the three-sized features output by the normalized residual block to generate a single-view multi-scale feature, specifically by upsampling the low-scale feature and concatenating it with the current-scale feature to generate a single-view multi-scale feature.

5. The method for detecting cracks in corroded reinforced concrete based on YOLOV11 according to claim 3, characterized in that, In the cross-view interaction module: The intra-view multi-scale self-attention layer includes 4 layers of VIT hierarchical Transformer layers, an SE channel attention module, a feed-forward network and a normalized residual block, and each layer of the VIT hierarchical Transformer layer contains an 8-head self-attention layer; The cross-scale cross-attention layer includes a bilinear interpolation layer, a convolutional layer, a 12-head cross-attention layer and a feed-forward network; The global semantic alignment layer includes a feature concatenation layer and three layers of global self-attention layers. The feature concatenation layer flattens the multi-scale features of all views into a unified sequence, and then adds view identity embeddings and scale identity embeddings to mark the feature sources. The global self-attention layer includes an eight-head self-attention layer, a feed-forward network, and a dynamic position encoding layer.

6. The method for detecting rusty reinforced concrete cracks based on YOLOV11 according to claim 3, characterized in that, In the two-stage dynamic fusion module: The scale dimension fusion module includes a feature concatenation layer, a double-hidden-layer FC weight generator, and a weighted fusion layer; The view dimension fusion module serializes the scale fusion feature of each view into tokens, and then uses three layers of VIT-style self-attention layers to complete cross-view feature fusion. The VIT-style self-attention layer includes an eight-head self-attention layer, a feed-forward network, layer normalization, and residual connections.

7. The method for detecting rusted reinforced concrete cracks based on YOLOV11 according to claim 3, wherein, In the auxiliary module: The hierarchical dynamic routing generates weights for local tokens through one layer of fully-connected layers and generates weights for global views through one layer of fully-connected layers; The cross-view contrast learning module calculates the NT-Xent contrast loss using a two-layer fully-connected projection head; The view consistency supervision module generates prediction values by combining the MSE loss and the contrast loss through a three-layer MLP prediction head.

8. The method for detecting rusty reinforced concrete cracks based on YOLOV11 according to claim 2, characterized in that, The DVT_FVS module includes: The preprocessing module: uses blind deconvolution motion blur kernel estimation and adaptive non-local mean denoising, combines scale normalization and data augmentation to complete the preprocessing of the input frames; The feature fusion module: embeds the optical flow into a latent representation through a ResNet34_CBAM encoder, and connects it with the poses of the real and virtual cameras; infers a new virtual camera pose through a BiLSTM+MobileViT block, and uses this virtual camera pose to generate a warping grid; Then, spatio-temporal feature fusion is achieved through a 3D convolutional module and an Hourglass network to predict the new virtual camera pose as a quaternion; the input frames are warped based on optical image stabilization and the virtual camera pose to generate stable frames.

9. The method for detecting rusted reinforced concrete cracks based on YOLOV11 according to claim 8, wherein, The training process of the DVT_FVS module adopts: the dynamic weight GradNorm strategy, and cooperatively optimizes the smooth loss in the BiLSTM stage, the tracking loss in the Transformer stage, the improved ArcFace angle loss in the Hourglass stage, and the reconstruction loss of the whole process.

10. A rusty reinforced concrete crack detection device based on YOLOV11, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the YOLOV11-based rusty reinforced concrete crack detection method according to any one of claims 1-9 when running the program instructions.

Citation Information

Patent Citations

  • Tunnel disease detection method based on multi-scale feature pyramid

    CN118781077A

  • Improved YOLOv11 crack identification and quantification method for strip mine area and slope thereof

    CN119540722A

  • Improved sewer internal defect detection method and system based on YOLOv11

    CN119540725A

  • Bridge multi-mode multi-target disease intelligent identification method and device under complex background and medium

    CN120071007A

Cited By

  • Soil moisture content prediction method and system based on BiLSTM-Transform dynamic weight hybrid architecture

    CN120597733A

  • Spatial perception multi-view anomaly detection method based on meta-view representation

    CN121074512A

  • Two-stage concrete crack segmentation method and system considering complex illumination conditions

    CN121258988A

  • Two-stage concrete crack segmentation method and system considering complex lighting conditions

    CN121258988B

  • Pavement pit and small roadblock detection method based on YOLOv8

    CN121353873A