Multi-task distillation-driven concrete bridge apparent disease detection and segmentation method

Through the multi-task distillation-driven method, combined with feature extraction module, detection network and segmentation network, knowledge migration loss and total mixing loss are used to improve the accuracy and efficiency of concrete bridge defect detection, and solve the problem of poor applicability of student models in multi-task scenarios.

CN120259175APending Publication Date: 2025-07-04KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510203971.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, student models cannot effectively extract in-depth and abstract features in concrete bridge defect detection, resulting in low detection accuracy and poor applicability in multi-task scenarios.

Method used

Using a multi-task distillation-driven method, the student model includes a sequential feature extraction module, a detection network and a segmentation network, combining knowledge migration loss and total mixing loss, the image features are efficiently extracted using the convolutional layer and the P_block module, and defect detection and crack segmentation are performed.

Benefits of technology

It significantly improves detection accuracy and computing efficiency, enhances model performance, and solves the problem of poor applicability of traditional student models in multi-task detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259175A_ABST
    Figure CN120259175A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task distillation-driven concrete bridge apparent disease detection and segmentation method, and the method comprises the steps: obtaining a to-be-detected concrete bridge image, inputting the concrete bridge image into a trained student model, so as to obtain a defect detection result image and / or a crack segmentation result image, the student model at least comprises a feature extraction module, a detection network and / or a segmentation network which are sequentially connected in series. According to the student model provided by the invention, the feature extraction module comprising the convolutional layer and the Pblock module is adopted, so that the image features can be efficiently extracted, and meanwhile, the parameter quantity is remarkably reduced. According to the design, the memory access times of the model can be effectively reduced, so that the calculation efficiency is improved, the model performance is enhanced, and the detection precision is remarkably improved. In addition, the student model can execute multiple tasks at the same time, and the problem that the student model is difficult to detect in a multi-task scene can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of detection technologies, and particularly to a method for detecting and segmenting the apparent diseases of a concrete bridge driven by multi-task distillation. Background Art

[0002] Concrete bridges are an important part of transportation infrastructure. With the increase of the service life, the structures of concrete bridges may have different degrees of defects, such as cracks, peeling, corrosion, etc. These defects not only affect the bearing capacity and service life of the bridges, but also directly threaten traffic safety. Therefore, it is particularly important to detect the surface defects of concrete bridges.

[0003] At present, the distillation learning strategy has been applied to the task of detecting concrete bridge defects. Traditional distillation learning methods mainly guide the learning of the student model by imitating the output (soft labels or probability distributions) of the teacher model, and the student model may not necessarily directly learn higher-level features or abstract knowledge from the teacher model. Although the student model can achieve certain performance when learning the output of the teacher model, due to the learning process of the student model being limited by the ability and architecture of the teacher model, the student model may not be able to extract more in-depth and abstract features from the teacher model, resulting in a lower accuracy of the detection results.

[0004] Based on this, there is an urgent need to provide a method for detecting and segmenting the apparent diseases of a concrete bridge driven by multi-task distillation to improve the accuracy of detecting the surface defects of concrete bridges.

[0005] The above content is only used to assist in understanding the technical solution of the present application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] Embodiments of the present application provide a method for detecting and segmenting the apparent diseases of a concrete bridge driven by multi-task distillation, aiming to solve the technical problem of the low accuracy of detecting the surface defects of concrete bridges. To achieve the above object, embodiments of the present application provide a method for detecting and segmenting the apparent diseases of a concrete bridge driven by multi-task distillation, including:

[0007] Obtain a concrete bridge image to be detected;

[0008] Input the concrete bridge image into a trained student model to obtain a defect detection result image and / or a crack segmentation result image; the student model at least includes a feature extraction module, a detection network, and / or a segmentation network connected in series in sequence; wherein,

[0009] The feature extraction module is used to extract image features from the concrete bridge image;

[0010] The detection network is used to identify defects in the concrete bridge image to generate the defect detection result image;

[0011] The segmentation network is used to identify cracks from the concrete bridge image and segment the area where the cracks are located to generate the crack segmentation result image; where

[0012] The feature extraction module is sequentially composed of a 3X3 Conv1 layer, a 3X3 Conv2 layer, a P_block1 module, a 3X3 Conv3 layer, a P_block2 module, a 3X3 Conv4 layer, a P_block3 module, a 3X3 Conv5 layer, and a P_block4 module in series. Among them, the stride of the 3X3 Conv1 layer to the 3X3 Conv5 layer is 2, and the number of channels starts from 2^4 and increases in powers of 2 in sequence;

[0013] The P_block1 module to the P_block4 module are each sequentially composed of a first branch, a second branch in parallel, and a first splicing module respectively connected to the outputs of the first branch and the second branch; among them, the first branch is sequentially composed of a 3X3 PConv layer, a BN layer, an h_swish activation function, an SE attention layer, a 1×1Conv6 layer, and a BN layer in series, and the second branch is used to receive the input and transmit the input to the first splicing module.

[0014] Optionally, the detection network is sequentially composed of an SPPELAN module, a second splicing module, a first upsampling module, a third splicing module, a 3X3 Conv7 layer, a first RepNCSPELAN4 module, a second upsampling module, a fourth splicing module, a second RepNCSPELAN4 module, and a detection module in series. Among them, the output of the first RepNCSPELAN4 module is also connected to the input of the detection module; the upsampling multiples of the first upsampling module and the second upsampling module are both 2; the stride of the 3X3Conv7 layer is 1, and the number of channels is 256; the splicing dimensions of the second splicing module to the fourth splicing module are all 1; where

[0015] The output of the P_block4 module in the feature extraction module is respectively connected to the input of the SPPELAN module and the input of the second splicing module, the output of the P_block3 module is connected to the input of the third splicing module, and the output of the P_block2 module is connected to the input of the fourth splicing module.

[0016] Optionally, the SPPELAN module consists of a 1X1 Conv8 layer, an SP1 layer, an SP2 layer, an SP3 layer connected in series in sequence, and a first Concat layer connected to the outputs of the 1X1 Conv8 layer, the SP1 layer, the SP2 layer, and the SP3 layer respectively, and a 1X1Conv9 layer connected to the output of the first Concat layer; the number of channels of the 1X1 Conv8 layer and the 1X1 Conv9 layer is 256, and the Kernel_size of the SP1 layer, the SP2 layer, and the SP3 layer in the SPPELAN module is 5, Stride = 1, Padding = 2 。

[0017] Optionally, both the first RepNCSPELAN4 module and the second RepNCSPELAN4 module consist of a 1X1 Conv10 layer, a Spilt layer, a first RepNCSP module, a 3X3 Conv11 layer connected in series in sequence, a 3X3 Conv12 layer connected to the output of the Spilt layer, a second RepNCSP module connected to the output of the 3X3 Conv12 layer, a second Concat layer connected to the outputs of the Spilt layer, the 3X3 Conv11 layer, and the 3X3 Conv12 layer respectively, and a 1X1 Conv13 layer connected to the output of the second Concat layer, where the output of the second RepNCSP module is also connected to the input of the first RepNCSP module, and the Spilt layer splits the input into two parts and transmits them to the second Concat layer respectively; the stride of the 1X1 Conv10 layer and the 1X1 Conv13 layer is 1, and the padding size is 0; the stride of the 3X3Conv11 layer and the 3X3 Conv12 layer is 1, and the padding size is 1.

[0018] Optionally, both the first RepNCSP module and the second RepNCSP module consist of a third branch and a fourth branch connected in parallel in sequence, and a third Concat layer connected to the outputs of the third branch and the fourth branch respectively, and a 1X1Conv14 layer connected to the output of the third Concat layer, where the stride of the 1X1Conv14 layer is 1 and the padding size is 0; the third branch consists of a 1X1 Conv15 layer and N RepNBottLeneck modules connected in series in sequence, and the fourth branch consists of a 1X1Conv16 layer; where the stride of the 1X1 Conv15 layer and the 1X1 Conv16 layer is 1, and the padding size is 0.

[0019] Optionally, the RepNBottLeneck module is successively composed of a fifth branch and a sixth branch connected in parallel; wherein the fifth branch is successively composed of a RepConvN module and a 1X1Conv17 layer connected in series, the stride of the 1X1Conv17 layer is 1, and the padding size is 0; the sixth branch is used to receive and transmit the input.

[0020] Optionally, the segmentation network is successively composed of a third upsampling module, a 3X3 DWConv1 layer, a fourth upsampling module, a 3X3 DWConv2 layer, a fifth upsampling module, a 3X3 DWConv3 layer, a sixth upsampling module, a C2f module, and a 3X3DWConv4 layer connected in series; wherein the upsampling multiples of the third upsampling module to the sixth upsampling module are all 2; the number of channels of the 3X3 DWConv1 layer to the 3X3DWConv3 layer starts from 2^6 and decreases successively by a power of 2, and the strides are 2, 1, 1 in turn, the number of channels of the 3X3DWConv4 layer is 4, and the stride is 1; the number of channels of the C2f module is 16; wherein,

[0021] The output of the second RepNCSPELAN4 module in the detection network is connected to the input of the third upsampling module.

[0022] Optionally, before the step of inputting the concrete bridge image into the trained student model to obtain a defect detection result image and / or a segmentation result image, the following steps are further included:

[0023] Performing a preprocessing operation on the concrete bridge image; wherein,

[0024] The preprocessing operation includes at least one of the following steps:

[0025] Performing normalization processing on the concrete bridge image, and the normalization processing at least includes dividing the pixels of the concrete bridge image by 255 and scaling the pixel values to between 0 and 1;

[0026] Performing at least one of the operations of enhancing the concrete bridge image, adjusting the brightness, enhancing the contrast, and rotating the image;

[0027] Performing a fusion operation on the concrete bridge image; wherein,

[0028] The fusion operation successively includes the following steps:

[0029] Randomly selecting four different images from the concrete bridge image;

[0030] Resizing the side length of the image to 640 and keeping the image ratio unchanged;

[0031] Crop the four different images with a crosshair at a random position in sequence;

[0032] Stitch the cropped images together to obtain the fused concrete bridge image;

[0033] Update the bbox coordinates of the fused concrete bridge image.

[0034] Optionally, the method further includes:

[0035] Obtain the dataset to be trained;

[0036] Perform the preprocessing operation on the dataset;

[0037] Input the dataset into a preset teacher model and the student model, and train the student model from the teacher model by the knowledge distillation method to obtain the trained student model.

[0038] Optionally, the student model includes a knowledge transfer loss and a total mixing loss; where

[0039] The knowledge transfer loss of the student model is:

[0040]

[0041] where m represents the batch size, j represents the index of the feature dimension, d represents the number of classes, i represents the i-th class, ε is a constant, P represents the prediction result of the student model, Q represents the output of the teacher model, and Q is used as the soft label to guide the learning of the student model;

[0042] where, if ρ and σ are set to 1, the knowledge transfer loss is:

[0043]

[0044] where and are the knowledge transfer losses of the object detection task and the pixel segmentation task respectively. The total mixing loss of the student model is:

[0045] L total = a(L det + L seg ) + (1 - a)L dil

[0046] where L det 、L seg and L dil represent the object detection loss, the pixel segmentation loss, and the knowledge transfer loss respectively.

[0047] A method for detecting and segmenting apparent diseases of concrete bridges driven by multi-task distillation. By obtaining the concrete bridge images to be detected and inputting the concrete bridge images into the trained student model, a defect detection result image and / or a crack segmentation result image can be obtained. The student model at least includes a feature extraction module, a detection network, and / or a segmentation network connected in series in sequence. The student model proposed in this application can efficiently extract image features and significantly reduce the number of parameters by adopting a feature extraction module including a convolutional layer and a P_block module. This design effectively reduces the number of model memory accesses, thereby improving the calculation efficiency, enhancing the model performance, and significantly improving the accuracy of the detection results. In addition, by setting the knowledge transfer loss and the total mixing loss of the student model and training multiple tasks of defect detection and crack segmentation from the teacher model through the distillation learning method, the defect detection and crack analysis capabilities of the student model can be improved, effectively solving the problem of poor applicability and difficult detection of traditional student models when dealing with multi-task problems. Description of the Drawings

[0048] Figure 1 It is a flowchart of the steps of a method for detecting and segmenting apparent diseases of concrete bridges driven by multi-task distillation proposed in this application;

[0049] Figure 2 It is a structural schematic diagram of the feature extraction module of the student model proposed in this application;

[0050] Figure 3 It is a structural schematic diagram of the P_block module proposed in this application;

[0051] Figure 4 It is a structural schematic diagram of the student model including a feature extraction module and a detection network proposed in this application;

[0052] Figure 5 It is a structural schematic diagram of the SPPELAN module proposed in this application;

[0053] Figure 6 It is a structural schematic diagram of the RepNCSPELAN4 module proposed in this application;

[0054] Figure 7 It is a structural schematic diagram of the student model including a feature extraction module, a detection network, and a segmentation network proposed in this application;

[0055] Figure 8 It is a schematic diagram of image cropping proposed in this application;

[0056] Figure 9 It is a flowchart of the steps of a method for detecting and segmenting apparent diseases of concrete bridges driven by multi-task distillation proposed in this application;

[0057] Figure 10 It is a training schematic diagram of the student model proposed in this application.

[0058] The realization of the purpose of this application, functional features and advantages will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments

[0059] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0060] To better understand the above technical solutions, the exemplary embodiments of this application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.

[0061] For ease of understanding, the following is an explanation of the English abbreviations used in the embodiments of this application:

[0062] Conv (Convolution, convolution);

[0063] PConv (Pointwise Convolution, point convolution);

[0064] DWConv (Depthwise Convolution, depth convolution);

[0065] Conv / PConv / DWConv(k, s, p), where k represents the convolution kernel size, s represents the stride, and p represents the padding size;

[0066] RepConvN (Reparameterization Convolution, reparameterization convolution), N represents the number of RepConv;

[0067] P_block (PortableNet_block, portable network module);

[0068] Concat (concatenate, splicing), in this application, Cat and Concat are synonymous and both refer to splicing;

[0069] SPPELAN (Spatial Pyramid Pooling Enhanced with ELAN, spatial pyramid pooling enhanced with ELAN);

[0070] Upsample (upsampling, abbreviated as Up);

[0071] RepNCSPELAN4 (Reparameterization Convolutional Spatial Pyramid Encoder - Decoder with Attention and Normalization, reparameterized CSPELAN), where N represents the number of CSPELANs;

[0072] Detect;

[0073] SP (Spatial Pyramid Pooling, also simply referred to as SPP); SP(Kernel_size, Stride, Padding), where Kernel_size represents the kernel size of SP, Stride represents the stride of SP, and Padding represents the padding pixel value of SP;

[0074] C2f (Cascaded Convolutional Features);

[0075] Spilt (split operation);

[0076] RepNCSP (Reparameterization Cross - Stage Partial, reparameterized CSP module), where N represents the number of CSPs;

[0077] RepNBottLeneck (Reparameterized Bottleneck Layer), where N represents the number of BottLenecks.

[0078] Example 1

[0079] Refer to Figure 1 , Figure 1 is a flowchart of the steps of a multi - task distillation - driven method for detecting and segmenting the apparent diseases of concrete bridges proposed in this application. A multi - task distillation - driven method for detecting and segmenting the apparent diseases of concrete bridges proposed in this application includes step S10 and step S20.

[0080] Step S10: Obtain the concrete bridge image to be detected;

[0081] In some embodiments, the concrete bridge to be detected can be photographed by a camera device to obtain the concrete bridge image to be detected. Among them, a camera device with high resolution can be selected to obtain a concrete bridge image with high resolution, thereby improving the accuracy of the detection result.

[0082] Step S20: Input the concrete bridge image into the trained student model to obtain a defect detection result image and / or a crack segmentation result image.

[0083] In this embodiment, the present application proposes a new student model that can be used for both defect detection and / or crack segmentation of concrete bridge images, which means that this student model can not only perform a single task of defect detection, but also perform a dual task of defect detection and crack segmentation simultaneously. This innovative student model can handle multiple tasks simultaneously, effectively solving the problem of poor applicability and difficult detection of traditional student models when dealing with multi-task problems.

[0084] In some embodiments, the student model proposed in the present application at least includes a feature extraction module, a detection network, and / or a segmentation network connected in series in sequence. The feature extraction module is used to extract image features from the input concrete bridge image. The detection network is used to identify defects in the concrete bridge image to generate a defect detection result image. The segmentation network is used to identify cracks in the concrete bridge image and segment the area where the cracks are located to generate a crack segmentation result image. The defects identified by the detection network include potential structural defects of the concrete bridge, such as cracks, holes, corrosion, spalling, etc.

[0085] In some embodiments, referring to Figure 2 , Figure 2 is a schematic structural diagram of the feature extraction module of the student model proposed in the present application. As shown in Figure 2 , the feature extraction module is sequentially composed of a 3X3 Conv1 layer, a 3X3 Conv2 layer, a P_block1 module, a 3X3Conv3 layer, a P_block2 module, a 3X3 Conv4 layer, a P_block3 module, a 3X3 Conv5 layer, and a P_block4 module connected in series. Among them, the stride of the 3X3 Conv1 layer to the 3X3 Conv5 layer is 2, and the number of channels starts from 2^4 and increases in powers of 2 in sequence. It should be noted that N in ConvN only represents the number of Conv, which is used to distinguish different convolutional layers. Similarly, N in P_blockN only represents the number of P_block, which is used to distinguish different P_block modules.

[0086] In some embodiments, the P_block1 module to the P_block4 module have the same composition structure. They are all P_block modules, only with different numbers to represent different module instances. Referring to Figure 3 , Figure 3 is a schematic structural diagram of the P_block module proposed in the present application. As shown in Figure 3As shown, the P_block module consists of a first branch, a second branch in parallel, and a first splicing module (corresponding to the Cat of Figure 3 ) connected to the outputs of the first branch and the second branch in sequence. Among them, the first branch is composed of a 3X3 PConv layer, a BN layer, an h_swish activation function, an SE attention layer, a 1×1Conv6 layer, and a BN layer connected in series. The second branch is used to receive the input and transmit the input to the first splicing module. It can be understood that the second branch has no modules and does not perform any processing on the input. After receiving the input, it directly transmits the input to the first splicing module. Figure 3 The second branch is used to receive the input and transmit the input to the first splicing module. It can be understood that the second branch has no modules and does not perform any processing on the input. After receiving the input, it directly transmits the input to the first splicing module.

[0087] In this embodiment, the student model uses this feature extraction module for feature extraction, which can greatly reduce the number of parameters of the model, thereby reducing the number of memory accesses of the model, enhancing the model performance, and achieving the purpose of improving the accuracy of the detection result.

[0088] Example 1, represent the concrete bridge image to be detected as F1, and its image size is 640x640x3. After F1 is input into the student model, it first undergoes feature extraction by the feature extraction module. The processing flow of the feature extraction module for the input image F1 is as follows:

[0089] 1. Use a 3x3 Conv1 to perform a convolution operation on the input image F1 with a scale of 640x640x3, a stride of 2, and 16 channels, to obtain an output feature map F2 with a scale of 320×320×16.

[0090] 2. Use a 3x3 Conv2 to perform a convolution operation on the feature map F2, with a stride of 2 and 32 channels, to obtain an output feature map F3 with a scale of 160×160×32.

[0091] 3. Use the P_block1 module to process the feature map F3: In the first branch, the feature map F3 undergoes a 3×3 PConv convolution operation, normalization by the BN layer, activation by the h_swish activation function, SE attention mechanism processing, a 1×1Conv6 convolution operation, and normalization by the BN layer, to obtain an output feature map F3a, which is then transmitted to the first splicing module Cat. In the second branch, the feature map F3 is transmitted to the first splicing module Cat. Finally, the first splicing module Cat splices the feature map F3a output by the first branch and the feature map F3 output by the second branch to obtain an output feature map F4 with a scale of 160×160×32.

[0092] 4. Use a 3x3 Conv3 to perform a convolution operation on the feature map F4, with a stride of 2 and 64 channels, to obtain an output feature map F5 with a scale of 80×80×64.

[0093] 5. Process the feature map F5 using the P_block2 module to obtain the output feature map F6 with a scale of 80×80×64. The processing flow of the P_block2 module is the same as that of the P_block1 module, which will not be elaborated here.

[0094] 6. Perform a convolution operation on the feature map F6 using a 3x3 Conv4 with a stride of 2 and 128 channels to obtain the output feature map F7 with a scale of 40×40×128.

[0095] 7. Process the feature map F7 using the P_block3 module to obtain the output feature map F8 with a scale of 40×40×128. Similarly, the processing flow of the P_block3 module is the same as that of the P_block1 module, which will not be elaborated here.

[0096] 8. Perform a convolution operation on the feature map F8 using a 3x3 Conv5 with a stride of 2 and 256 channels to obtain the output feature map F9 with a scale of 20×20×256.

[0097] 9. Process the feature map F9 using the P_block4 module to obtain the output feature map F10 with a scale of 20×20×256. Similarly, the processing flow of the P_bloc4 module is the same as that of the P_block1 module, which will not be elaborated here.

[0098] In the technical solution provided in this embodiment, by obtaining a concrete bridge image to be detected and inputting the concrete bridge image into a trained student model to obtain a defect detection result image and / or a crack segmentation result image, the student model at least includes a feature extraction module, a detection network, and / or a segmentation network connected in series in sequence. The student model proposed in this application can perform multiple task detections. By adopting a feature extraction module including a convolutional layer and a P_block module, it can efficiently extract image features and significantly reduce the number of parameters. This design effectively reduces the number of model memory accesses, thereby improving the calculation efficiency, enhancing the model performance, and significantly improving the accuracy of the detection results.

[0099] In some embodiments, the detection network is sequentially composed of an SPPELAN module, a second splicing module, a first upsampling module, a third splicing module, a 3X3 Conv7 layer, a first RepNCSPELAN4 module, a second upsampling module, a fourth splicing module, a second RepNCSPELAN4 module, and a detection module connected in series. The output of the first RepNCSPELAN4 module is also connected to the input of the detection module. The upsampling multiples of the first upsampling module and the second upsampling module are both 2. The stride of the 3X3 Conv7 layer is 1, and the number of channels is 256. The splicing dimensions of the second splicing module to the fourth splicing module are all 1. Among them, the output of the P_block4 module in the feature extraction module is respectively connected to the input of the SPPELAN module and the input of the second splicing module, the output of the P_block3 module is connected to the input of the third splicing module, and the output of the P_block2 module is connected to the input of the fourth splicing module. In addition, the detection module is used to perform a Detect operation on the output images of the first RepNCSPELAN4 module and the second RepNCSPELAN4 module to obtain a defect detection result image. The specific implementation manner of the Detect operation in this embodiment is not limited, and those skilled in the art can flexibly select according to the existing technology, such as a method based on bounding box regression.

[0100] Specifically, referring to Figure 4 , Figure 4 is a schematic structural diagram of the student model including a feature extraction module and a detection network proposed in this application. As Figure 4 shown, the detection network is sequentially composed of SPPELAN, Cat2 (the second splicing module), Upsample1 (the first upsampling module), Cat3 (the third splicing module), Conv7, RepNCSPELAN4_1 (the first RepNCSPELAN4 module), Upsample2 (the second upsampling module), Cat4 (the fourth splicing module), RepNCSPELAN4_2 (the second RepNCSPELAN4_2 module), and Detect (the detection module) connected in series. Figure 4 Detect is not shown in

[0101] Figure 5 Figure 5 , Figure 5 is a schematic structural diagram of the SPPELAN module proposed in this application. As Figure 5As shown, the SPPELAN module consists of a 1X1 Conv8 layer, an SP1 layer, an SP2 layer, an SP3 layer in series, and a first Concat layer (corresponding to Concat1 in Figure 5 ) connected to the outputs of the 1X1 Conv8 layer, the SP1 layer, the SP2 layer, and the SP3 layer, and a 1X1 Conv9 layer connected to the output of the first Concat layer. The number of channels of the 1X1 Conv8 layer and the 1X1 Conv9 layer is 256. In the SPPELAN module, the Kernel_size (kernel size) of the SP1 layer, the SP2 layer, and the SP3 layer is 5, the Stride (stride) is 1, and the Padding (padding pixel value) is 2. It should be noted that N in SPN only represents the number of SPs, used to distinguish different SP instances. In fact, the composition structures of the SP1 layer, the SP2 layer, and the SP3 layer are the same.

[0102] Furthermore, referring to Figure 6 , Figure 6 is the structural schematic diagram of the RepNCSPELAN4 module proposed in this application. As shown in figure (a) of Figure 6 , the composition structures of the first RepNCSPELAN4 module and the second RepNCSPELAN4 module are the same, and both consist of a 1X1 Conv10 layer, a Spilt layer, a first RepNCSP module, a 3X3 Conv11 layer in series, a 3X3 Conv12 layer connected to the output of the Spilt layer, a second RepNCSP module connected to the output of the 3X3Conv12 layer, a second Concat layer (corresponding to Concat2 in figure (a)) connected to the outputs of the Spilt layer, the 3X3Conv11 layer, and the 3X3 Conv12 layer, and a 1X1 Conv13 layer connected to the output of the second Concat layer. The output of the second RepNCSP module is also connected to the input of the first RepNCSP module. The Spilt layer splits the input into two parts and transmits them separately. One part is passed to the second Concat layer after subsequent processing, and the other part is directly passed to the second Concat layer without processing. The stride of the 1X1Conv10 layer and the 1X1 Conv13 layer is 1, and the padding size is 0. The stride of the 3X3 Conv11 layer and the 3X3 Conv12 layer is 1, and the padding size is 1.

[0103] It can be understood that the Split layer splits the input into two parts and then transmits them separately. One part is passed to the second Concat layer after subsequent processing, and the other part is directly passed to the second Concat layer without processing.

[0104] As shown in Figure 6As shown in Figure (b), the first RepNCSP module and the second RepNCSP module have the same composition structure. They are both composed of a third branch and a fourth branch in parallel, a third Concat layer (corresponding to Concat3 in Figure (b)) connected to the outputs of the third branch and the fourth branch respectively, and a 1X1Conv14 layer connected to the output of the third Concat layer. The stride of the 1X1Conv14 layer is 1, and the padding size is 0. The third branch is sequentially composed of a 1X1 Conv15 layer and N RepNBottLeneck modules connected in series. The fourth branch is composed of a 1X1 Conv16 layer. The strides of the 1X1 Conv15 layer and the 1X1 Conv16 layer are both 1, and the padding sizes are both 0.

[0105] As Figure 6 shown in Figure (c), the RepNBottLeneck module is sequentially composed of a fifth branch and a sixth branch in parallel. The fifth branch is sequentially composed of a RepConvN module and a 1X1Conv17 layer connected in series, and the stride of the 1X1Conv17 layer is 1, and the padding size is 0. The sixth branch is used to receive and transmit the input. It can be understood that the sixth branch has no modules and does not perform any processing on the input. After receiving the input, it directly transmits the input to the next module. It should be noted that N in RepNBottLeneck represents the number of BottLenecks, and N in RepConvN represents the number of RepConvs.

[0106] Example 2: Taking the concrete bridge image F1 to be detected in Example 1 as an example, after the input image F1 is subjected to feature extraction by the feature extraction module, it enters the detection network for defect detection. The processing flow of the feature extraction module can refer to Example 1 and will not be elaborated here. The processing process of the detection network is as follows:

[0107] 1. Use the SPPELAN module to perform operations on the feature map F10 output by the P_block4 module in the feature extraction module. The number of channels is 256, and the output feature map F11 with a scale of 20×20×256 is obtained.

[0108] 2. Use the second splicing module to splice F11 and F10 to obtain the output feature map F12 with a scale of 20×20×512. The concatenate function based on PyTorch is: torch.cat((F10, F11), 1).

[0109] 3. Use the first upsampling module to perform upsampling on the feature map F12. The upsampling factor is 2, and the output feature map F13 with a scale of 40x40x512 is obtained.

[0110] 4. Use the third splicing module to splice the feature map F13 and the feature map F8 output by the P_block3 module in the feature extraction module to obtain an output feature map F14 with a scale of 40×40×640. The concatenate function based on PyTorch is: torch.cat((F8, F13), 1).

[0111] 5. Perform a convolution operation on F14 using a 3x3 Conv7 with a stride of 1 and 256 channels to obtain an output feature map F15 with a scale of 40×40×256.

[0112] 6. Use the first RepNCSPELAN4 module to operate on the feature map F15 with 256 channels to obtain an output feature map F16 with a scale of 40×40×256. The processing flow of the RepNCSPELAN4 module can be referred to Figure 6 (Figure (a), which will not be elaborated here.)

[0113] 7. Use the second upsampling module to upsample the feature map F16 with an upsampling factor of 2 to obtain an output feature map F17 with a scale of 80×80×256.

[0114] 8. Use the fourth splicing module to splice the feature map F17 and the feature map F6 output by the P_block2 module in the feature extraction module to obtain an output feature map F18 with a scale of 80×80×320. The concatenate function based on PyTorch is: torch.cat((F6, F17), 1).

[0115] 9. Use the second RepNCSPELAN4 module to operate on the feature map F18 to obtain an output feature map F19 with a scale of 80×80×512. The processing flow of the second RepNCSPELAN4 module is the same as that of the first RepNCSPELAN4 module, which will not be elaborated here.

[0116] 10. Use the detection module to perform a Detect operation on the feature map F16 output by the first RepNCSPELAN4 module and the feature map F19 output by the second RepNCSPELAN4 module to obtain a defect detection result image, which shows the location of the defect.

[0117] In some embodiments, the segmentation network is sequentially composed of a third upsampling module, a 3X3 DWConv1 layer, a fourth upsampling module, a 3X3 DWConv2 layer, a fifth upsampling module, a 3X3 DWConv3 layer, a sixth upsampling module, a C2f module, and a 3X3 DWConv4 layer in series. Specifically, the upsampling multiples of the third to sixth upsampling modules are all 2. The number of channels of the 3X3 DWConv1 layer to the 3X3 DWConv3 layer starts from 2^6 and decreases successively by powers of 2, and the strides are 2, 1, 1 in sequence. The number of channels of the 3X3 DWConv4 layer is 4, and the stride is 1 。 The number of channels of the C2f module is 16. The output of the second RepNCSPELAN4 module in the detection network is connected to the input of the third upsampling module, and the output of the second RepNCSPELAN4 module in the detection network is used as the input of the segmentation network.

[0118] Specifically, referring to Figure 7 , Figure 7 is a schematic structural diagram of the student model including a feature extraction module, a detection network, and a segmentation network proposed in this application. As Figure 7 shown, the segmentation network is sequentially composed of Upsample3 (the third upsampling module), a 3X3 DWConv1 layer, Upsample4 (the fourth upsampling module), a 3X3 DWConv2 layer, Upsample5 (the fifth upsampling module), a 3X3 DWConv3 layer, Upsample6 (the sixth upsampling module), a C2f module, and a 3X3 DWConv4 layer in series. It can be understood that the result output by the 3X3 DWConv4 layer is the crack segmentation result image.

[0119] In this embodiment, by constructing a detection network and a segmentation network, the student model can be used to simultaneously perform defect detection and / or crack segmentation on concrete bridge images, which means that the student model can not only perform a single task of defect detection, but also perform dual tasks of defect detection and crack segmentation at the same time. This innovative student model can handle multiple tasks simultaneously and can effectively solve the problem of poor applicability and difficult detection of traditional student models when dealing with multi-task problems.

[0120] Example 3: Taking the input image F1 in Example 1 as an example, the input image F1 is subjected to feature extraction by the feature extraction module. Then it is subjected to defect detection by the detection network. Finally, it is subjected to crack segmentation by the segmentation network. The processing flow of the feature extraction module can refer to Example 1, and the processing flow of the detection network can refer to Example 2, which will not be elaborated here. The processing flow of the segmentation network is as follows:

[0121] 1. Use the third upsampling module to perform upsampling on the feature map F output by the second RepNCSPELAN4 module in the detection network 19 with an upsampling factor of 2 to obtain an output feature map F20 with a scale of 160×160×512.

[0122] 2. Use a 3×3 DWConv1 to perform depth convolution on the feature map F20 with 64 channels and a stride of 2 to obtain an output feature map F21 with a scale of 80×80×64.

[0123] 3. Use the fourth upsampling module to perform upsampling on the feature map F21 to obtain an output feature map F22 with a scale of 160×160×64.

[0124] 4. Use a 3×3 DWConv2 to perform depth convolution on the feature map F22 with 32 channels and a stride of 1 to obtain an output feature map F23.

[0125] 5. Use the fifth upsampling module to perform upsampling on the feature map F23 to obtain an output feature map F24 with a scale of 320×320×32.

[0126] 6. Use a 3×3 DWConv3 to perform depth convolution on the feature map F24 with 16 channels and a stride of 1 to obtain an output feature map F25 with a scale of 320×320×16.

[0127] 7. Use the sixth upsampling module to perform upsampling on the feature map F25 to obtain an output feature map F26.

[0128] 8. Use the C2f module to operate on the feature map F26 with 16 channels to obtain an output feature map F27 with a scale of 640×640×16.

[0129] 9. Use a 3x3 DWConv4 to perform depth convolution on the feature map F26 to obtain a crack segmentation result image with a scale of 3.

[0130] Example 2

[0131] In some embodiments, before the step of inputting the concrete bridge image into the trained student model to obtain a defect detection result image and / or a segmentation result image, it further includes: performing a preprocessing operation on the concrete bridge image.

[0132] In some embodiments, the preprocessing operation may include at least one of the following steps:

[0133] Performing normalization processing on the concrete bridge image, where the normalization processing at least includes dividing the pixels of the concrete bridge image by 255 to scale the pixel values to between 0 and 1;

[0134] Perform at least one of the operations of enhancing the concrete bridge image, adjusting the brightness, enhancing the contrast, and rotating the image;

[0135] Perform a fusion operation on the concrete bridge image.

[0136] In some embodiments, the fusion operation may include: randomly selecting four different images from the concrete bridge image. Then resize the side length of the image to 640 while keeping the image ratio unchanged. After that, crop the four different images in turn with a crosshair at a random position. Thus, splice the four cropped images to obtain the fused concrete bridge image. Finally, update the bbox coordinates of the fused concrete bridge image. As an example, refer to Figure 8 , Figure 8 which is the schematic diagram of image cropping proposed in this application. As shown in Figure 8 (a) figure, four concrete bridge images are respectively cropped by the crosshair at a random position, so as to obtain four different images, and splice the corresponding positions of the four different images, thus obtaining Figure 8 the concrete bridge image shown in (b) figure.

[0137] Embodiment III

[0138] Refer to Figure 9 , Figure 9 which is the flowchart of the steps of a multi-task distillation-driven method for detecting and segmenting the apparent diseases of a concrete bridge proposed in this application. A multi-task distillation-driven method for detecting and segmenting the apparent diseases of a concrete bridge proposed in this application further includes steps S30 to S50.

[0139] Step S30: Obtain the dataset to be trained;

[0140] Step S40: And / or perform the preprocessing operation on the dataset;

[0141] In this embodiment, the preprocessing operation can refer to the relevant embodiments given above, and will not be elaborated here.

[0142] Step S50: Input the dataset into a preset teacher model and the student model, and train the student model from the teacher model by the knowledge distillation method to obtain the trained student model.

[0143] In this embodiment, the teacher model can be any teacher model in the prior art. For example, models such as AOYOLO, YOLOv7, and YOLOv9. Refer to Figure 10 , Figure 10 which is the training schematic diagram of the student model proposed in this application. As shown in Figure 10As shown, the student model simultaneously performs defect detection ( Figure 10 Task 1 in Figure 10 ) and crack segmentation ( Task 2 in

[0144] ) through knowledge distillation from the teacher model, and finally obtains a student model capable of simultaneously performing defect detection and crack segmentation tasks. In the prior art, both the teacher model and the student model are trained using the cross-entropy loss function. Among them, the loss function of the student model depends on both hard targets and the output probability distribution of the teacher model, and is optimized recursively. This method has good effects in single-task scenarios, but when faced with applications that need to handle multiple tasks, such as multi-task learning, the adaptability of the single-task distillation strategy is poor.

[0145]

[0146]

[0147] In some embodiments, it is well-known in the art that the CS divergence uses the logarithmic operator to quantify the distance between two domains. Inspired by the CS divergence, this application innovatively designs the knowledge transfer loss function of the student model using the CS divergence. The knowledge transfer loss function is:

[0148]

[0149]

[0150] where m represents the batch size, j represents the index of the feature dimension, d represents the number of classes, i represents the i-th class, ε is a constant, P represents the prediction result of the student model, Q represents the output of the teacher model, and Q serves as the soft label to guide the learning of the student model. and are the knowledge transfer losses of the object detection task and the pixel segmentation task respectively.

[0151] In some embodiments, in order to effectively complete the knowledge transfer from the large teacher model to the lightweight student model, this application designs a total hybrid loss for the student model to supervise the distillation learning of the student model, which means that the student model can simultaneously learn object detection knowledge and segmentation parsing knowledge from the teacher model, thereby greatly improving the detection efficiency of the student model. The total hybrid loss of the student model is:

[0152] L total = a(L det + L seg ) + (1 - a)L dil

[0153] Among them, L det , L seg and L dil respectively represent the object detection loss, the pixel segmentation loss, and the knowledge transfer loss.

[0154] In the technical solution provided in this embodiment, through the knowledge transfer loss and the total mixing loss of the student model, the distillation learning of the model is supervised through the total mixing loss, so that the student model can learn object detection knowledge and segmentation parsing knowledge from the teacher model at the same time, thereby greatly improving the detection efficiency and the performance of the student model in multiple task scenarios.

[0155] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.

[0156] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0157] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions for causing a terminal device (which can be a computer, a mobile phone, a tablet computer) to execute the methods described in the various embodiments of the present application.

[0158] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A multi-task distillation-driven method for detecting and segmenting the apparent diseases of concrete bridges, characterized in that, Including: Obtain a concrete bridge image to be detected; Input the concrete bridge image into a trained student model to obtain a defect detection result image and / or a crack segmentation result image; the student model at least includes a feature extraction module, a detection network, and / or a segmentation network connected in series in sequence; where The feature extraction module is used to extract image features from the concrete bridge image; The detection network is used to identify defects in the concrete bridge image to generate the defect detection result image; The segmentation network is used to identify cracks from the concrete bridge image and segment the area where the cracks are located to generate the crack segmentation result image; where The feature extraction module is sequentially composed of a 3X3 Conv1 layer, a 3X3 Conv2 layer, a P_block1 module, a 3X3 Conv3 layer, a P_block2 module, a 3X3 Conv4 layer, a P_block3 module, a 3X3 Conv5 layer, and a P_block4 module connected in series. Among them, the stride of the 3X3 Conv1 layer to the 3X3 Conv5 layer is 2, and the number of channels starts from 2^4 and increases successively in powers of 2; The P_block1 module to the P_block4 module are each sequentially composed of a first branch, a second branch connected in parallel, and a first splicing module respectively connected to the outputs of the first branch and the second branch; among them, the first branch is sequentially composed of a 3X3 PConv layer, a BN layer, an h_swish activation function, an SE attention layer, and a 1×1Conv6 layer, a BN layer connected in series, and the second branch is used to receive the input and transmit the input to the first splicing module.

2. The method according to claim 1, wherein The detection network is sequentially composed of an SPPELAN module, a second splicing module, a first upsampling module, a third splicing module, a 3X3 Conv7 layer, a first RepNCSPELAN4 module, a second upsampling module, a fourth splicing module, a second RepNCSPELAN4 module, and a detection module connected in series. Among them, the output of the first RepNCSPELAN4 module is also connected to the input of the detection module; the upsampling multiples of the first upsampling module and the second upsampling module are both 2; the stride of the 3X3 Conv7 layer is 1, and the number of channels is 256; the splicing dimensions of the second splicing module to the fourth splicing module are all 1; where The output of the P_block4 module in the feature extraction module is respectively connected to the input of the SPPELAN module and the input of the second splicing module, the output of the P_block3 module is connected to the input of the third splicing module, and the output of the P_block2 module is connected to the input of the fourth splicing module.

3. The method according to claim 2, wherein The SPPELAN module consists of a 1X1 Conv8 layer, an SP1 layer, an SP2 layer, an SP3 layer connected in series in sequence, a first Concat layer connected to the outputs of the 1X1 Conv8 layer, the SP1 layer, the SP2 layer, and the SP3 layer respectively, and a 1X1 Conv9 layer connected to the output of the first Concat layer; the number of channels of the 1X1 Conv8 layer and the 1X1 Conv9 layer is 256, and the Kernel_size = 5, Stride = 1, Padding = 2 for the SP1 layer, the SP2 layer, and the SP3 layer in the SPPELAN module 。 4. The method according to claim 3, wherein The first RepNCSPELAN4 module and the second RepNCSPELAN4 module are both sequentially composed of a 1X1 Conv10 layer in series, a Spilt layer, a first RepNCSP module, a 3X3 Conv11 layer, a 3X3 Conv12 layer connected to the output of the Spilt layer, a second RepNCSP module connected to the output of the 3X3 Conv12 layer, a second Concat layer respectively connected to the outputs of the Spilt layer, the 3X3 Conv11 layer and the 3X3 Conv12 layer, and a 1X1 Conv13 layer connected to the output of the second Concat layer. The output of the second RepNCSP module is also connected to the input of the first RepNCSP module. The Spilt layer splits the input into two parts and transmits them to the second Concat layer respectively. The stride of the 1X1Conv10 layer and the 1X1 Conv13 layer is 1, and the padding size is 0. The stride of the 3X3Conv11 layer and the 3X3 Conv12 layer is 1, and the padding size is 1.

5. The method according to claim 4, wherein The first RepNCSP module and the second RepNCSP module are both sequentially composed of a third branch and a fourth branch in parallel, a third Concat layer respectively connected to the outputs of the third branch and the fourth branch, and a 1X1Conv14 layer connected to the output of the third Concat layer. The stride of the 1X1Conv14 layer is 1, and the padding size is 0. The third branch is sequentially composed of a 1X1 Conv15 layer and N RepNBottLeneck modules in series. The fourth branch is composed of a 1X1 Conv16 layer. The stride of the 1X1 Conv15 layer and the 1X1 Conv16 layer is 1, and the padding size is 0.

6. The method according to claim 5, characterized in that, The RepNBottLeneck module is sequentially composed of a fifth branch and a sixth branch in parallel. The fifth branch is sequentially composed of a RepConvN module and a 1X1Conv17 layer in series. The stride of the 1X1Conv17 layer is 1, and the padding size is 0. The sixth branch is used to receive and transmit the input.

7. The method according to claim 1, wherein The segmentation network is sequentially composed of a third upsampling module, a 3X3DWConv1 layer, a fourth upsampling module, a 3X3 DWConv2 layer, a fifth upsampling module, a 3X3 DWConv3 layer, a sixth upsampling module, a C2f module, and a 3X3 DWConv4 layer in series. The sampling multiples of the third upsampling module to the sixth upsampling module are all 2. The number of channels of the 3X3 DWConv1 layer to the 3X3DWConv3 layer starts from 2^6 and decreases in powers of 2 in sequence, and the strides are 2, 1, 1 in sequence. The number of channels of the 3X3DWConv4 layer is 4, and the stride is 1. The number of channels of the C2f module is 16. The output of the second RepNCSPELAN4 module in the detection network is connected to the input of the third upsampling module.

8. The method according to claim 1, characterized in that Before the step of inputting the concrete bridge image into the trained student model to obtain a defect detection result image and / or a segmentation result image, the following steps are further included: Performing a preprocessing operation on the concrete bridge image; wherein, The preprocessing operation includes at least one of the following steps: Performing a normalization process on the concrete bridge image, and the normalization process at least includes dividing the pixels of the concrete bridge image by 255 to scale the pixel values to between 0 and 1; Performing at least one operation of enhancement operation, brightness adjustment, contrast enhancement, and image rotation on the concrete bridge image; Performing a fusion operation on the concrete bridge image; wherein, The fusion operation sequentially includes the following steps: Randomly selecting four different images from the concrete bridge image; Resizing the side length of the image to 640 while keeping the image ratio unchanged; Cropping the four different images respectively with a crosshair at a random position in sequence; Stitching the cropped images to obtain the fused concrete bridge image; Updating the bbox coordinates of the fused concrete bridge image.

9. The method according to claim 8, characterized in that, The method further includes: Obtaining a dataset to be trained; Performing the preprocessing operation on the dataset; Inputting the dataset into a preset teacher model and the student model, and training the student model from the teacher model through the knowledge distillation method to obtain the trained student model.

10. The method according to claim 9, wherein The student model includes a knowledge transfer loss and a total mixed loss; wherein, The knowledge transfer loss of the student model is: Wherein, m represents the batch size, j represents the index of the feature dimension, d represents the number of categories, i represents the i-th category, ε is a constant, P represents the prediction result of the student model, Q represents the output of the teacher model, and Q is used as the soft label to guide the learning of the student model; Wherein, if ρ and σ are set to 1, the knowledge transfer loss is: Among them, and are the knowledge transfer losses of the object detection task and the pixel segmentation task respectively. The total mixed loss of the student model is: L total = a(L det + L seg ) + (1 - a)L dil Among them, L det , L seg and L dil respectively represent the object detection loss, the pixel segmentation loss, and the knowledge transfer loss.