Carbon brake disc defect detection method based on DUFG-Net model

Through the detection method based on the DUFG-Net model, combined with the multi-directional feature enhancement network and the U2-Net network, the problem of inability to effectively detect internal defects of the carbon brake disc and handle small defects in the prior art is solved, and efficient and accurate defect detection effect is achieved.

CN120070344APending Publication Date: 2025-05-30XIDIAN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510105297.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art cannot effectively detect internal defects of the brake disc in carbon brake disc defect detection, and lacks targetedness when dealing with small defects, resulting in high missed detection rates and false detection rates.

Method used

Using the detection method based on the DUFG-Net model, the internal defect images of the carbon brake disc are collected through X-ray images, and the multi-directional feature enhancement network is used to extract texture features from four directions, and deep learning is combined with the U2-Net network and the discriminator network to optimize the generator's processing ability to the edge.

Benefits of technology

It realizes efficient detection of internal and surface defects of carbon brake discs, improves the detection accuracy of micro defects, reduces the leakage detection rate and error detection rate, and meets the strict quality inspection standards of carbon brake discs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070344A_ABST
    Figure CN120070344A_ABST
Patent Text Reader

Abstract

The invention discloses a carbon brake disc defect detection method based on a DUFG-Net model, a generated sample set is a collected carbon brake disc defect image which is an X-ray image, the internal quality condition of a carbon brake disc can be observed, texture features are extracted from four directions through a multi-direction feature enhancement network, and a carbon brake disc defect detection result is obtained. And structure information in the image of the carbon brake disc is fully utilized. In a shallow decoder of the U2-Net network, shallow features are extracted through a convolutional layer, up-sampling is adjusted to be in the size of an original image, an edge image is obtained and generated through an activation function, the edge image is generated through continuous confrontation of the U2-Net network and a discriminator, and the processing capacity of a generator to edges is optimized through supervised learning of the edge image and a real edge image. The method is suitable for detecting the surface defects of the carbon brake disc, and can be used for detecting cracks or tiny edge areas in the internal defects of the carbon brake disc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of mechanical defect diagnosis, and further relates to a method for detecting carbon brake disc defects based on the DUFG-Net model in the technical field of brake disc fault diagnosis. The present invention can be used to detect defects in the quality of carbon brake discs. Background Art

[0002] Installing a carbon brake disc with defects on a vehicle may lead to brake failure or performance degradation, thereby affecting driving safety. Therefore, it is necessary to strictly detect the carbon brake disc before leaving the factory to ensure its braking effect and eliminate any potential safety hazards. However, the current detection technologies are not sufficient. Manual inspection checks for surface defects by visual observation or simple tools, but is ineffective for detecting small or internal defects and has low efficiency. Although machine vision inspection has been applied, there are still deficiencies in the development of these methods in the technical field of brake disc fault diagnosis. Firstly, most of the existing technologies are used for fault diagnosis of brake disc surface defects. Secondly, the current target detection algorithms are mostly designed for general industrial products, and key detail information is lost in the low-resolution feature maps due to continuous convolution and downsampling during the feature extraction process, resulting in high rates of missed detection and false detection when detecting small defects and failing to meet the strict quality detection standards of carbon brake discs.

[0003] Changchun University of Technology discloses a method for surface defect detection of automotive brake discs in its patent document "A Method for Surface Defect Detection of Automotive Brake Discs Based on Deep Learning" (Application No. CN 202311584766.4, Publication No. CN 117541568 A). The implementation steps of this method are as follows: 1. Build an image acquisition device to obtain the surface image of the automotive brake disc; 2. Preprocess and annotate the obtained surface image of the automotive brake disc to establish a surface defect dataset of the automotive brake disc; 3. Construct a surface defect detection model for the automotive brake disc based on the surface defect dataset of the automotive brake disc; 4. Train the model using the training set and validation set and select the optimal detection model according to the evaluation index; 5. Validate the model using the test set to determine whether the model performance meets the requirements; 6. Use the defect detection model that meets the requirements to detect the surface of the automotive brake disc and output the detection result to achieve intelligent defect detection and recognition. This method can improve the efficiency and accuracy of surface defect detection of automotive brake discs, reduce labor costs, and can quickly adapt to the surface defect detection of new products, shortening the development cycle. However, the deficiencies of this method still exist. On the one hand, since the training set in this method only contains the surface images of automotive brake discs, the constructed defect detection model is more suitable for surface defect detection of automotive brake discs and cannot detect internal defects of the brake discs. On the other hand, since the brake disc surface defect detection model of this method uses the YOLOv5 network and lacks pertinence in dealing with tiny defects, small-scale or fuzzy-edge defects are missed. Summary of the Invention

[0004] The purpose of the present invention is to provide a carbon brake disc defect detection method based on the DUFG-Net model for the deficiencies of the above-mentioned existing technologies, to solve the problems that the existing brake disc defect detection model cannot be applied to the surface quality defect detection of automotive brake discs and the existing technology lacks pertinence in dealing with tiny defects.

[0005] The technical idea for achieving the object of the present invention is as follows: The defective carbon brake disc images collected in the present invention are X-ray images, which are not restricted by the part shape and can observe the internal quality of the carbon brake disc. Through the multi-directional feature enhancement network, texture features are extracted from four directions, making full use of the structural information in the carbon brake disc images, so as to solve the problem that the existing brake disc defect detection model is more suitable for detecting surface defects of automotive brake discs. The multi-directional feature enhancement network includes a feature extraction layer, a Gabor filter, an adaptive weighting module, a multi-head self-attention mechanism, and a feature fusion and output module connected in sequence. The feature extraction layer extracts multi-scale information from the input features, captures features with different receptive fields through multi-scale convolution, and efficiently extracts direction-sensitive features using depthwise separable convolution to form a preliminary feature representation. The Gabor filter simulates the direction selectivity of the human visual system, enhances the feature response of the defective area from four different directions of 0°, 45°, 90°, and 135°, and extracts the direction-sensitive features on the surface of the brake disc. The adaptive weighting module dynamically adjusts the weights according to the importance of different direction features, enhances the representation ability of key features, and suppresses redundant information at the same time. The multi-head self-attention mechanism captures the long-range dependence relationship between features through global feature interaction learning, improving the sensitivity to tiny defects. The feature fusion and output module integrates the multi-directional and multi-scale features output by each module into a unified representation, and restores the resolution to the same as the input image through upsampling, providing high-quality input features for the subsequent detection network. In the present invention, shallow features are extracted through a convolutional layer in the shallow decoder of the U2-Net network, upsampled to the original image size, and passed through an activation function to obtain a generated edge map. The generated edge map optimizes the edge processing ability of the generator through supervised learning with the real edge map. The discriminator of the present invention consists of a lightweight convolutional network, and the constraint on the generated edge map enhances the clarity and authenticity of the edge, so as to solve the problem that the existing technology lacks pertinence in dealing with tiny defects.

[0006] To achieve the above object, the specific implementation steps of the present invention are as follows:

[0007] The carbon brake disc is imaged by an X-ray digital scanner to obtain a defective carbon brake disc image, and a DUFG-Net model is constructed to learn the defect features of the image; the steps of this method include the following:

[0008] Step 1, generate a training set and a test set:

[0009] K black-and-white carbon brake disc images with defects obtained by the X-ray digital scanner are used to form a sample set, where K≥200, and the images in the sample set include at least five types of defects; after the images in the sample set are processed and labeled, according to the ratio of 8:2, the images in the sample set and their corresponding labels are divided into a training set and a test set;

[0010] Step 2: Build a DUFG-Net model with a multi-directional feature enhancement network, a U2-Net network, and a discriminator network connected in sequence, and set the parameters;

[0011] Step 3: Train the DUFG-Net model;

[0012] Step 4: Input the test set into the trained DUFG-Net model and output the detection results of carbon brake disc defects.

[0013] Furthermore, the steps to generate the training set and the test set are as follows:

[0014] First step: Inside the X-ray digital scanner, take pictures of the carbon brake disc from directly above the circular surface of the carbon brake disc to obtain carbon brake disc X-ray images that can observe the internal quality status of the carbon brake disc. The images are black and white images. Manually screen the carbon brake disc X-ray images to obtain K carbon brake disc defect images with defects, where K ≥ 200, including at least five types of defects;

[0015] Second step: Crop the carbon brake disc defect images to obtain sub-images of size 256×256, with a 20% overlapping area between two sub-images;

[0016] Third step: Perform segmentation mask annotation on the carbon brake disc sub-images and save them in the COCO annotation format;

[0017] Fourth step: Combine the carbon brake disc sub-images and their corresponding labels to form a training sample set M and a test sample set K, with a ratio of 8:2.

[0018] Furthermore, the five types of defects include: circular hole defects, high-density inclusions, cracks, porosity and uneven material distribution, and delamination.

[0019] Furthermore, the segmentation mask annotation refers to: annotating each pixel in each image, marking the pixel values of the non-defect part as 0 and the pixel values of the defect part as 1. The output mask image is a single-channel grayscale image.

[0020] Furthermore, the structure of the multi-directional feature enhancement network is in sequence: a feature extraction layer, a multi-directional Gabor filter, an adaptive weighting module, a multi-head self-attention mechanism, and a feature fusion and output module;

[0021] The structure of the described feature extraction layer is as follows: the first convolutional layer, the batch normalization layer, the RELU activation layer, the parallel second convolutional layer, third convolutional layer, and fourth convolutional layer, the fifth convolutional layer, the depth convolutional layer, and the pointwise convolutional layer; the convolutional kernel sizes of the first to fifth convolutional layers, the depth convolutional layer, and the pointwise convolutional layer of the feature extraction layer are set to 3×3, 3×3, 5×5, 7×7, 1×1, 3×3, 1×1 in sequence, the number of convolutional kernels is set to 64, 64, 64, 64, 128, 128, 128 in sequence, the stride is set to 1, and the padding is set to 1, 1, 2, 3, 0, 1, 0 in sequence; the batch normalization layer is implemented using BatchNorm2d; the RELU activation layer is implemented using the RELU function;

[0022] The structure of the described multi-directional Gabor filter is as follows: a directionally adjustable Gabor convolutional kernel and a direction splicing and fusion sub-module; the convolutional kernel size of the directionally adjustable Gabor convolutional kernel of the multi-directional Gabor filter is set to 5×5, the number of directions is set to 4, which are 0°, 45°, 90°, and 135° in sequence, and the two-dimensional Gabor kernel is set to: G(x,y;λ,θ,φ,σ,γ) = exp(-(〖x'〗^2 + γ^2〖y'〗^2) / (2σ^2))cos(2πx' / λ + φ) where x' = xcosθ + ysinθ, y' = -xsinθ + ycosθ, where λ is the wavelength, which controls the sensitivity of the filter to image texture, θ is the direction angle, which are 0°, 45°, 90°, and 135° in sequence, φ is the phase deviation, σ is the standard deviation of the Gaussian envelope, and γ is the spatial aspect ratio;

[0023] The structure of the described direction splicing and fusion sub-module is as follows: a feature splicing layer, a convolutional layer, a RELU activation layer, and a batch normalization layer; the number of channels of the feature splicing layer of the direction splicing and fusion sub-module is set to 512, the convolutional kernel size of the convolutional layer is set to 1×1, the number of convolutional kernels is set to 512, the RELU activation layer is implemented using the RELU function, and the batch normalization layer is implemented using BatchNorm2d;

[0024] The structure of the described adaptive weighting module is as follows: a global average pooling layer, a first fully connected layer, and a second fully connected layer; the number of neurons of the first and second fully connected layers of the adaptive weighting module are set to 64 and 4 in sequence, and the activation functions are set to: RELU, Softmax; the structure of the described multi-head self-attention mechanism is as follows: a first linear transformation layer, a batch normalization layer, and a second linear transformation layer; the number of channels of the first and second linear transformation layers of the multi-head self-attention mechanism are both set to 128, and the batch normalization layer is implemented using BatchNorm2d;

[0025] The structure of the feature fusion and output module is as follows in sequence: convolutional layer, batch normalization layer, activation layer; the convolution kernel size of the convolutional layer of the feature fusion and output module is set to 1×1, the number of convolution kernels is set to 64, the batch normalization layer is implemented by BatchNorm2d, and the activation layer is implemented by the RELU function.

[0026] Further, the batch normalization layer implemented by BatchNorm2d means that the mean and variance of the mini-batch data are normalized by the following formula:

[0027]

[0028] where x i represents the pixel value of the input feature map, μ represents the mean of the current batch, σ 2 represents the variance of the current batch, ∈ represents the numerical stability term, and γ and β represent the learnable parameters.

[0029] Further, the structure of the U2-Net network is as follows in sequence: the 1st decoder, the 2nd decoder, the 3rd decoder, the 4th decoder, the 5th decoder, the 6th decoder, the 1st encoder, the 2nd encoder, the 3rd encoder, the 4th encoder, the 5th encoder, Concatenation layer;

[0030] The structures of the first decoder and the first encoder are exactly the same, and their structures are as follows in sequence: the first convolutional layer, the first batch normalization layer, the first activation layer, the second convolutional layer, the second batch normalization layer, the second activation layer, the first downsampling layer, the third convolutional layer, the third batch normalization layer, the third activation layer, the second downsampling layer, the fourth convolutional layer, the fourth batch normalization layer, the fourth activation layer, the third downsampling layer, the fifth convolutional layer, the fifth batch normalization layer, the fifth activation layer, the fourth downsampling layer, the sixth convolutional layer, the sixth batch normalization layer, the sixth activation layer, the fifth downsampling layer, the seventh convolutional layer, the seventh batch normalization layer, the seventh activation layer, the eighth convolutional layer, the eighth batch normalization layer, the eighth activation layer, the ninth convolutional layer, the ninth batch normalization layer, the ninth activation layer, the first upsampling layer, the tenth convolutional layer, the tenth batch normalization layer, the tenth activation layer, the second upsampling layer, the eleventh convolutional layer, the eleventh batch normalization layer, the eleventh activation layer, the third upsampling layer, the twelfth convolutional layer, the twelfth batch normalization layer, the twelfth activation layer, the fourth upsampling layer, the thirteenth convolutional layer, the thirteenth batch normalization layer, the thirteenth activation layer, the fifth upsampling layer, the fourteenth convolutional layer, the fourteenth batch normalization layer, the fourteenth activation layer; the kernel sizes of the first to fourteenth convolutional layers in the first decoder and the first encoder are all set to: 3×3, and the numbers of convolutional kernels are set in sequence to: 64, 64, 32, 16, 8, 4, 2, 2, 2, 4, 8, 16, 32, 64. The first to fourteenth batch normalization layers are all implemented using BatchNorm2d, the first to fourteenth activation layers are all implemented using the RELU function, the sampling multiples of the first to fifth downsampling layers are all set to 2, and the sampling multiples of the first to fifth upsampling layers are all set to ;

[0031] The structures of the second decoder and the second encoder are exactly the same, and their structures are reduced compared to the first decoder and the first encoder by the fifth downsampling layer, the seventh convolutional layer, the seventh batch normalization layer, the seventh activation layer, the fifth upsampling layer, the fourteenth convolutional layer, the fourteenth batch normalization layer, and the fourteenth activation layer;

[0032] The kernel sizes of the 12 convolutional layers in the second decoder and the second encoder are all set to: 3×3, and the numbers of convolutional kernels are set in sequence to: 64, 64, 32, 16, 8, 4, 4, 4, 8, 16, 32, 64;

[0033] The structures of the third decoder and the third encoder are exactly the same, and their structures are reduced compared to the second decoder and the second encoder by the fourth downsampling layer, the sixth convolutional layer, the sixth batch normalization layer, the sixth activation layer, the fourth upsampling layer, the thirteenth convolutional layer, the thirteenth batch normalization layer, and the thirteenth activation layer;

[0034] Set the convolutional kernel sizes of the convolutional layers in the third decoder and the third encoder to: 3×3, and set the numbers of convolutional kernels to: 64, 64, 32, 16, 8, 8, 8, 16, 32, 64 in sequence;

[0035] The structures of the fourth decoder and the fourth encoder are exactly the same. Compared with the third decoder and the third encoder, their structures reduce the third downsampling layer, the fifth convolutional layer, the fifth batch normalization layer, the fifth activation layer, the third upsampling layer, the twelfth convolutional layer, the twelfth batch normalization layer, and the twelfth activation layer;

[0036] Set the convolutional kernel sizes of the convolutional layers in the fourth decoder and the fourth encoder to 3×3, and set the numbers of convolutional kernels to: 64, 64, 32, 16, 16, 16, 32, 64 in sequence;

[0037] The structures of the fifth and sixth decoders and the fifth encoder are exactly the same. Their structures are in sequence: the first convolutional layer, the first batch normalization layer, the first activation layer, the second convolutional layer, the second batch normalization layer, the second activation layer, the third convolutional layer, the third batch normalization layer, the third activation layer, the fourth convolutional layer, the fourth batch normalization layer, the fourth activation layer, the fifth convolutional layer, the fifth batch normalization layer, the fifth activation layer, the sixth convolutional layer, the sixth batch normalization layer, the sixth activation layer, the seventh convolutional layer, the seventh batch normalization layer, the seventh activation layer, the eighth convolutional layer, the eighth batch normalization layer, the eighth activation layer;

[0038] Set the convolutional kernel sizes of the first to eighth convolutional layers in the fifth and sixth decoders and the fifth encoder to 3×3, set the numbers of convolutional kernels to 64, implement the first to eighth batch normalization layers with BatchNorm2d, and implement the first to eighth activation layers with the RELU function.

[0039] Furthermore, the structure of the discriminator network is in sequence: the input layer, the first convolutional layer, the first pooling layer, the second convolutional layer, the second pooling layer, the third convolutional layer, the third pooling layer, the first fully connected layer, the second fully connected layer; set the number of input channels of the input layer to 1, and set the size of the passable image to 256×256; set the convolutional kernel sizes of the first to third convolutional layers to 3×3, and set the numbers of convolutional kernels to: 64, 128, 256 in sequence; set the size of the pooling kernel of the pooling layer to 2×2, set the pooling stride to 2, and set the numbers of neurons in the first and second fully connected layers to: 1024, 1 in sequence.

[0040] Furthermore, the training of the DUFG-Net model means inputting the training set into the DUFG-Net model for forward propagation, and iteratively updating the weight parameters of the U2-Net network through the cross-entropy loss and the Dice loss until the total loss function of the model converges, and obtaining the trained DUFG-Net model.

[0041] Further, the total loss function of the model includes: cross-entropy loss L BCE , Dice loss L Dice , generation loss and adversarial loss and the update formulas for updating θ t , ε t are as follows:

[0042]

[0043]

[0044] where N represents the total number of pixels, y i is the true mask value of the i-th pixel of the sample image, is the generated mask value of the i-th pixel of the sample image, log represents the logarithmic operation, D(y) is the output of the discriminator for the true mask, taking the value of 1, is the output of the discriminator for the generated mask, taking the value in [0,1], and respectively represent the loss values of the U2-Net network and the discriminator network, θ t-1 , ε t-1 respectively represent the weight parameters of the U2-Net network and the discriminator network at the (t - 1)-th iteration, and α represents the learning rate.

[0045] Compared with the prior art, the present invention has the following advantages:

[0046] First, the sample set generated by the present invention is the X-ray image of the carbon brake disc defect collected, which can observe the internal quality condition of the carbon brake disc. Through the multi-directional feature enhancement network, texture features are extracted from four directions, making full use of the structural information in the carbon brake disc image, overcoming the problem that the prior art brake disc defect detection model is more suitable for the surface defect detection of automotive brake discs. The present invention is not only applicable to the detection of carbon brake disc surface defects, but also can be used to detect internal defects of carbon brake discs, and is applicable to carbon brake discs installed on various vehicles.

[0047] Second, the present invention extracts shallow features through the convolutional layer in the shallow decoder of the U2-Net network, upsamples and adjusts them to the original image size, and obtains the generated edge map through the activation function. Through the continuous confrontation between the U2-Net network and the discriminator, the generated edge map is supervised and learned with the true edge map to optimize the edge processing ability of the generator, overcoming the problem of lack of pertinence in the prior art for dealing with tiny defects, and greatly improving the detection accuracy of cracks or fine edge regions in the carbon brake disc defect detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is the implementation flowchart of the embodiments of the present invention;

[0049] Figure 2 is the structural schematic diagram of the DUFG-Net model constructed in the embodiments of the present invention;

[0050] Figure 3 is the structural schematic diagram of the multi-directional feature enhancement network constructed in the embodiments of the present invention;

[0051] Figure 4 is the structural schematic diagram of the U2-Net network constructed in the embodiments of the present invention;

[0052] Figure 5 is the structural schematic diagram of the discriminator network constructed in the embodiments of the present invention. Detailed implementation manners

[0053] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0054] Refer to Figure 1 for a further detailed description of the implementation steps of the embodiments of the present invention.

[0055] Step 1, generate a training set and a test set.

[0056] Inside the X-ray digital scanner, take pictures of the aircraft carbon brake disc directly above the circular surface of the aircraft carbon brake disc to obtain X-ray images of the aircraft carbon brake disc that can observe the internal quality status of the aircraft carbon brake disc. The images are black and white images. Manually screen the X-ray images of the aircraft carbon brake disc to obtain 200 defective images of the aircraft carbon brake disc with defects, including at least five types of defects.

[0057] The five types of defects include: circular hole defects, high-density inclusions, cracks, porosity and uneven material distribution, delamination.

[0058] Crop the defective images of the carbon brake disc to obtain sub-images with a size of 256×256, and there is a 20% overlapping area between two sub-images.

[0059] Perform segmentation mask annotation on the sub-images of the carbon brake disc and save them in the COCO annotation format.

[0060] The segmentation mask annotation refers to: annotating each pixel in each image, marking the pixel values of the non-defective parts as 0 and the pixel values of the defective parts as 1, and the output mask image is a single-channel grayscale image.

[0061] Combine the sub-images of the carbon brake disc and their corresponding labels to form a training sample set M and a test sample set K, with a ratio of 8:2.

[0062] Step 2: Build a DUFG-Net model with a multi-directional feature enhancement network, a U2-Net network, and a discriminator network connected in sequence, and set parameters, as Figure 2 shown.

[0063] Refer to Figure 3 for a further detailed description of the multi-directional feature enhancement network constructed in the present invention.

[0064] The structure of the multi-directional feature enhancement network is successively: a feature extraction layer, a multi-directional Gabor filter, an adaptive weighting module, a multi-head self-attention mechanism, and a feature fusion and output module.

[0065] The structure of the feature extraction layer is successively: a first convolutional layer, a batch normalization layer, a RELU activation layer, parallel second, third, and fourth convolutional layers, a fifth convolutional layer, a depth convolutional layer, and a pointwise convolutional layer; the kernel sizes of the first to fifth convolutional layers, the depth convolutional layer, and the pointwise convolutional layer of the feature extraction layer are successively set to: 3×3, 3×3, 5×5, 7×7, 1×1, 3×3, 1×1; the number of kernels is successively set to: 64, 64, 64, 64, 128, 128, 128; the stride is set to 1 for all; the padding is successively set to 1, 1, 2, 3, 0, 1, 0; the batch normalization layer is implemented using BatchNorm2d; the RELU activation layer is implemented using the RELU function.

[0066] The statement that the batch normalization layer is implemented using BatchNorm2d means that the mean and variance of the mini-batch data are normalized by the following formula:

[0067]

[0068] where x i represents the pixel value of the input feature map, μ represents the mean of the current batch, σ 2 represents the variance of the current batch, ∈ represents a numerical stability term, and γ and β represent learnable parameters.

[0069] The described multi-directional Gabor filter structure is as follows: a directionally adjustable Gabor convolution kernel and a direction splicing and fusion sub-module; the convolution kernel size of the directionally adjustable Gabor convolution kernel of the multi-directional Gabor filter is set to 5×5, the number of directions is set to 4, which are 0°, 45°, 90°, and 135° in sequence, and the two-dimensional Gabor kernel is set as: G(x,y;λ,θ,φ,σ,γ) = exp(-(〖x'〗^2 + γ^2〖y'〗^2) / (2σ^2))cos(2πx' / λ + φ), where x' = xcosθ + ysinθ, y' = -xsinθ + ycosθ. Here, λ is the wavelength, which controls the sensitivity of the filter to image texture, θ is the direction angle, which are 0°, 45°, 90°, and 135° in sequence, φ is the phase deviation, σ is the standard deviation of the Gaussian envelope, and γ is the spatial aspect ratio.

[0070] The described direction splicing and fusion sub-module structure is as follows: a feature splicing layer, a convolution layer, a RELU activation layer, and a batch normalization layer; the number of channels of the feature splicing layer of the direction splicing and fusion sub-module is set to 512, the convolution kernel size of the convolution layer is set to 1×1, the number of convolution kernels is set to 512, the RELU activation layer is implemented using the RELU function, and the batch normalization layer is implemented using BatchNorm2d.

[0071] The described adaptive weighted module structure is as follows: a global average pooling layer, a first fully connected layer, and a second fully connected layer; the number of neurons in the first and second fully connected layers of the adaptive weighted module is set to 64 and 4 in sequence, and the activation functions are set to: RELU and Softmax; the described multi-head self-attention mechanism structure is as follows: a first linear transformation layer, a batch normalization layer, and a second linear transformation layer; the number of channels of the first and second linear transformation layers of the multi-head self-attention mechanism is set to 128, and the batch normalization layer is implemented using BatchNorm2d.

[0072] The described feature fusion and output module structure is as follows: a convolution layer, a batch normalization layer, and an activation layer; the convolution kernel size of the convolution layer of the feature fusion and output module is set to 1×1, the number of convolution kernels is set to 64, the batch normalization layer is implemented using BatchNorm2d, and the activation layer is implemented using the RELU function.

[0073] Refer to Figure 4 , and a further detailed description of the U2-Net network constructed by the present invention is given.

[0074] The structure of the described U2-Net network is as follows: a first decoder, a second decoder, a third decoder, a fourth decoder, a fifth decoder, a sixth decoder, a first encoder, a second encoder, a third encoder, a fourth encoder, a fifth encoder, and a Concatenation layer.

[0075] The structures of the first decoder and the first encoder are exactly the same, and their structures are as follows in sequence: the first convolutional layer, the first batch normalization layer, the first activation layer, the second convolutional layer, the second batch normalization layer, the second activation layer, the first downsampling layer, the third convolutional layer, the third batch normalization layer, the third activation layer, the second downsampling layer, the fourth convolutional layer, the fourth batch normalization layer, the fourth activation layer, the third downsampling layer, the fifth convolutional layer, the fifth batch normalization layer, the fifth activation layer, the fourth downsampling layer, the sixth convolutional layer, the sixth batch normalization layer, the sixth activation layer, the fifth downsampling layer, the seventh convolutional layer, the seventh batch normalization layer, the seventh activation layer, the eighth convolutional layer, the eighth batch normalization layer, the eighth activation layer, the ninth convolutional layer, the ninth batch normalization layer, the ninth activation layer, the first upsampling layer, the tenth convolutional layer, the tenth batch normalization layer, the tenth activation layer, the second upsampling layer, the eleventh convolutional layer, the eleventh batch normalization layer, the eleventh activation layer, the third upsampling layer, the twelfth convolutional layer, the twelfth batch normalization layer, the twelfth activation layer, the fourth upsampling layer, the thirteenth convolutional layer, the thirteenth batch normalization layer, the thirteenth activation layer, the fifth upsampling layer, the fourteenth convolutional layer, the fourteenth batch normalization layer, the fourteenth activation layer; the kernel sizes of the first to fourteenth convolutional layers in the first decoder and the first encoder are all set to 3×3, and the numbers of convolutional kernels are set in sequence as: 64, 64, 32, 16, 8, 4, 2, 2, 2, 4, 8, 16, 32, 64. The first to fourteenth batch normalization layers are all implemented by BatchNorm2d, the first to fourteenth activation layers are all implemented by the RELU function, the sampling multiples of the first to fifth downsampling layers are all set to 2, and the sampling multiples of the first to fifth upsampling layers are all set to 。

[0076] The structures of the second decoder and the second encoder are exactly the same, and their structures are reduced by the fifth downsampling layer, the seventh convolutional layer, the seventh batch normalization layer, the seventh activation layer, the fifth upsampling layer, the fourteenth convolutional layer, the fourteenth batch normalization layer, and the fourteenth activation layer compared with the first decoder and the first encoder.

[0077] The kernel sizes of the 12 convolutional layers in the second decoder and the second encoder are all set to 3×3, and the numbers of convolutional kernels are set in sequence as: 64, 64, 32, 16, 8, 4, 4, 4, 8, 16, 32, 64.

[0078] The structures of the third decoder and the third encoder are exactly the same, and their structures are reduced by the fourth downsampling layer, the sixth convolutional layer, the sixth batch normalization layer, the sixth activation layer, the fourth upsampling layer, the thirteenth convolutional layer, the thirteenth batch normalization layer, and the thirteenth activation layer compared with the second decoder and the second encoder.

[0079] Set the convolutional kernel size of the convolutional layers in the third decoder and the third encoder to 3×3, and set the number of convolutional kernels to: 64, 64, 32, 16, 8, 8, 8, 16, 32, 64 in sequence.

[0080] The structures of the fourth decoder and the fourth encoder are exactly the same. Compared with the third decoder and the third encoder, their structures reduce the third downsampling layer, the fifth convolutional layer, the fifth batch normalization layer, the fifth activation layer, the third upsampling layer, the twelfth convolutional layer, the twelfth batch normalization layer, and the twelfth activation layer.

[0081] Set the convolutional kernel size of the convolutional layers in the fourth decoder and the fourth encoder to 3×3, and set the number of convolutional kernels to: 64, 64, 32, 16, 16, 16, 32, 64 in sequence.

[0082] The structures of the fifth and sixth decoders and the fifth encoder are exactly the same. Their structures are in sequence: the first convolutional layer, the first batch normalization layer, the first activation layer, the second convolutional layer, the second batch normalization layer, the second activation layer, the third convolutional layer, the third batch normalization layer, the third activation layer, the fourth convolutional layer, the fourth batch normalization layer, the fourth activation layer, the fifth convolutional layer, the fifth batch normalization layer, the fifth activation layer, the sixth convolutional layer, the sixth batch normalization layer, the sixth activation layer, the seventh convolutional layer, the seventh batch normalization layer, the seventh activation layer, the eighth convolutional layer, the eighth batch normalization layer, the eighth activation layer.

[0083] Set the convolutional kernel size of the first to eighth convolutional layers in the fifth and sixth decoders and the fifth encoder to 3×3, set the number of convolutional kernels to 64, implement the first to eighth batch normalization layers using BatchNorm2d, and implement the first to eighth activation layers using the RELU function.

[0084] Refer to Figure 5 , and make a further detailed description of the discriminator network constructed by the present invention.

[0085] The structure of the discriminator network is in sequence: the input layer, the first convolutional layer, the first pooling layer, the second convolutional layer, the second pooling layer, the third convolutional layer, the third pooling layer, the first fully connected layer, the second fully connected layer; set the number of input channels of the input layer to 1, and set the passable image size to 256×256; set the convolutional kernel size of the first to third convolutional layers to 3×3, and set the number of convolutional kernels to: 64, 128, 256 in sequence; set the size of the pooling kernel of the pooling layer to 2×2, set the pooling stride to 2, and set the number of neurons in the first and second fully connected layers to: 1024, 1 in sequence.

[0086] Step 3, train the DUFG-Net model.

[0087] Input the training set into the DUFG-Net model for forward propagation.

[0088] Randomly select 32 from the training sample set M as the input of the DUFG-Net model for forward propagation to obtain the defect detection results corresponding to 32 training samples.

[0089] The feature extraction layer in the multi-directional feature enhancement network generates a multi-scale feature map x 1 , and the multi-directional Gabor filter extracts the edge and texture features in four directions (0°, 45°, 90°, 135°) from the feature map x 1 to generate a multi-directional feature map x 2 . The adaptive weighting module assigns weights to each direction to enhance the attention to key directions, and the multi-head self-attention mechanism dynamically weights and fuses the directional features of the feature map x 2 to obtain the feature map x 3 . The feature fusion and output module downsamples the fused feature map x 3 to output a feature map x 4 .

[0090] The first decoder in the U2-Net network downsamples the feature map x 4 to obtain the feature map x 4 , the second decoder downsamples the feature map x 5 to obtain the feature map x 6 , the third decoder downsamples the feature map x 6 to obtain the feature map x 7 , the fourth decoder downsamples the feature map x 7 to obtain the feature map x 8 , the fifth decoder convolves the feature map x 8 to obtain the feature map x 9 , the sixth decoder convolves the feature map x 9 to obtain the feature map x 10 , the fifth encoder upsamples the feature map x 9 and the feature map x 10 by bilinear interpolation to obtain an upsampled feature map x 11 , the fourth encoder upsamples the feature map x 8 and the feature map x 11 by bilinear interpolation to obtain an upsampled feature map x 12 , the third encoder upsamples the feature map x 7 and the feature map x 12 by bilinear interpolation to obtain the feature map x 13 , the second encoder upsamples the feature map x 6 and the feature map x 13 by bilinear interpolation to obtain the feature map x14 , the first encoder processes the feature map x 5 and the feature map x 14 is upsampled by bilinear interpolation to obtain the feature map x 15 , the feature map x 10 , the feature map x 11 , the feature map x 12 , the feature map x 13 , the feature map x 14 and the feature map x 15 are convolved with a 3*3 filter and then upsampled back to the original image size by linear interpolation, concatenated using Concat, and passed through the Sigmoid function to obtain 32 aircraft carbon brake disc defect detection results.

[0091] In the discriminator network, the first convolutional layer, the first pooling layer, the second convolutional layer, the second pooling layer, the third convolutional layer, and the third pooling layer sequentially perform convolution and pooling on the feature map x 5 to obtain the high-dimensional feature map feature map x 6 , and the high-dimensional feature map x 6 is flattened and input into the fully connected layer to map the high-dimensional feature map to a scalar value

[0092] The weights of the U2-Net network are iteratively updated using the cross-entropy loss and the Dice loss until the total loss function of the model converges, resulting in the trained DUFG-Net model.

[0093] The total loss function of the model includes: the cross-entropy loss L BCE , the Dice loss L Dice , the generation loss and the adversarial loss and the update formulas for updating θ t , ε t are as follows:

[0094]

[0095]

[0096] where N represents the total number of pixels, y i is the true mask value of the i-th pixel of the sample image, is the generated mask value of the i-th pixel of the sample image, log represents the logarithm operation, D(y) is the output of the discriminator for the true mask, taking the value of 1, is the output of the discriminator for the generated mask, taking the value in [0,1], and respectively represent the loss values of the U2-Net network and the discriminator network, θ t-1 , εt-1 respectively represent the weight parameters of the U2-Net network and the discriminator network in the (t-1)-th iteration, and α represents the learning rate.

[0097] Step 4: Input the test set into the trained DUFG-Net model to output the detection results of the defects of the aircraft carbon brake disc.

[0098] Although the present invention has been described in detail with general descriptions and specific embodiments in this specification, on the basis of the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. A carbon brake disc defect detection method based on DUFG-Net model, characterized in that: The carbon brake disc is imaged by an X-ray digital scanner to obtain a defect image of the carbon brake disc, and a DUFG-Net model is constructed to learn defect features of the image; the steps of the method include the following: Step 1: Generate training and test sets: The K images of black and white carbon brake discs with defects acquired by X-ray digital scanners are composed into a sample set, K ≥ 200, and the images in the sample set contain at least five types of defects; after processing and annotating the images in the sample set, the images in the sample set and their corresponding labels are divided into a training set and a test set according to a ratio of 8:2; Step 2: Build a DUFG-Net model that is connected in sequence by a multi-directional feature enhancement network, a U2-Net network, and a discriminator network, and set the parameters; Step 3, train the DUFG-Net model; Step 4: Input the test set into the trained DUFG-Net model and output the detection results of carbon brake disc defects.

2. The detection method according to claim 1, characterized in that: The steps to generate the training set and test set described in step 1 are as follows: The first step is to photograph the carbon brake disc from the top of the circular surface of the carbon brake disc in an X-ray digital scanner to obtain an X-ray image of the carbon brake disc that can observe the internal quality of the carbon brake disc. The image is a black and white image. The carbon brake disc X-ray image is manually screened to obtain K defect images of the carbon brake disc with defects, K ≥ 200, including at least five defect types; In the second step, the carbon brake disc defect image is cropped to obtain a sub-image of size 256×256, and there is a 20% overlap area between the two sub-images; The third step is to annotate the carbon brake disc image with segmentation masks and save it in COCO annotation format. In the fourth step, the carbon brake disc images and their corresponding labels are combined into a training sample set M and a test sample set K, with a ratio of 8:

2.

3. The detection method according to claim 2, characterized in that: The five defect types include: circular hole defects, high-density inclusions, cracks, looseness, uneven material distribution, and delamination.

4. The detection method according to claim 2, characterized in that: The segmentation mask annotation refers to: annotating each pixel in each image, annotating the pixel value of the non-defective part as 0, and the pixel value of the defective part as 1, and the output mask image is a single-channel grayscale image.

5. The detection method according to claim 1, characterized in that: The structure of the multi-directional feature enhancement network described in step 2 is: feature extraction layer, multi-directional Gabor filter, adaptive weighting module, multi-head self-attention mechanism and feature fusion and output module; The structure of the feature extraction layer is: the first convolution layer, the batch normalization layer, the RELU activation layer, the parallel second convolution layer, the third convolution layer, the fourth convolution layer, the fifth convolution layer, the depth convolution layer, and the point-by-point convolution layer; the convolution kernel sizes of the first to fifth convolution layers, the depth convolution layer, and the point-by-point convolution layer of the feature extraction layer are set to: 3×3, 3×3, 5×5, 7×7, 1×1, 3×3, 1×1, respectively, and the number of convolution kernels is set to: 64, 64, 64, 64, 128, 128, 128, respectively, the step sizes are all set to 1, and the padding is set to 1, 1, 2, 3, 0, 1, 0 respectively; the batch normalization layer is implemented by BatchNorm2d; the RELU activation layer is implemented by the RELU function; The structure of the multi-directional Gabor filter is: a directional Gabor convolution kernel and a directional splicing and fusion submodule; the convolution kernel size of the directional adjustable Gabor convolution kernel of the multi-directional Gabor filter is set to: 5×5, the number of directions is set to 4, which are 0°, 45°, 90°, and 135° respectively, and the two-dimensional Gabor kernel is set to: G(x,y;λ,θ,φ,σ,γ)=exp(-(〖x'〗^2+γ^2〖y'〗^2) / (2σ^2))cos(2πx' / λ+φ), x'=xcosθ+ysinθ, y'=-xsinθ+ycosθ, wherein λ is the wavelength, which controls the sensitivity of the filter to image texture, θ is the direction angle, which is: 0°, 45°, 90°, and 135° respectively, φ is the phase deviation, σ is the standard deviation of the Gaussian envelope, and γ is the spatial aspect ratio; The structure of the directional splicing and fusion submodule is: feature splicing layer, convolution layer, RELU activation layer, batch normalization layer; the number of channels of the feature splicing layer of the directional splicing and fusion submodule is set to 512, the convolution kernel size of the convolution layer is set to 1×1, the number of convolution kernels is set to 512, the RELU activation layer is implemented by using the RELU function, and the batch normalization layer is implemented by using BatchNorm2d; The structure of the adaptive weighted module is: global average pooling layer, first fully connected layer, second fully connected layer; the number of neurons of the first and second fully connected layers of the adaptive weighted module are set to: 64, 4, respectively, and the activation function is set to: RELU, Softmax; the structure of the multi-head self-attention mechanism is: first linear transformation layer, batch normalization layer, second linear transformation layer; the number of channels of the first and second linear transformation layers of the multi-head self-attention mechanism are both set to 128, and the batch normalization layer is implemented using BatchNorm2d; The structure of the feature fusion and output module is: convolution layer, batch normalization layer, activation layer; the convolution kernel size of the convolution layer of the feature fusion and output module is set to 1×1, the number of convolution kernels is set to 64, the batch normalization layer is implemented using BatchNorm2d, and the activation layer is implemented using RELU function.

6. The detection method according to claim 5, characterized in that: The batch normalization layer is implemented using BatchNorm2d, which means that the mean and variance of the small batch data are normalized by the following formula: Among them, x i Describes the pixel value of the input feature map, μ represents the mean of the current batch, σ 2 represents the variance of the current batch, ∈ represents a numerical stability term, and γ and β represent learnable parameters.

7. The detection method according to claim 6, characterized in that: The structure of the U2-Net network in step 2 is the first decoder, the second decoder, the third decoder, the fourth decoder, the fifth decoder, the sixth decoder, the first encoder, the second encoder, the third encoder, the fourth encoder, the fifth encoder, and the concatenation layer. The structure of the first decoder is exactly the same as that of the first encoder, and the structure is as follows: the first convolution layer, the first batch normalization layer, the first activation layer, the second convolution layer, the second batch normalization layer, the second activation layer, the first downsampling layer, the third convolution layer, the third batch normalization layer, the third activation layer, the second downsampling layer, the fourth convolution layer, the fourth batch normalization layer, the fourth activation layer, the third downsampling layer, the fifth convolution layer, the fifth ... third activation layer, the second downsampling layer, the fourth convolution layer, the fourth batch normalization layer, the fourth activation layer, the third downsampling layer, the fifth convolution layer, the fifth batch Normalization layer, 5th activation layer, 4th downsampling layer, 6th convolution layer, 6th batch normalization layer, 6th activation layer, 5th downsampling layer, 7th convolution layer, 7th batch normalization layer, 7th activation layer, 8th convolution layer, 8th batch normalization layer, 8th activation layer, 9th convolution layer, 9th batch normalization layer, 9th activation layer, 1st upsampling layer, 10th convolution layer, 10th batch normalization layer, 10th activation layer, 2nd upsampling layer The first decoder and the first encoder are configured with the following convolution kernel sizes: 3×3, 64, 32, 16, 8, 4, 2, 2, 2, 4, 8, 16, 32, 64, respectively. The first to 14th batch normalization layers are implemented by BatchNorm2d, the first to 14th activation layers are implemented by RELU function, the sampling multiples of the first to 5 downsampling layers are set to 2, and the sampling multiples of the first to 5 upsampling layers are set to The structures of the second decoder and the second encoder are exactly the same, and compared with the first decoder and the first encoder, the structures thereof reduce the 5th downsampling layer, the 7th convolutional layer, the 7th batch normalization layer, the 7th activation layer, the 5th upsampling layer, the 14th convolutional layer, the 14th batch normalization layer, and the 14th activation layer; The convolution kernel sizes of the 12 convolutional layers in the second decoder and the second encoder are set to 3×3, and the number of convolution kernels is set to 64, 64, 32, 16, 8, 4, 4, 4, 8, 16, 32, 64 in sequence; The structures of the third decoder and the third encoder are exactly the same, and compared with the second decoder and the second encoder, the structure thereof reduces the fourth downsampling layer, the sixth convolutional layer, the sixth batch normalization layer, the sixth activation layer, the fourth upsampling layer, the thirteenth convolutional layer, the thirteenth batch normalization layer, and the thirteenth activation layer; The convolution kernel size of the convolution layer in the third decoder and the third encoder is set to 3×3, and the number of convolution kernels is set to 64, 64, 32, 16, 8, 8, 8, 16, 32, 64 in sequence; The structures of the fourth decoder and the fourth encoder are exactly the same, and compared with the third decoder and the third encoder, the structures thereof reduce the third downsampling layer, the fifth convolutional layer, the fifth batch normalization layer, the fifth activation layer, the third upsampling layer, the twelfth convolutional layer, the twelfth batch normalization layer, and the twelfth activation layer; The convolution kernel size of the convolution layer in the 4th decoder and the 4th encoder is set to 3×3, and the number of convolution kernels is set to 64, 64, 32, 16, 16, 16, 32, 64 in sequence; The structures of the 5th and 6th decoders are exactly the same as the 5th encoder, and their structures are: the 1st convolutional layer, the 1st batch normalization layer, the 1st activation layer, the 2nd convolutional layer, the 2nd batch normalization layer, the 2nd activation layer, the 3rd convolutional layer, the 3rd batch normalization layer, the 3rd activation layer, the 4th convolutional layer, the 4th batch normalization layer, the 4th activation layer, the 5th convolutional layer, the 5th batch normalization layer, the 5th activation layer, the 6th convolutional layer, the 6th batch normalization layer, the 6th activation layer, the 7th convolutional layer, the 7th batch normalization layer, the 7th activation layer, the 8th convolutional layer, the 8th batch normalization layer, the 8th activation layer; The convolution kernel sizes of the 1st to 8th convolutional layers in the 5th and 6th decoders and the 5th encoder are all set to 3×3, the number of convolution kernels is all set to 64, the 1st to 8th batch normalization layers are all implemented using BatchNorm2d, and the 1st to 8th activation layers are all implemented using RELU function.

8. The detection method according to claim 7, characterized in that: The structure of the discriminator network described in step 2 is input layer, 1st convolutional layer, 1st pooling layer, 2nd convolutional layer, 2nd pooling layer, 3rd convolutional layer, 3rd pooling layer, 1st fully connected layer, 2nd fully connected layer; the number of input channels of the input layer is set to 1, and the passable image size is set to 256×256; the convolution kernel sizes of the 1st to 3rd convolutional layers are all set to 3×3, and the number of convolution kernels is set to: 64, 128, 256, respectively; the size of the pooling kernel of the pooling layer is set to 2×2, the pooling step is set to 2, and the number of neurons in the 1st and 2nd fully connected layers are set to: 1024, 1, respectively.

9. The detection method according to claim 8, characterized in that: The training DUFG-Net model described in step 3 refers to inputting the training set into the DUFG-Net model for forward propagation, iteratively updating the weight parameters of the U2-Net network through cross entropy loss and Dice loss until the total loss function of the model converges, and obtaining a trained DUFG-Net model.

10. The detection method according to claim 9, characterized in that: The total loss function of the model includes: cross entropy loss L BCE 、Dice loss L Dice , generate loss and combat loss And for θ t , ε t The update formula for updating is: Where N represents the total number of pixels, y i is the true mask value of the i-th pixel of the sample image, is the generated mask value of the i-th pixel of the sample image, log represents the logarithmic operation, D(y) is the output of the discriminator for the true mask, and its value is 1. is the output of the discriminator to generate the mask, with a value of [0,1], and Represent the loss values ​​of the U2-Net network and the discriminator network, θ t-1 , ε t-1 They represent the weight parameters of the U2-Net network and the discriminator network in the t-1th iteration respectively, and α represents the learning rate.

Citation Information

Patent Citations

  • Product surface defect detection model and detection method based on deformable convolution

    CN111739001A

  • Watermark removing method and device, equipment, medium and product

    CN116363000A

  • Ceramic tile surface defect segmentation method based on improved U2-Net

    CN116645514A

  • Automobile brake disc surface defect detection method based on deep learning

    CN117541568A

  • Improved UNet fan blade crack detection method fused with multi-direction strip convolution

    CN118552506A