Lightweight industrial defect detection method
By constructing the C2P-YOLO model, combining the lightweight upsampling module and the dual attention mechanism, the problem of large model scale and slow detection speed in wind power tower defect detection is solved, and high-precision and efficient defect detection are achieved.
Patent Information
- Application Number
- CN202510219135.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
AI Technical Summary
In the detection of wind power tower defects, the problem of large model scale, slow detection speed and difficulty in achieving high-precision detection in complex and dense defect scenarios.
A lightweight industrial defect detection method is proposed, and feature extraction and object detection are performed by constructing a C2P-YOLO model, including input network, backbone network, neck network and head network, combined with a lightweight upsampling module and a dual attention mechanism.
It significantly improves the ability to detect tiny defects and nuances of the surface of wind power towers, improves detection speed and accuracy, and reduces the calculation amount and memory cost.
Smart Images

Figure CN120219804A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and specifically to a lightweight industrial defect detection method. Background Art
[0002] Wind power generation technology has developed vigorously with the increasing global demand for clean energy. However, the growth in the scale and quantity of wind power generation equipment has also brought serious damage problems. As the support structure of the entire wind power equipment, the wind power tower barrel is often eroded by various external factors such as wind force, ultraviolet radiation, and temperature changes in the harsh natural environment, resulting in problems such as corrosion and cracks on the surface of the wind power tower, affecting its stability and safety. At present, the defect detection of wind power tower barrels mainly relies on manual detection, which is time-consuming and laborious and has limited accuracy, making it difficult to meet the requirements of modern industry for efficient and accurate detection.
[0003] Object detection technology is one of the research hotspots in the field of computer vision and has been widely used in practical application scenarios such as face recognition, vehicle tracking, and industrial defect detection. Object detection technology can detect the defects existing in the images of the surface of the wind power tower barrel, thus helping practitioners to simply and efficiently complete the maintenance work of the wind power tower barrel. Industrial object detection methods can be divided into two categories: traditional object detection and object detection based on deep learning. However, with the development of technology and data, traditional object detection algorithms can no longer meet the application requirements in industry due to low time efficiency, poor robustness, and accuracy, and object detection technology based on deep learning has become the focus of research in the field of computer vision.
[0004] At present, researchers have proposed various algorithms for industrial defect object detection technology based on deep learning. A Faster R-CNN model based on feature fusion and cascade detection network can accurately detect the types and positions of defects on the surface of steel, but the large scale and complex structure of its model lead to great limitations in actual application and poor detection speed; an industrial defect detection model based on lightweight YOLOv4 introduces depthwise separable convolution in the Neck part, greatly reducing the number of model parameters and improving the detection speed, but this algorithm is difficult to perform high-precision classification and detection for complex and dense defects of wind power tower barrels.
[0005] Therefore, developing a lightweight and high-precision industrial defect detection model has become an urgent need in the current industrial field. Summary of the Invention
[0006] (1) Technical Problems to be Solved
[0007] In view of the deficiencies of the prior art, the present application provides a lightweight industrial defect detection method.
[0008] (2) Technical Solutions
[0009] To solve the above problems, the present application provides the following technical solutions:
[0010] A lightweight industrial defect detection method, comprising:
[0011] Obtaining surface image data of a wind power tower barrel;
[0012] Preprocessing the surface image data of the wind power tower barrel and establishing an industrial defect image dataset; the preprocessing at least includes image enhancement, normalization, and size adjustment;
[0013] Constructing C2P-YOLO, where the C2P-YOLO includes an input network Input, a backbone network Backbone, a neck network Neck, and a head network Head;
[0014] Training the C2P-YOLO with the industrial defect image dataset and optimizing the C2P-YOLO through a loss function;
[0015] Applying the trained and optimized C2P-YOLO to the wind power tower barrel scenario for real-time defect detection.
[0016] Preferably, the obtaining of the surface image data of the wind power tower barrel specifically includes: collecting surface defect images of the wind power tower barrel through an industrial camera.
[0017] Preferably, the constructing of the C2P-YOLO specifically includes:
[0018] Determining input data information and receiving the preprocessed surface image data of the wind power tower barrel through the input network Input;
[0019] Designing the backbone network Backbone, extracting features through a C2P module, and improving the feature extraction ability through a multi-branch stacking module;
[0020] Designing the neck network Neck and performing feature recombination and fusion through a CARAFE upsampling module;
[0021] Designing the head network Head and adaptively integrating local features and global dependencies through a dual attention mechanism.
[0022] Preferably, the designing of the backbone network Backbone, extracting features through a C2P module, and improving the feature extraction ability through a multi-branch stacking module specifically includes:
[0023] Dividing the initial feature map into the first branch and the second branch;
[0024] The initial feature map enters the first branch, passes through the standard convolutional layer to generate a first intermediate feature map, and the first intermediate feature map enters the Bottleneck unit to reduce the number of channels by half through 1×1 convolution, and then doubles the number of channels through 3×3 convolution to extract features to obtain the first branch feature map;
[0025] The initial feature map enters the second branch, passes through the standard convolutional layer to obtain the second branch feature map;
[0026] The first branch feature map and the second branch feature map are concatenated in channels to obtain a first fused feature map; the first fused feature map is input into the lightweight convolution Pconv to reduce the convolution operation to obtain the output feature map.
[0027] Preferably, the lightweight convolution Pconv sets the partial rate to 1 / x, and only applies conventional convolution to 1 / x of the input channels for spatial feature extraction, keeping the remaining channels unchanged, reducing the number of floating-point operation counts FLOPs.
[0028] Preferably, the neck network Neck is designed to perform feature recombination and fusion through the CARAFE upsampling module, specifically including:
[0029] The CARAFE upsampling module adaptively learns and predicts the upsampling kernel through the prediction unit, and uses the feature recombination unit to complete the upsampling operation of the feature map;
[0030] The shapes of the input and output feature maps in the prediction unit are:
[0031] I = W×H×C (1)
[0032] O = αW×αH×C (2)
[0033] In formulas (1) and (2), I is the input feature map, O is the output feature map, W is the width of the feature map, H is the length of the feature map, C is the number of channels of the feature map, and α is the upsampling ratio;
[0034] In the upsampling kernel prediction, the size of the upsampling kernel is k up , the number of channels of the feature map is reduced to C through a 1×1 convolutional layer m , the upsampling kernel is predicted through a convolutional layer of size ke×ke, and a new upsampling kernel is obtained by unfolding in the channel dimension. The size of the new upsampling kernel is: αH×αW×k up ×k up ;
[0035] The new upsampling kernel is normalized using the softmax function. For each pixel point of the output feature map, it is mapped back to the input feature map, and k centered on the pixel point is takenup × k up and perform a dot product operation with the predicted new upsampling kernel to obtain the upsampling result;
[0036] The calculation formula of the weighted sum operator in the feature recombination unit is:
[0037]
[0038] In formula (3), r is the radius of the convolution kernel, X(i + n, j + m) is the pixel value in the input feature map, i and j are the position coordinates of the current pixel point, n and m are the offsets of the current pixel point, W is the weight, and X l ” is the value of the recombined pixel point.
[0039] Preferably, the head network Head is designed to adaptively integrate local features and global dependencies through a dual attention mechanism, specifically including:
[0040] i. Use second-order attention pooling for feature fusion to aggregate features in the entire input space. The mathematical calculation expression is:
[0041] G bilinear (A, B) = AB T = ∑ i a i b i T (4)
[0042] In formula (4), a i and b i are the feature vectors from feature maps A and B respectively. Gbilinear is the output of the second-order attention pooling, that is, a set containing visual primitives;
[0043] ii. Distribute the aggregated features to each position of the input. The subsequent convolutional layer can perceive global information. The mathematical calculation expression of the feature distribution is:
[0044] z i = ∑ j v ij g j = G gather (X)v i (5)
[0045] In formula (5), z i is the new local feature at position i, v i is the local feature vector at position i, g j is the feature vector aggregated from the first step, G gather (X) is the aggregation function, and X is the input feature map;
[0046] iii. Combine the calculations in steps i and ii. The mathematical calculation expression of the dual attention mechanism is as follows:
[0047] Z = F distr (G gather (X), V) = G gather (X) ⊙ softmax(ρ(X; W ρ )) (6)
[0048] In formula (6), Z is the output tensor, V is the attention weight vector normalized by the softmax function, and ρ(X; W ρ ) is used to generate the attention weight, and ⊙ represents element-wise multiplication.
[0049] Preferably, training the C2P-YOLO with the industrial defect image dataset and optimizing the C2P-YOLO with a loss function specifically include:
[0050] Load the processed industrial defect image dataset into memory, and divide the industrial defect image dataset into a training set: validation set ratio of 7:3;
[0051] Initialize C2P-YOLO and set the training hyperparameters, where the training hyperparameters at least include the learning rate, batch size, number of iterations, and optimizer;
[0052] Input the image into C2P-YOLO for feature extraction and object detection, and output the prediction results;
[0053] Calculate the loss function value based on the difference between the prediction results and the true labels; update the parameters through backpropagation;
[0054] After each training epoch, evaluate the model performance using the validation set and save the optimal model parameters.
[0055] (III) Beneficial Effects
[0056] Compared with the prior art, the present application provides a lightweight industrial defect detection method, which has the following beneficial effects:
[0057] 1. The industrial defect detection model adaptively integrates local features and global dependencies through a lightweight upsampling module and a dual attention mechanism module to help the model fuse deep and shallow features to detect tiny defects on the surface of the wind turbine tower and the subtle differences between defects, significantly improving the detection ability of tiny defects.
[0058] 2. In the industrial defect detection model, the C2P module integrates the change of gradient into the feature map from beginning to end according to the idea of splitting the flow extracted by CSPNet, which not only enhances the learning ability of convolution, but also ensures the accuracy while reducing the computational amount, and reduces the computational bottleneck and memory cost.
[0059] 3. The industrial defect detection model replaces ordinary convolution with lightweight convolution Pconv, which can extract spatial features more effectively while reducing redundant calculations and access volume.
[0060] Some of the additional aspects and advantages of this application will be given in the following description, some will become obvious from the following description, or will be learned through the practice of this application. Description of the Drawings
[0061] The above and / or additional aspects and advantages of this application will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:
[0062] Figure 1 is a schematic flowchart of a lightweight industrial defect detection method of this application;
[0063] Figure 2 is the overall structure diagram of a lightweight industrial defect detection method of this application;
[0064] Figure 3 is the structure diagram of the C2P module of a lightweight industrial defect detection method of this application;
[0065] Figure 4 is the process diagram of the lightweight convolution Pconv of a lightweight industrial defect detection method of this application. Detailed Embodiments
[0066] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts shall fall within the protection scope of this application.
[0067] The terms "first" and "second" in the specification and claims of this application may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, the meaning of "a plurality" is two or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the front and rear related objects.
[0068] In the description of the present application, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present application.
[0069] In the description of the present application, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", "connected to" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0070] Please refer to Figures 1 - 4 , the present application provides a new technical solution: a lightweight industrial defect detection method, including:
[0071] Obtain the surface image data of the wind power tower barrel;
[0072] Preprocess the surface image data of the wind power tower barrel and establish an industrial defect image dataset; the preprocessing at least includes image enhancement, normalization, and size adjustment;
[0073] Construct C2P-YOLO, and the C2P-YOLO includes an input network Input, a backbone network Backbone, a neck network Neck, and a head network Head;
[0074] Train the C2P-YOLO with the industrial defect image dataset and optimize the C2P-YOLO through a loss function;
[0075] Apply the trained and optimized C2P-YOLO to the wind power tower barrel scenario for real-time defect detection.
[0076] In the present invention, the obtaining of the surface image data of the wind power tower barrel specifically includes: collecting the surface defect images of the wind power tower barrel through an industrial camera.
[0077] In the present invention, the construction of the C2P-YOLO specifically includes:
[0078] Determine the input data information, and receive the pre - processed surface image data of the wind turbine tower through the input network Input;
[0079] Design the backbone network Backbone, perform feature extraction through the C2P module, and improve the feature extraction ability through the multi - branch stacking module;
[0080] Design the neck network Neck, and perform feature recombination and fusion through the CARAFE up - sampling module;
[0081] Design the head network Head, and adaptively integrate local features and global dependencies through the dual attention mechanism.
[0082] In the present invention, the design of the backbone network Backbone, performing feature extraction through the C2P module and improving the feature extraction ability through the multi - branch stacking module specifically includes:
[0083] Split the initial feature map into the first branch and the second branch;
[0084] The initial feature map enters the first branch, passes through the standard convolutional layer to generate the first intermediate feature map, and the first intermediate feature map enters the Bottleneck unit. The number of channels is reduced by half through 1×1 convolution, and then doubled through 3×3 convolution to extract features and obtain the first - branch feature map;
[0085] The initial feature map enters the second branch, passes through the standard convolutional layer to obtain the second - branch feature map;
[0086] Perform concat channel splicing on the first - branch feature map and the second - branch feature map to obtain the first fusion feature map; input the first fusion feature map into the lightweight convolution Pconv to reduce the convolution operation and obtain the output feature map.
[0087] In the present invention, the lightweight convolution Pconv sets the partial rate to 1 / x, only applies conventional convolution to 1 / x part of the input channels for spatial feature extraction, keeps the remaining channels unchanged, and reduces the number of floating - point operation counts FLOPs.
[0088] In the present invention, the design of the neck network Neck, performing feature recombination and fusion through the CARAFE up - sampling module specifically includes:
[0089] The CARAFE up - sampling module adaptively learns and predicts the up - sampling kernel through the prediction unit, and uses the feature recombination unit to complete the up - sampling operation of the feature map;
[0090] The shapes of the input and output feature maps in the prediction unit are:
[0091] I = W × H × C (1)
[0092] O = αW × αH × C (2)
[0093] In formulas (1) and (2), I is the input feature map, O is the output feature map, W is the width of the feature map, H is the length of the feature map, C is the number of channels of the feature map, and α is the upsampling ratio;
[0094] In upsampling kernel prediction, the upsampling kernel size is k up , reduce the number of channels of the feature map to C through a 1×1 convolutional layer m , predict the upsampling kernel through a convolutional layer of size ke×ke, and expand it in the channel dimension to obtain a new upsampling kernel, and the size of the new upsampling kernel is: αH × αW × k up ×k up ;
[0095] Use the softmax function to normalize the new upsampling kernel. For each pixel point of the output feature map, map it back to the input feature map, take a k up ×k up region, and perform a dot product operation with the predicted new upsampling kernel to obtain the upsampling result;
[0096] The calculation formula of the weighted sum operator in the feature recombination unit is:
[0097]
[0098] In formula (3), r is the radius of the convolutional kernel, X(i + n, j + m) is the pixel value in the input feature map, i and j are the position coordinates of the current pixel point, n and m are the offsets of the current pixel point, W is the weight, and X l ” is the value of the recombined pixel point.
[0099] In the present invention, the head network Head is designed to adaptively integrate local features and global dependencies through a dual attention mechanism, specifically including:
[0100] i. Use second-order attention pooling for feature fusion to aggregate features of the entire input space, and the mathematical calculation expression is:
[0101] G bilinear (A, B) = AB T = ∑ i a i b i T (4)
[0102] In formula (4), a i and b iThey are the feature vectors from feature maps A and B respectively, and Gbilinear is the output of the second-order attention pooling, that is, a set containing visual primitives;
[0103] ii. Distribute the aggregated features to each position of the input. The subsequent convolutional layer can perceive global information. The mathematical calculation expression of the feature distribution is:
[0104] z i = ∑ j v ij g j = G gather (X)v i (5)
[0105] In formula (5), z i is the new local feature at position i, v i is the local feature vector at position i, g j is the feature vector aggregated from the first step, G gather (X) is the aggregation function, and X is the input feature map;
[0106] iii. Combine the calculations in steps i and ii. The mathematical calculation expression of the dual attention mechanism is:
[0107] Z = F distr (G gather (X), V) = G gather (X) ⊙ softmax(ρ(X; W ρ )) (6)
[0108] In formula (6), Z is the output tensor, V is the attention weight vector normalized by the softmax function, ρ(X; W ρ ) is used to generate attention weights, and ⊙ represents element-wise multiplication.
[0109] In a specific embodiment, formula (4) represents aggregating the interaction information between the input feature maps A and B through second-order attention pooling to generate a feature set containing global visual primitives; formula (5) represents distributing the feature vector g j aggregated from the first step to each position i to generate a new local feature z i , so that the subsequent convolutional layer can perceive global information.
[0110] In a specific embodiment, the prediction unit includes a channel compressor, a content encoder, and a kernel normalizer; the channel compressor is a 1×1 convolution that reduces the channels of the input feature map to reduce subsequent computational complexity; the content encoder is a kernel of size k up × k upThe upsampling kernel uses different upsampling kernels at different positions, and the shape of the upsampling kernel is αH×αW×k up ×k up , that is, for the feature map output by the channel compressor, a convolution layer with ke×ke is used to predict the upsampling kernel to obtain a new upsampling kernel; the kernel normalizer normalizes the upsampling kernel using the softmax function.
[0111] In the present invention, training the C2P-YOLO with the industrial defect image dataset and optimizing the C2P-YOLO through a loss function specifically include:
[0112] Loading the processed industrial defect image dataset into memory and dividing the industrial defect image dataset into a training set: a validation set in a ratio of 7:3;
[0113] Initializing C2P-YOLO and setting training hyperparameters, where the training hyperparameters at least include learning rate, batch size, number of iterations, and optimizer;
[0114] Inputting the image into C2P-YOLO for feature extraction and object detection, and outputting the prediction result;
[0115] Calculating the loss function value according to the difference between the prediction result and the true label; updating the parameters through backpropagation;
[0116] After each training cycle, evaluating the model performance using the validation set and saving the optimal model parameters.
[0117] In a specific embodiment, to evaluate the detection performance of the C2P-YOLO model, the C2P-YOLO model is compared with current popular defect detection methods. In this paper, the NEU-DET dataset is used to compare with the classical detection model Faster RCNN+FPN and YOLO series detection models respectively. The experimental result comparison table is shown in Table 1; in Table 1, P is the accuracy rate, R is the recall rate, mAP@0.5 is the mean average precision calculated under the condition that the IoU threshold is 0.5, and Para / M is the number of model parameters.
[0118] Table 1 Experimental comparison result table
[0119] Model P% R% mAP%@0.5 Para / M Faster RCNN+FPN 70.4 68.1 71.3 97.7 YOLOX 69.1 71.5 73.5 193 YOLOV5X 75.4 73.7 77.6 164 YOLOV7 75.9 69.3 75.7 78.3 YOLOV8l 70.3 70.4 72.5 83.5 CSL-YOLO 88.0 75.1 84.9 58.4
[0120] It can be seen from Table 1 that the map score of the CXP-YOLO model reaches 84.9%, which has an improvement of nearly 10% compared with most mainstream defect detection algorithms, and the detection performance is very prominent.
[0121] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0122] Although the embodiments of this application have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of this application, and the scope of this application is defined by the appended claims and their equivalents.
Claims
1. A lightweight industrial defect detection method, characterized in that: include: Obtaining wind turbine tower surface image data; Preprocessing the wind turbine tower surface image data and establishing an industrial defect image data set; the preprocessing at least includes image enhancement, normalization and size adjustment; Construct C2P-YOLO, wherein C2P-YOLO includes an input network Input, a backbone network Backbone, a neck network Neck, and a head network Head; The C2P-YOLO is trained by the industrial defect image dataset, and the C2P-YOLO is optimized by a loss function; The trained and optimized C2P-YOLO is applied to the wind tower scenario for real-time defect detection.
2. A lightweight industrial defect detection method according to claim 1, characterized in that: The obtaining of wind turbine tower surface image data specifically includes: collecting wind turbine tower surface defect images by an industrial camera.
3. A lightweight industrial defect detection method according to claim 1, characterized in that: The construction of C2P-YOLO specifically includes: Determine input data information, and receive pre-processed wind tower surface image data through the input network Input; Design the backbone network Backbone, extract features through the C2P module, and improve the feature extraction capability through the multi-branch stacking module; Design the neck network Neck, and perform feature reorganization and fusion through the CARAFE upsampling module; The head network Head is designed to adaptively integrate local features and global dependencies through a dual attention mechanism.
4. A lightweight industrial defect detection method according to claim 3, characterized in that: The backbone network Backbone is designed to perform feature extraction through the C2P module and improve the feature extraction capability through the multi-branch stacking module, specifically including: Splitting the initial feature map into the first branch and the second branch; The initial feature map enters the first branch, passes through the standard convolution layer to generate a first intermediate feature map, the first intermediate feature map enters the Bottleneck unit, reduces the number of channels by half through 1×1 convolution, and then doubles the number of channels through 3×3 convolution, extracts features to obtain a first branch feature map; The initial feature map enters the second branch and passes through the standard convolution layer to obtain a second branch feature map; The first branch feature map and the second branch feature map are concat-channeled to obtain a first fused feature map; the first fused feature map is input into the lightweight convolution Pconv to reduce the convolution operation to obtain an output feature map.
5. A lightweight industrial defect detection method according to claim 4, characterized in that: The lightweight convolution Pconv sets the partial rate to 1 / x, applies conventional convolution only to the 1 / x part of the input channel for spatial feature extraction, keeps the remaining channels unchanged, and reduces the number of floating-point operations FLOPs.
6. A lightweight industrial defect detection method according to claim 3, characterized in that: The design of the neck network Neck, through the CARAFE upsampling module to perform feature reorganization and fusion, specifically includes: The CARAFE upsampling module adaptively learns and predicts the upsampling kernel through the prediction unit, and uses the feature recombination unit to complete the upsampling operation of the feature map; The shapes of the input and output feature maps in the prediction unit are: I=W×H×C (1) O=αW×αH×C (2) In formulas (1) and (2), I is the input feature map, O is the output feature map, W is the width of the feature map, H is the length of the feature map, C is the number of channels of the feature map, and α is the upsampling ratio; In the upsampling kernel prediction, the upsampling kernel size is k up , the number of feature map channels is reduced to C through a 1×1 convolution layer m , predict the upsampling kernel through the convolution layer of size ke×ke, and expand it in the channel dimension to obtain a new upsampling kernel, the size of which is: αH×αW×k up ×k up ; The new upsampling kernel is normalized using the softmax function. For each pixel in the output feature map, it is mapped back to the input feature map and k pixels centered on the pixel are taken. up ×k up The area is then dot-producted with the predicted new upsampling kernel to obtain the upsampling result. The calculation formula of the price weight and operator in the feature recombination unit is: In formula (3), r is the radius of the convolution kernel, X(i+n,j+m) is the pixel value in the input feature map, i and j are the position coordinates of the current pixel, n and m are the offsets of the current pixel, W is the weight, and X l ” is the value of the reorganized pixel.
7. A lightweight industrial defect detection method according to claim 3, characterized in that: The design of the head network Head adaptively integrates local features and global dependencies through a dual attention mechanism, specifically including: i. Use second-order attention pooling for feature fusion to aggregate the features of the entire input space. The mathematical calculation expression is: G bilinear (A,B)=AB T =Σ i a i b i T (4) In formula (4), a i and b i are the feature vectors from feature maps A and B respectively, and Gbilinear is the output of the second-order attention pooling, that is, a set of visual primitives; ii. Distribute the aggregated features to each position of the input. The subsequent convolutional layer can perceive the global information. The mathematical calculation expression of feature distribution is: z i =Σ j v ij g j =G gather (X)v i (5) In formula (5), z i is the new local feature at position i, v i is the local eigenvector of position i, g j is the feature vector obtained from the first step of aggregation, G gather (X) is the aggregation function, X is the input feature map; iii. Combining steps i and ii, the mathematical expression of the dual attention mechanism is: Z=F distr (G gather (X),V)=G gather (X)⊙softmax(ρ(X;W ρ )) (6) In formula (6), Z is the output tensor, V is the attention weight vector normalized by the softmax function, and ρ(X; W ρ ) is used to generate attention weights, and ⊙ represents element-wise multiplication.
8. A lightweight industrial defect detection method according to claim 1, characterized in that: The training of the C2P-YOLO by the industrial defect image dataset and the optimization of the C2P-YOLO by the loss function specifically include: Loading the processed industrial defect image dataset into a memory, and dividing the industrial defect image dataset into a training set and a validation set in a ratio of 7:3; Initialize C2P-YOLO and set training hyperparameters, wherein the training hyperparameters include at least learning, batch size, number of iterations, and optimizer; Input the image into C2P-YOLO, perform feature extraction and target detection, and output the prediction result; Calculate the loss function value based on the difference between the predicted result and the true label; update the parameters through back propagation; After each training cycle, the model performance is evaluated using the validation set and the optimal model parameters are saved.