A steel pipe defect detection system and method based on image analysis
By combining the ResNet model and YOLOv8 model in the steel pipe defect detection system, deep feature and edge information in the image are extracted and fused, the problems of limited complex defect recognition capabilities and low image acquisition efficiency in the prior art are solved, and high accuracy and high efficiency steel pipe defect detection are achieved.
Patent Information
- Application Number
- CN202411428925.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-14
AI Technical Summary
In the prior art, the vision-based steel pipe defect detection system does not process it when acquiring images, and cannot fully extract and utilize deep features in the image, resulting in limited recognition ability of complex defects and low image acquisition efficiency, which makes it impossible to adapt to real-time detection scenarios.
Using a steel pipe defect detection system based on image analysis, the steel pipe images are collected in real time through the image acquisition module and pre-processed. Deep features are extracted in combination with the ResNet model, and edge pixel points are generated through canny edge detection, feature fusion and recalibration are performed, and defect location and type are finally identified in the YOLOv8 model.
It improves the accuracy and efficiency of defect detection, enhances the ability to identify different types of defects, realizes comprehensive inspection of complex defects, promotes the performance improvement of the overall solution, and realizes the intelligence of steel pipe defect detection.
Smart Images

Figure CN119048488B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and in particular to a steel pipe defect detection system and method based on image analysis. Background Art
[0002] In industrial production, steel pipes are an important basic material, and their quality directly affects the safety and performance of the products. During the production process, cracks, pits and other defects may appear on the surface of steel pipes. These defects not only affect the appearance, but also have serious safety hazards. Therefore, the rapid and accurate detection of steel pipe surface defects has become an urgent problem to be solved. Traditional steel pipe inspection mainly relies on manual inspection and simple detection tools. This method is not only time-consuming and labor-intensive, but also limited by manual experience and judgment. The detection accuracy is difficult to guarantee, and it can only cover a limited detection area, and it is impossible to achieve comprehensive inspection and actual monitoring of the entire pipeline system. Today, image analysis technology is becoming more and more mature, and detecting steel pipe defects through image acquisition and recognition has become a new and feasible method.
[0003] In the prior art, publication number CN109668907A discloses a vision-based steel pipe defect detection system, which takes a picture of the outer surface of the steel pipe through a steel pipe outer surface detection camera, and takes a picture of the inner surface of the steel pipe through a steel pipe inner surface detection camera, stores the image in an image storage module, and uses a computer to mobilize an image analysis module to analyze the collected image to identify the surface defects of the steel pipe.
[0004] The main problems with the above solution are: it relies on simple camera acquisition when collecting images, does not process the images, and cannot fully extract and utilize the deep features in the images, resulting in limited ability to recognize complex defects and difficulty in adapting to defects of different types and forms; and the image acquisition efficiency is low and cannot adapt to scenarios that require real-time detection.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not constitute the prior art that is already known to one of ordinary skill in the art. Summary of the invention
[0006] The object of the present invention is to provide a steel pipe defect detection system and method based on image analysis to solve the problems raised in the above background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A steel pipe defect detection system based on image analysis, specifically comprising:
[0009] An image acquisition module is used to acquire the steel pipe image in real time, scale the steel pipe image to a size of 224×224, convert it into a grayscale image, normalize the grayscale values of all pixels in the grayscale image, generate a first recognition image, and establish a coordinate system in the first recognition image;
[0010] The model building module is used to collect a dataset of steel pipe images with known defects, mark the steel pipe defects on each image, generate a dataset of steel pipe images with marked defects, use the dataset of steel pipe images with marked defects as input and the location of the defects as labels, and train the ResNet model.
[0011] A feature extraction module is used to input the first recognition image into the trained ResNet model to generate a ResNet feature map, perform canny edge detection on the first recognition image to generate edge pixel points of the first recognition image, and extract edge pixel points to generate a second recognition image;
[0012] A feature fusion module is used to perform weighted fusion of the ResNet feature map and the second recognition image to generate a comprehensive feature map;
[0013] The defect detection module is used to add the SE module after each convolutional layer in the YOLOv8 model, generate channel descriptors through global average pooling, generate attention weights through a two-layer fully connected network, recalibrate the comprehensive feature map through the attention weights, generate a recalibrated feature map, and input the recalibrated feature map into the YOLOv8 model to identify the defect location and type.
[0014] Furthermore, in the first recognition image, a plane rectangular coordinate system is established with the column where the leftmost pixel is located as the y-axis and the row where the bottommost pixel is located as the x-axis to ensure that each pixel has a unique plane coordinate.
[0015] Furthermore, the formula for converting the steel pipe image into a grayscale image is:
[0016] H=0.299·R+0.587·G+0.114·B
[0017] Among them, H represents the grayscale value of the pixel, R represents the red channel value, G represents the green channel value, and B represents the blue channel value.
[0018] Furthermore, the principle on which the first recognition image is subjected to canny edge detection is as follows:
[0019] For each pixel in the first recognition image, the matrix consisting of the pixel and its neighboring pixels is convolved with the horizontal template and vertical template of the Prewitt operator to generate the grayscale difference of the pixel in the horizontal and vertical directions. The formula is:
[0020]
[0021] Among them, P X represents the horizontal template of the Prewitt operator, P Y represents the vertical template of the Prewitt operator, G x Represents the horizontal difference of the pixel, G y Represents the vertical difference of the pixel point, (x, y) represents the plane coordinates of the pixel point;
[0022] The gradient amplitude of each pixel is generated based on the grayscale difference in the horizontal and vertical directions. The formula is:
[0023]
[0024] Among them, G(x,y) represents the gradient amplitude of the pixel point with coordinates (x,y), G x represents the horizontal difference of the pixel point, G y Indicates the vertical difference of the pixel;
[0025] The edge threshold is preset. When the gradient amplitude of a pixel point is higher than the edge threshold, the pixel point is retained as an edge pixel point, otherwise the pixel point is discarded.
[0026] Furthermore, the formula for generating the comprehensive feature map is:
[0027] K(x,y)=w1·F(x,y)+w2·G(x,y)
[0028] Among them, K(x,y) represents the feature value of the pixel with coordinates (x,y), F(x,y) represents the activation value of the pixel with coordinates (x,y) in the ResNet feature map, w1 represents the weight of the activation value, w2 represents the weight of the gradient amplitude, w1+w2=1 and w1=w2.
[0029] Furthermore, the formula for generating attention weights is:
[0030] s=σ(W2·δ(W1·z))
[0031] Among them, s represents the attention weight, σ represents the sigmoid activation function, δ represents the ReLU activation function, W1 represents the fully connected weight matrix that reduces the channel descriptor, W2 represents the fully connected weight matrix that restores the number of channels, and z represents the channel descriptor.
[0032] Furthermore, the formula for recalibrating the comprehensive feature map is:
[0033]
[0034] in, represents the eigenvalue of the pixel with coordinates (x, y) after recalibration, and K(x, y) represents the eigenvalue of the pixel with coordinates (x, y).
[0035] The present invention also provides a steel pipe defect detection method based on image analysis, which is performed by the above-mentioned steel pipe defect detection system based on image analysis, and the specific steps include:
[0036] Step 1: collect the steel pipe image in real time, scale the steel pipe image to 224×224 size, convert it into a grayscale image, normalize the grayscale values of all pixels in the grayscale image, generate a first recognition image, and establish a coordinate system in the first recognition image;
[0037] Step 2: Collect a dataset of steel pipe images with known defects, mark the steel pipe defects on each image, generate a dataset of steel pipe images with marked defects, use the dataset of steel pipe images with marked defects as input, and use the defect locations as labels to train the ResNet model;
[0038] Step 3: Input the first recognition image into the trained ResNet model to generate a ResNet feature map, perform canny edge detection on the first recognition image, generate edge pixels of the first recognition image, and extract edge pixels to generate a second recognition image;
[0039] Step 4: Perform weighted fusion of the ResNet feature map and the second recognition image to generate a comprehensive feature map;
[0040] Step 5: Add the SE module after each convolutional layer in the YOLOv8 model, generate channel descriptors through global average pooling, generate attention weights through a two-layer fully connected network, recalibrate the comprehensive feature map through the attention weights, generate a recalibrated feature map, and input the recalibrated feature map into the YOLOv8 model to identify the defect location and type.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] The present invention effectively improves the accuracy and efficiency of defect detection by combining deep learning models and image processing technology. The training can learn the ResNet model with rich defect features, improve the recognition ability of different types of defects, and use the deeper network structure and residual connection of the ResNet model to better capture the deep features in the image and reduce the problem of gradient disappearance, thereby improving the training effect and detection accuracy of the ResNet model; and the feature map generated by the ResNet model can extract the deep features in the steel pipe and capture complex defect patterns, while the canny edge detection provides sensitivity to significant edges. The two are weighted and fused to generate a comprehensive feature map, so that the detection process has both deep feature information and edge detail information, which greatly improves the comprehensiveness and accuracy of defect detection and promotes the performance improvement of the overall solution.
[0043] The present invention also combines the SE module with the YOLOv8 model to generate a channel descriptor, which enhances the YOLOv8 model's attention to important features, suppresses irrelevant and minor features, makes defect positioning more accurate, and combines laser sensors to automatically identify and locate defects, thereby realizing intelligent steel pipe defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a schematic diagram of a system module of an embodiment of the present invention;
[0045] Figure 2 The figure is a schematic diagram of the method flow of an embodiment of the present invention. DETAILED DESCRIPTION
[0046] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments.
[0047] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0048] Example:
[0049] See also Figure 1 , the present invention provides a technical solution:
[0050] A steel pipe defect detection system based on image analysis, specifically comprising:
[0051] An image acquisition module is used to acquire the steel pipe image in real time, scale the steel pipe image to a size of 224×224, convert it into a grayscale image, normalize the grayscale values of all pixels in the grayscale image, generate a first recognition image, and establish a coordinate system in the first recognition image;
[0052] In this embodiment, the steel pipe is scanned in real time by a laser sensor to obtain point cloud data of the steel pipe, and the point cloud is projected onto a two-dimensional plane to obtain an image of the steel pipe. Compared with directly collecting images, collection through laser sensors is not affected by light and point cloud data can be converted into two-dimensional images at any angle, providing rich perspectives.
[0053] In this embodiment, the formula for converting the steel pipe image into a grayscale image is:
[0054] H=0.299·R+0.587·G+0.114·B
[0055] Among them, H represents the grayscale value of the pixel, R represents the red channel value, G represents the green channel value, and B represents the blue channel value.
[0056] The formula for normalizing the grayscale value of pixels is:
[0057]
[0058] Among them, H0 represents the result after the grayscale value of the pixel is normalized. In the grayscale image, min and max correspond to the lowest grayscale value and the lowest grayscale value, respectively, min = 0, max = 255;
[0059] In the first recognition image, a plane rectangular coordinate system is established with the column where the leftmost pixel is located as the y-axis and the row where the bottommost pixel is located as the x-axis to ensure that each pixel has a unique plane coordinate.
[0060] The model building module is used to collect a dataset of steel pipe images with known defects, mark the steel pipe defects on each image, generate a dataset of steel pipe images with marked defects, use the dataset of steel pipe images with marked defects as input and the location of the defects as labels, and train the ResNet model.
[0061] In this embodiment, a large number of steel pipe images containing known defects are collected, the images are scaled to 224×224 and converted into grayscale images, and the defect areas in the grayscale images are annotated by Labelmg. The ResNet model uses ResNet-50, and the established ResNet-50 model structure is:
[0062] Input layer: used to receive the grayscale image of the steel pipe with marked defects;
[0063] Initial convolution layer: 7×7 convolution kernel, stride 2, 64 input channels, activated by ReLU function;
[0064] Residual module: contains 16 residual blocks in four stages: the first stage includes 3 residual blocks, each with 64 filters; the second stage includes 4 residual blocks, each with 128 filters; the third stage includes 6 residual blocks, each with 256 filters; the fourth stage includes 3 residual blocks, each with 512 filters;
[0065] Global average pooling layer: used to compress the spatial dimension of each channel of the feature map into a single value to generate a global feature vector;
[0066] Fully connected layer: locates defects based on the extracted features.
[0067] A feature extraction module is used to input the first recognition image into the trained ResNet model to generate a ResNet feature map, perform canny edge detection on the first recognition image to generate edge pixel points of the first recognition image, and extract edge pixel points to generate a second recognition image;
[0068] In this embodiment, the principle of performing canny edge detection on the first recognition image is:
[0069] For each pixel in the first recognition image, the matrix consisting of the pixel and its neighboring pixels is convolved with the horizontal template and vertical template of the Prewitt operator to generate the grayscale difference of the pixel in the horizontal and vertical directions. The formula is:
[0070]
[0071] Among them, P X represents the horizontal template of the Prewitt operator, P Y represents the vertical template of the Prewitt operator, G x Represents the horizontal difference of the pixel, G y Represents the vertical difference of the pixel point, (x, y) represents the plane coordinates of the pixel point;
[0072] The gradient amplitude of each pixel is generated based on the grayscale difference in the horizontal and vertical directions. The formula is:
[0073]
[0074] Among them, G(x,y) represents the gradient amplitude of the pixel point with coordinates (x,y), G x represents the horizontal difference of the pixel point, G y Indicates the vertical difference of the pixel;
[0075] The edge threshold is preset. When the gradient amplitude of a pixel point is higher than the edge threshold, the pixel point is retained as an edge pixel point, otherwise the pixel point is discarded.
[0076] The principle of the preset edge threshold is: combine the gradient amplitudes of all pixels to generate a histogram of the gradient amplitude, and select the initial edge threshold in the range of 70% to 90% of the histogram. A higher percentile helps to suppress noise and reduce false positives, but may miss subtle edges. A lower percentile can capture more details, but may cause too much noise to be recognized as edges. Therefore, after selecting the initial edge threshold, the edge pixels of the first recognition image are extracted, and the experts in the field judge the extraction effect of the edge pixels, and adjust the edge threshold according to the extraction effect of the edge pixels until the edge extracted according to the edge threshold meets the requirements;
[0077] A feature fusion module is used to perform weighted fusion of the ResNet feature map and the second recognition image to generate a comprehensive feature map;
[0078] In this embodiment, the formula for generating the comprehensive feature map is:
[0079] K(x,y)=w1·F(x,y)+w2·G(x,y)
[0080] Among them, K(x,y) represents the feature value of the pixel with coordinates (x,y), F(x,y) represents the activation value of the pixel with coordinates (x,y) in the ResNet feature map, w1 represents the weight of the activation value, w2 represents the weight of the gradient amplitude, w1+w2=1 and w1=w2.
[0081] The activation value of a pixel in the ResNet feature map represents the output value of each pixel obtained by calculating the input image through the ResNet model and activating it through the activation function. Each activation value reflects the importance of the corresponding pixel feature.
[0082] The comprehensive feature map generated by fusing the ResNet feature map and the second recognition image can simultaneously reflect the roughness or smoothness of the steel pipe surface, the defect shape, the brightness change of the defect part and the edge contour of the defect on the steel pipe surface. The weight is allocated by the importance of the ResNet feature map and the second recognition image. In order to consider the characteristics reflected by the two images at the same time, the same weight is set for the two, w1=w2=0.5. After calculating the characteristic value of each pixel, the pixel coordinates are arranged according to the pixel feature group to generate a comprehensive feature map.
[0083] The defect detection module is used to add the SE module after each convolutional layer in the YOLOv8 model, generate channel descriptors through global average pooling, generate attention weights through a two-layer fully connected network, recalibrate the comprehensive feature map through the attention weights, generate a recalibrated feature map, and input the recalibrated feature map into the YOLOv8 model to identify the defect location and type.
[0084] In this embodiment, the SE module represents a Squeeze-and-Excitation module, which is used to enhance the representation capability of neural network features;
[0085] The recalibrated feature map with known defect location and type is used as input. After the labels of defect location and type are marked on the feature map, it is input into the YOLOv8 model for training. The YOLOv8 model with the input value as the recalibrated feature map and the output value as the defect location and type is obtained.
[0086] The structure of the YOLOv8 model is:
[0087] Backbone network: contains convolutional layers and activation functions, which are used to extract basic features of the input image;
[0088] SE module: located after each convolutional layer, used to generate attention weights;
[0089] Feature pyramid network: used to process feature information of different scales;
[0090] Detection head: used to generate a bounding box to locate the defect position of the steel pipe and analyze the defect type.
[0091] In this embodiment, the formula for generating the attention weight is:
[0092] s=σ(w2·δ(W1·z))
[0093] Among them, s represents the attention weight, σ represents the sigmoid activation function, δ represents the ReLU activation function, W1 represents the fully connected weight matrix that reduces the channel descriptor, W2 represents the fully connected weight matrix that restores the number of channels, and z represents the channel descriptor.
[0094] The purpose of attention weight is to emphasize important features and suppress unimportant features, thereby enhancing the expressiveness of feature maps. In the SE module, attention weight is achieved by weighting the channels. W1 is a matrix that reduces the channel descriptor to reduce complexity and capture important features. The dimension is Where C represents the number of channels, r represents the scaling ratio, and r = 8 or 16, W2 is used to restore the number of channels and make the attention weight match the number of channels of the feature map, and the dimension is
[0095] The formula for recalibrating the comprehensive feature map is:
[0096]
[0097] in, represents the eigenvalue of the pixel with coordinates (x, y) after recalibration, and K(x, y) represents the eigenvalue of the pixel with coordinates (x, y).
[0098] The comprehensive feature map is recalibrated according to the attention weight, so that the feature map pays more attention to the key features of steel pipe defects. The feature map of each channel is weighted according to its corresponding attention weight. Channels with higher weights are amplified and channels with lower weights are weakened, highlighting the anomalies and defects on the surface of the steel pipe, reducing the interference of unimportant features, and allowing the model to focus on important local information.
[0099] See also Figure 2 The present invention also provides a steel pipe defect detection method based on image analysis, which is performed by the above-mentioned steel pipe defect detection system based on image analysis, and the specific steps include:
[0100] Step 1: collect the steel pipe image in real time, scale the steel pipe image to 224×224 size, convert it into a grayscale image, normalize the grayscale values of all pixels in the grayscale image, generate a first recognition image, and establish a coordinate system in the first recognition image;
[0101] Step 2: Collect a dataset of steel pipe images with known defects, mark the steel pipe defects on each image, generate a dataset of steel pipe images with marked defects, use the dataset of steel pipe images with marked defects as input, and use the defect locations as labels to train the ResNet model;
[0102] Step 3: Input the first recognition image into the trained ResNet model to generate a ResNet feature map, perform canny edge detection on the first recognition image, generate edge pixels of the first recognition image, and extract edge pixels to generate a second recognition image;
[0103] Step 4: Perform weighted fusion of the ResNet feature map and the second recognition image to generate a comprehensive feature map;
[0104] Step 5: Add the SE module after each convolutional layer in the YOLOv8 model, generate channel descriptors through global average pooling, generate attention weights through a two-layer fully connected network, recalibrate the comprehensive feature map through the attention weights, generate a recalibrated feature map, and input the recalibrated feature map into the YOLOv8 model to identify the defect location and type.
[0105] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0106] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. Those skilled in the art may appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software methods depends on the specific application and design constraints of the technical solution.
[0107] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0108] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. A steel pipe defect detection system based on image analysis, characterized in that: Specifically include: An image acquisition module is used to acquire the steel pipe image in real time, scale the steel pipe image to a size of 224×224, convert it into a grayscale image, normalize the grayscale values of all pixels in the grayscale image, generate a first recognition image, and establish a coordinate system in the first recognition image; The model building module is used to collect a dataset of steel pipe images with known defects, mark the steel pipe defects on each image, generate a dataset of steel pipe images with marked defects, use the dataset of steel pipe images with marked defects as input and the location of the defects as labels, and train the ResNet model. A feature extraction module is used to input the first recognition image into the trained ResNet model to generate a ResNet feature map, perform canny edge detection on the first recognition image to generate edge pixel points of the first recognition image, and extract edge pixel points to generate a second recognition image; A feature fusion module is used to perform weighted fusion of the ResNet feature map and the second recognition image to generate a comprehensive feature map; The defect detection module is used to add the SE module after each convolutional layer in the YOLOv8 model, generate channel descriptors through global average pooling, generate attention weights through a two-layer fully connected network, recalibrate the comprehensive feature map through the attention weights, generate a recalibrated feature map, and input the recalibrated feature map into the YOLOv8 model to identify the defect location and type; SE module stands for Squeeze-and-Excitation module, which is used to enhance the representation ability of neural network features; The recalibrated feature map with known defect locations and types is used as input. After the labels of the defect locations and types are marked on the feature map, it is input into the YOLOv8 model for training, and a YOLOv8 model is obtained whose input value is the recalibrated feature map and whose output value is the defect location and type.
2. The steel pipe defect detection system based on image analysis according to claim 1, characterized in that: In the first recognition image, a plane rectangular coordinate system is established with the column where the leftmost pixel is located as the y-axis and the row where the bottommost pixel is located as the x-axis to ensure that each pixel has a unique plane coordinate.
3. The steel pipe defect detection system based on image analysis according to claim 1, characterized in that: The formula used in the image acquisition module to convert the steel pipe image into a grayscale image is: H=0.299·R+0.587·G+0.114·B Among them, H represents the grayscale value of the pixel, R represents the red channel value, G represents the green channel value, and B represents the blue channel value.
4. The steel pipe defect detection system based on image analysis according to claim 1, characterized in that: The principle on which the canny edge detection is performed on the first recognition image in the feature extraction module is based on: For each pixel in the first recognition image, the matrix consisting of the pixel and its neighboring pixels is convolved with the horizontal template and vertical template of the Prewitt operator to generate the grayscale difference of the pixel in the horizontal and vertical directions. The formula is: Among them, P X represents the horizontal template of the Prewitt operator, P Y represents the vertical template of the Prewitt operator, G x Represents the horizontal difference of the pixel, G y represents the vertical difference of the pixel point, and (x, y) represents the plane coordinates of the pixel point; The gradient amplitude of each pixel is generated based on the grayscale difference in the horizontal and vertical directions. The formula is: Among them, G(x, y) represents the gradient amplitude of the pixel point with coordinates (x, y), G x represents the horizontal difference of the pixel point, G y Indicates the vertical difference of the pixel; The edge threshold is preset. When the gradient amplitude of a pixel point is higher than the edge threshold, the pixel point is retained as an edge pixel point, otherwise the pixel point is discarded.
5. The steel pipe defect detection system based on image analysis according to claim 4 is characterized in that: The formula for generating the comprehensive feature map is: K(x,y)=w1·F(x,y)+w2·G(x,y) Among them, K(x, y) represents the feature value of the pixel with coordinates (x, y), F(x, y) represents the activation value of the pixel with coordinates (x, y) in the ResNet feature map, w1 represents the weight of the activation value, w2 represents the weight of the gradient amplitude, w1+w2=1 and w1=w2.
6. The steel pipe defect detection system based on image analysis according to claim 1, characterized in that: The formula for generating attention weights in the defect detection module is: s=σ(W2·δ(W1·z)) Among them, s represents the attention weight, σ represents the sigmoid activation function, δ represents the ReLU activation function, W1 represents the fully connected weight matrix that reduces the channel descriptor, W2 represents the fully connected weight matrix that restores the number of channels, and z represents the channel descriptor.
7. The steel pipe defect detection system based on image analysis according to claim 1, characterized in that: The formula for recalibrating the comprehensive feature map in the defect detection module is: in, represents the eigenvalue of the pixel with coordinates (x, y) after recalibration, and K(x, y) represents the eigenvalue of the pixel with coordinates (x, y).
8. A steel pipe defect detection method based on image analysis, characterized in that: The method is performed by the steel pipe defect detection system based on image analysis according to any one of claims 1 to 7, and the specific steps include: Step 1: collect the steel pipe image in real time, scale the steel pipe image to 224×224 size, convert it into a grayscale image, normalize the grayscale values of all pixels in the grayscale image, generate a first recognition image, and establish a coordinate system in the first recognition image; Step 2: Collect a dataset of steel pipe images with known defects, mark the steel pipe defects on each image, generate a dataset of steel pipe images with marked defects, use the dataset of steel pipe images with marked defects as input, and use the defect locations as labels to train the ResNet model; Step 3: Input the first recognition image into the trained ResNet model to generate a ResNet feature map, perform canny edge detection on the first recognition image, generate edge pixels of the first recognition image, and extract edge pixels to generate a second recognition image; Step 4: Perform weighted fusion of the ResNet feature map and the second recognition image to generate a comprehensive feature map; Step 5: Add the SE module after each convolutional layer in the YOLOv8 model, generate channel descriptors through global average pooling, generate attention weights through a two-layer fully connected network, recalibrate the comprehensive feature map through the attention weights, generate a recalibrated feature map, and input the recalibrated feature map into the YOLOv8 model to identify the defect location and type.
Citation Information
Patent Citations
Vision-based steel pipe defect detecting system
CN109668907A
Visual building crack recognition method based on attention mechanism and ResNet fusion
CN112734739A
Weld joint surface defect detection method and system, electronic equipment and storage medium
CN116823819A