Construction method of surface defect detection model, mechanical arm grabbing classification method, computer system and computer readable storage medium

By fusing the features of RGB images and heat maps in the surface defect detection model, and using the bar-enhanced convolution module and attention mechanism, the problem of poor detection of elongated defects in traditional models is solved, achieving higher detection accuracy.

CN120070319AActive Publication Date: 2025-05-30ZHENGZHOU UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411995534.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Traditional target detection models are more suitable for regular-shaped targets, and lack targeting detection of slender defects such as cracks and scratches, and are prone to missed or missed detection.

Method used

A surface defect detection model construction method is adopted to integrate the features of RGB images and heat maps through a dual-branch backbone network, and use the bar enhancement convolution module and attention mechanism to enhance the capture ability of bar features.

Benefits of technology

This method can detect elongated defects more accurately, reduce the occurrence of missed or missed detection, and improve the accuracy of defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070319A_ABST
    Figure CN120070319A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of industrial image defect detection, and particularly relates to a construction method of a surface defect detection model, a mechanical arm grabbing classification method, a computer system and a computer readable storage medium. Comprising the following steps: 1) through a backbone network, obtaining initial feature maps of different scales of each layer corresponding to an RGB image and a thermodynamic diagram, and carrying out feature fusion; obtaining an adjustment feature map of each layer through a neck network; 2) processing different adjustment feature maps through a bar-shaped enhanced convolution module, and inputting the processed adjustment feature maps into different detection heads; updating the model parameters and then repeating the steps 1)-2) until iteration is stopped; the processing mode of the bar-shaped enhanced convolution module comprises the following steps: adjusting the feature map to obtain a horizontal enhanced feature map and a vertical enhanced feature map through a horizontal bar-shaped convolution layer and a vertical bar-shaped convolution layer respectively; and fusing the horizontal and vertical enhanced feature maps with the features obtained by processing the attention mechanism element by element, fusing the two groups of fusion results, and then fusing the fused results with the adjusted feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial image defect detection, and particularly relates to a method for constructing a surface defect detection model, a robotic arm grasping and classification method, a computer system, and a computer-readable storage medium. Background Art

[0002] With the continuous evolution of industrial production, steel, as a key structural material, is widely used in various infrastructure and engineering projects. However, during the manufacturing and processing of steel products, they are often affected by various factors, resulting in diverse defects on their surfaces, which seriously damages the quality and service life of steel products. Therefore, the rapid and accurate detection of steel surface defects has become a crucial link to ensure product quality and safety.

[0003] Combining robotic arm grasping with machine vision enables the robotic arm grasping system to have the capabilities of intelligent perception and intelligent decision-making, and can improve the adaptability of the robotic arm grasping system in complex environments. The robotic arm vision guidance system is a robotic arm motion control system that does not require manual participation and performs defect detection through computer vision detection combined with neural networks. Currently, in the industrial field, mainstream robotic arms usually complete the grasping and classification tasks of target objects through complex point-to-point teaching. Although the teaching method has advantages such as easy operation and low cost, this method may waste a lot of time. To solve the problem of difficult classification caused by whether the object to be grasped has defects, enabling the robotic arm to autonomously detect whether industrial products have defects and then achieve sorting has extremely important research value.

[0004] CAM visualizes the confidence or importance of different regions in an image in the form of a color-coded heatmap, where brightness indicates that the region has a higher response to the network and makes a greater contribution. The heatmap generated by CAM intuitively shows the focus and regions of interest of the model, providing convenience for understanding. Therefore, by borrowing the idea of CAM to generate a heatmap, with the help of the generated heatmap, the missing defect location information in the RGB image can be supplemented, and through deep learning methods, combining the features of the heatmap and the RGB image to achieve defect detection can effectively improve the recall rate of defects, thereby achieving more accurate defect detection results.

[0005] Currently, although the defect detection method based on deep learning shows significant advantages, it faces various restrictions. The defect detection method based on deep learning mainly faces two major problems: First, due to objective reasons such as lighting conditions and imaging device limitations during the data collection process, and factors such as the irregular boundaries of defect targets, some defects exhibit weak features and low contrast, which makes defect detection more difficult and often results in missed detections. Second, different from traditional object detection categories, there are a large number of slender defects in the defect categories, such as cracks and scratches. These types of defects have extreme aspect ratios, while traditional object detection models are more suitable for objects with regular shapes and lack pertinence in detecting slender defects, making it easier to miss or misdetect. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for constructing a surface defect detection model, a robotic arm grasping and classification method, a computer system, and a computer-readable storage medium, which are used to solve the problem that in the prior art, traditional object detection models are more suitable for objects with regular shapes and lack pertinence in detecting slender defects, making it easier to miss or misdetect.

[0007] To achieve the above purpose, the present invention provides a method for constructing a surface defect detection model, including:

[0008] 1) Input the RGB images and heat map samples in the training set into the surface defect detection model to be trained. Through the dual-branch backbone network of the surface defect detection model, initial feature maps of different scales corresponding to the RGB images and heat maps are obtained respectively, and the initial feature maps corresponding to the RGB images and heat maps of the same scale are subjected to feature fusion; then, through the neck network of the defect detection model, adjusted feature maps of each layer are obtained according to the fusion result;

[0009] 2) After processing different adjusted feature maps through the bar enhancement convolution module respectively, input them into different detection heads of the head network of the defect detection model to obtain the detection results of defect positions and classifications; after updating the parameters of the defect detection model through the loss function, repeat 1)-2) until the iteration stops to complete the construction of the defect detection model;

[0010] The method for processing the adjusted feature maps through the bar enhancement convolution module includes:

[0011] The adjusted feature maps are respectively passed through at least one horizontal strip convolution layer and at least one vertical strip convolution layer to obtain horizontal and vertical enhanced feature maps; the horizontal and vertical strip convolution layers respectively contain convolution kernels of sizes 1×k and k×1, where k is greater than 1; the horizontal enhanced feature map is fused with the feature obtained by processing it through the attention module, the vertical enhanced feature map is fused with the feature obtained by processing it through the attention module, and the two groups of fusion results are fused and then fused with the adjusted feature map itself.

[0012] Further, the way of fusing the horizontal enhanced feature map with the feature obtained by processing it through the attention module includes:

[0013] The horizontal enhanced feature map, the feature obtained by processing the horizontal enhanced feature map through the spatial attention module, and the feature obtained by processing the horizontal enhanced feature map through the channel attention module are multiplied element by element for fusion.

[0014] Further, the way of fusing the vertical enhanced feature map with the feature obtained by processing it through the attention module includes:

[0015] The vertical enhanced feature map, the feature obtained by processing the vertical enhanced feature map through the spatial attention module, and the feature obtained by processing the vertical enhanced feature map through the channel attention module are multiplied element by element for fusion.

[0016] Further, the way of fusing the initial feature maps corresponding to the RGB image and the heat map of the same scale includes:

[0017] After the result of element-wise subtraction of the initial feature maps corresponding to the RGB image and the heat map of the same scale is processed through the spatial attention module, it is respectively multiplied element by element with the initial feature maps corresponding to the RGB image and the heat map of the same scale, and then added element by element and processed through a convolution layer containing a 1×1 convolution kernel.

[0018] Further, the way of respectively obtaining the initial feature maps of different scales corresponding to the RGB image and the heat map through the dual-branch backbone network of the surface defect detection model includes:

[0019] Downsample the heatmap layer by layer through the branch corresponding to the heatmap in the dual-branch backbone network to obtain initial feature maps of different scales corresponding to the heatmap; downsample the RGB image layer by layer through the branch corresponding to the RGB image in the dual-branch backbone network until the downsampling result is the smallest initial feature map with a scale greater than the upper threshold corresponding to the set scale range; after fusing the smallest initial feature map with the initial feature map corresponding to the heatmap of the same scale, continue to downsample layer by layer. Each subsequent downsampling is performed on the fusion result of the downsampling result of the previous layer and the initial feature map corresponding to the heatmap of the same scale. According to the results of layer-by-layer downsampling, initial feature maps of different scales corresponding to the heatmap are obtained.

[0020] Further, at least one of the convolutional kernels in the at least one horizontal strip convolutional layer and the at least one vertical strip convolutional layer is a depth convolution.

[0021] Further, the lower threshold corresponding to the set scale range is the scale of the feature map obtained after downsampling the input image of the defect detection model by 5 layers, and the upper threshold corresponding to the set scale range is the scale of the feature map obtained after downsampling the input image of the defect detection model by 3 layers.

[0022] Further, both the horizontal strip convolutional layer and the vertical strip convolutional layer are less than or equal to 3.

[0023] Further, both branches in the dual-branch backbone network of the defect detection model adopt the MobileNetV3 network.

[0024] The above technical solution provides a new method for constructing a surface defect detection model. The beneficial effects include: on the basis of the feature map that fuses the features in the heatmap and the features in the RGB image at the same time, before the detection head performs localization and classification, first process the feature map to be input into the detection head through horizontal strip convolution and vertical strip convolution respectively to obtain a horizontal enhanced feature map and a vertical enhanced feature map. The horizontal enhanced feature map more accurately captures the features of strip defects in the horizontal direction, and the vertical enhanced feature map more accurately captures the features of strip defects in the vertical direction. Fusing the two can obtain the features containing the global information of strip defects in both the horizontal and vertical directions at the same time, enhancing the learning ability of the neural network for strip features; then fuse with the feature map (adjusted feature map) to be input into the detection head itself to ensure that the local details of strip defects are not lost; finally, combine the attention mechanism to enhance the attention ability of the defect detection model to the region most relevant to the features of strip defects in the image, enhancing the capture effect of strip convolution on strip features; therefore, the defect detection model obtained by this method can focus on more accurate detection of slender defects, reducing the occurrence of missed detection or misdetection.

[0025] The present invention also provides a robotic arm grasping and classification method, including: acquiring a product image to be detected, inputting the product image to be detected into a constructed defect detection model, and obtaining detection results of the defect position and defect type; sending corresponding instructions to the controller of the robotic arm according to the detection results of the defect position and defect type, so that after receiving the instructions, the controller of the robotic arm controls the robotic arm to perform grasping and classification actions corresponding to the detection results of the defect position and defect type;

[0026] The defect detection model is constructed by the above-mentioned method for constructing a surface defect detection model.

[0027] The above-mentioned robotic arm grasping and classification method of the present invention can achieve the same beneficial effects as the above-mentioned method for constructing a surface defect detection model.

[0028] The present invention also provides a computer system, where the processor is used to execute executable program instructions, and the executable program instructions are used to be executed to implement the above-mentioned robotic arm grasping and classification method.

[0029] The above-mentioned computer system of the present invention can achieve the same beneficial effects as the above-mentioned robotic arm grasping and classification method.

[0030] The present invention also provides a computer-readable storage medium, where computer program instructions are stored in the storage medium, and the computer program instructions are used to implement the above-mentioned robotic arm grasping and classification method when being executed.

[0031] The above-mentioned computer-readable storage medium of the present invention can achieve the same beneficial effects as the above-mentioned robotic arm grasping and classification method. Description of the Drawings

[0032] Figure 1 It is a schematic diagram of the structural principle of the surface defect detection model in the embodiment of the method for constructing a surface defect detection model of the present invention;

[0033] Figure 2 It is a schematic diagram of the principle of the bar enhancement convolution module of the surface defect detection model in the embodiment of the method for constructing a surface defect detection model of the present invention;

[0034] Figure 3 It is a schematic diagram of the principle of feature fusion of the initial feature maps corresponding to the RGB image and the heat map of the same scale in the embodiment of the method for constructing a surface defect detection model of the present invention;

[0035] Figure 4 It is an example diagram of the defect detection result of the surface defect detection model in the embodiment of the method for constructing a surface defect detection model of the present invention.

[0036] Figure 5This is a schematic diagram of the principle of the robotic arm grasping and classification method in the embodiment of the robotic arm grasping and classification method of the present invention. Detailed implementation manners

[0037] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0038] Embodiment of the method for constructing a surface defect detection model

[0039] This embodiment provides a technical solution for a method for constructing a surface defect detection model. By means of strip convolution, the capture of global features of strip defects in the horizontal and vertical directions is enhanced, and in combination with an attention mechanism, the attention ability of the defect detection model to the region in the image that is most relevant to the features of strip defects is enhanced, that is, the capture effect of strip convolution on strip features is enhanced, so that the constructed surface defect detection model better adapts to the shape of strip defects, and more accurate detection effects can be obtained for strip defects in different directions, effectively reducing the missed detection and false detection of strip defects.

[0040] Specifically, referring to Figure 1 , the method specifically includes:

[0041] 1) Input the RGB images and heat map samples in the training set into the surface defect detection model to be trained (which can also be simply referred to as the defect detection model). Through the dual-branch backbone network of the surface defect detection model, initial feature maps of different scales corresponding to the RGB images and heat maps are obtained respectively, and feature fusion is performed on the initial feature maps corresponding to the RGB images and heat maps of the same scale ( Figure 1 The Fusion module in is an algorithm module for performing feature fusion on the initial feature maps corresponding to the RGB images and heat maps of the same scale, which can be implemented by the special feature fusion method of this embodiment or by other existing feature fusion methods); then through the neck network of the defect detection model, adjusted feature maps of each layer are obtained according to the fusion result;

[0042] 2) After processing different adjusted feature maps through the strip enhancement convolution module respectively, input them into different detection heads of the head network of the defect detection model to obtain detection results of defect positions and classifications; update the parameters of the defect detection model through the loss function, and then repeat 1)-2) until the iteration stops to complete the construction of the defect detection model;

[0043] Referring to Figure 2 , the manner of processing the adjusted feature maps through the strip enhancement convolution module includes:

[0044] The adjusted feature map is respectively passed through at least one horizontal strip convolutional layer and at least one vertical strip convolutional layer to obtain horizontal and vertical enhanced feature maps; the horizontal and vertical strip convolutional layers respectively contain convolutional kernels of sizes 1×k and k×1, where k>1; the horizontal enhanced feature map is fused with the feature obtained by processing it through the attention module, the vertical enhanced feature map is fused with the feature obtained by processing it through the attention module, and the two sets of fusion results are fused and then fused with the adjusted feature map itself.

[0045] After analysis, traditional convolutional neural networks usually perform convolutional operations using square convolutional kernels, which perform well in many scenarios. However, convolutional kernels of this shape limit the receptive field of the network, causing it to focus on square local detail information. That is, in the presence of long strip-shaped targets in the defect dataset, square convolutional kernels are difficult to capture long-distance feature information and may also introduce feature information of irrelevant regions, reducing the detection accuracy; although using a larger convolutional kernel can increase the receptive field, it may increase the computational amount and learning of background features; using deformable convolutions to adaptively adjust the shape and position of the convolutional kernel can better capture object details, but it will lead to an increase in computational complexity and parameters, reducing the detection speed; therefore, the method for constructing the surface defect detection model in this embodiment, on the basis of the feature map that simultaneously fuses the features in the heat map and the features in the RGB image, before the detection head performs localization and classification, first processes the feature map to be input into the detection head through horizontal strip convolution and vertical strip convolution respectively to obtain horizontal and vertical enhanced feature maps. The horizontal enhanced feature map more accurately captures the features of strip-shaped defects in the horizontal direction, and the vertical enhanced feature map more accurately captures the features of strip-shaped defects in the vertical direction. Fusing the two can simultaneously obtain the features containing global information of strip-shaped defects in both the horizontal and vertical directions, enhancing the neural network's learning ability for strip features; then it is fused with the feature map (adjusted feature map) to be input into the detection head itself to ensure that the local details of the strip-shaped defects are not lost; finally, combined with the attention mechanism, the attention ability of the defect detection model to the region most relevant to the features of the strip-shaped defects in the image is enhanced, enhancing the capture effect of the strip convolution on the strip features; therefore, the defect detection model obtained by this method can focus on more accurate detection of slender defects, reducing the occurrence of missed detections or false detections.

[0046] Among them, as Figure 1 shown by the blue background area bounded by the dashed box in

[0047] The ways to obtain the initial feature maps of different scales corresponding to the RGB image and the heat map respectively through the dual-branch backbone network of the surface defect detection model include: Figure 1 in the branch corresponding to the heat map in the dual-branch backbone network (i.e., the branch corresponding to the initial feature maps such asFigure 1 in the Heatmap) for downsampling layer by layer ( Figure 1 The MB module in represents the sampling algorithm corresponding to downsampling. Specifically, in this embodiment, it is the convolutional structure and sampling algorithm in the MobileNetv3 Block), and initial feature maps of different scales corresponding to the heatmap are obtained Through the branch corresponding to the RGB image in the dual-branch backbone network (i.e., Figure 1 in the branch corresponding to the initial feature maps such as) the RGB image is downsampled layer by layer until the downsampling result is the smallest initial feature map whose scale is greater than the upper threshold corresponding to the set scale range; in this embodiment, since the upper threshold corresponding to the set scale range is the scale of the feature map obtained by downsampling the input image of the defect detection model by 3 layers. For example, the corresponding scale, so this smallest initial feature map is before is directly obtained by downsampling the RGB image, is also directly obtained by downsampling ; Next, this smallest initial feature map is fused with the initial feature map corresponding to the heatmap of the same scale After fusion (in this embodiment, the way to fuse this smallest initial feature map with the initial feature map corresponding to the heatmap of the same scale is to add them element by element. In other embodiments, other existing fusion methods can be used), then continue to downsample layer by layer. Each subsequent downsampling is performed on the fusion result of the downsampling result of the previous layer and the initial feature map corresponding to the heatmap of the same scale. For example, for and After fusion, the downsampling result is obtained. The obtained downsampling result is then fused with the initial feature map corresponding to the heatmap of the same scale And for and After fusion, the downsampling result is obtained; Then, according to the results of layer-by-layer downsampling (including the and obtained by the layer-by-layer downsampling method of direct layer-by-layer downsampling and the and obtained by the layer-by-layer downsampling method of fusing the initial feature maps corresponding to the heatmaps of the same scale and then downsampling) the initial feature maps of different scales corresponding to the heatmap can be obtained

[0048] Thus, the feature information of the RGB image and the heat map can be fully extracted and fused through the above backbone network structure; after analysis, the RGB image and the heat map carry information on different aspects of the product surface respectively, and the heat map can provide more accurate position information for the RGB image. Therefore, the effective fusion of these two information sources can achieve more comprehensive and rich information acquisition; the design of the backbone network structure enables it to effectively extract the respective feature information of the RGB image and the heat map, and fuse the two at the feature level during the initial feature extraction process, so as to achieve better performance and effects in the detection task. Therefore, it can overcome the differences between different information sources, thus providing more powerful support for the performance of the deep learning model in complex localization and classification tasks.

[0049] Moreover, in this embodiment, both branches of the dual-branch backbone network of the defect detection model adopt the MobileNetV3 network, that is, a lightweight network is used to form the backbone network for feature extraction, reducing the number of parameters of the backbone network, thereby improving the training efficiency and reducing the training cost. In other embodiments, backbone networks with other structures can also be used.

[0050] To balance the situation where the feature map scale is too small and a large amount of feature information is lost due to too deep downsampling levels, and the feature map scale is too large and too much useless noise and redundant information are introduced due to too shallow downsampling levels, in this embodiment, the feature fusion of the initial feature maps corresponding to the RGB image and the heat map of the same scale is the feature fusion of the initial feature maps corresponding to the RGB image and the heat map of the same scale within a set scale range; and the lower threshold corresponding to the set scale range is the scale of the feature map obtained after 5-layer downsampling of the image input to the defect detection model, and the upper threshold corresponding to it is the scale of the feature map obtained after 3-layer downsampling of the image input to the defect detection model. Usually, the scale of the feature map obtained after 5-layer downsampling of the image input to the defect detection model is 20×20 (only the scale of a single feature map), and the scale of the feature map obtained after 3-layer downsampling of the image input to the defect detection model is 80×80 (only the scale of a single feature map). Selecting the scale of the feature map used for defect localization and defect classification of the detection head within this range can take into account the effect of relatively more useful feature information and relatively less interfering noise and redundant information; that is, selecting the scale and the adjusted feature map P 3 the same as the initial feature map C 3 obtained by 3-layer downsampling of the image input to the defect detection model by the scale and the backbone network 4 the same as the adjusted feature map P 4and the initial feature map C obtained after 5 - layer downsampling of the images of the input defect detection model by the scale and backbone network 5 the same adjusted feature map P 5 They are respectively input into different detection heads to obtain the detection results of defect localization and classification. It should be noted that the number of detection heads adopted by the surface defect detection model in this embodiment needs to be determined according to the number of feature maps selected for defect localization and defect classification of the detection heads.

[0051] In addition, in this embodiment, after the backbone network extracts the initial feature map, a special feature fusion method is adopted for the initial feature maps corresponding to the RGB image and the heat map of the same scale, that is, the specially designed Figure 1 Fusion module in; Refer to Figure 3 , the method of feature fusion for the initial feature maps corresponding to the RGB image and the heat map of the same scale includes:

[0052] The initial feature map corresponding to the RGB image of the same scale and the initial feature map corresponding to the heat map The result of element - by - element subtraction After being processed by the spatial attention module, they are respectively multiplied element - by - element with the initial feature maps corresponding to the RGB image and the heat map of the same scale, and then added element - by - element and processed by a convolutional layer containing a 1×1 convolutional kernel (i.e., Figure 3 Conv1×1 in), to obtain the feature fusion result C i . Since the feature fusion of the initial feature maps corresponding to the RGB image and the heat map of the same scale is mainly to enhance the spatial features of weak defects, only the spatial attention module is used for processing. In this embodiment, since the spatial attention module is mostly implemented by max - pooling (MaxPooling), average - pooling (AvgPooling) and convolutional layers with large - scale convolutional kernels, the process of being processed by the spatial attention module here is represented as Figure 3 MaxPool and AvgPool in and the subsequent Conv7×7.

[0053] Specifically, the above - mentioned method of feature fusion for the initial feature maps corresponding to the RGB image and the heat map of the same scale can be expressed by the following formula:

[0054]

[0055] In the formula, SA refers to the processing of the spatial attention (Spatial Attention) module, and Conv 1×1 refers to the convolutional processing with a 1×1 convolutional kernel (i.e., Figure 3(in Conv1×1); denotes element-wise addition, denotes element-wise multiplication.

[0056] In addition, in this embodiment, the neck network part for feature fusion adopts a structure similar to YOLOv5, and improves the representation ability of the model through a top-down and bottom-up two-way fusion mechanism. Since the structure of the adopted neck network belongs to the prior art, it will not be elaborated here.

[0057] Specifically, referring to Figure 2 , in the bar enhancement convolution module of this embodiment, the ways of fusing the horizontal enhancement feature map with the feature obtained by processing it through the attention module include:

[0058] Multiply the horizontal enhancement feature map, the feature obtained by processing the horizontal enhancement feature map through the spatial attention module, and the feature obtained by processing the horizontal enhancement feature map through the channel attention module element-wise for fusion; fuse after adopting the spatial attention module and the channel attention module. Since the spatial attention focuses on the spatial dimensions in the image, that is, the width and height of the image. Its main purpose is to identify important regions in the image and assign higher weights to these regions. This mechanism helps the model focus on the key visual elements in the image and ignore the background or other unimportant regions. Therefore, the spatial attention mechanism enables the model to focus on the most important spatial positions in the defect image for the detection task; while the channel attention focuses on the feature channel dimension of the image, that is, the set of different features; its purpose is to identify which feature channels are important and enhance the feature responses of these channels. Therefore, the channel attention mechanism enables the model to focus on the most important channel features in the defect image for the classification task. These two attention mechanisms can be combined with each other to jointly improve the performance of the model.

[0059] In other embodiments, only one of the spatial attention module or the channel attention module can also be used to process the horizontal enhancement feature map and then fuse it with the horizontal enhancement feature map.

[0060] Similarly, referring to Figure 2 , the ways of fusing the vertical enhancement feature map with the feature obtained by processing it through the attention module include:

[0061] Multiply the vertical enhancement feature map, the feature obtained by processing the vertical enhancement feature map through the spatial attention module, and the feature obtained by processing the vertical enhancement feature map through the channel attention module element-wise for fusion. In other embodiments, only one of the spatial attention module or the channel attention module can also be used to process the vertical enhancement feature map and then fuse it with the vertical enhancement feature map.

[0062] ThroughFigure 1 The SECM module shown by the green background area bounded by the dashed box in the figure is Figure 2 Taking the bar enhanced convolution module in i as an example, the specific example of the method for processing the adjusted feature map P is as follows:

[0063] The adjusted feature map P i respectively obtains the horizontal enhanced feature map through at least one horizontal bar convolution layer and at least one vertical bar convolution layer and the vertical enhanced feature map where the horizontal and vertical bar convolution layers respectively contain convolution kernels of sizes 1×k and k×1, k>1, and the Sigmoid activation function is used after the convolution kernel; in this embodiment, 2 horizontal bar convolution layers and 2 vertical bar convolution layers are taken. In fact, 1 horizontal bar convolution layer and 1 vertical bar convolution layer can also achieve the effect of the bar enhanced convolution module at the principle level. Through experimental analysis, the effect of taking 2 horizontal bar convolution layers and 2 vertical bar convolution layers is the best, while in the case of taking more than 3 horizontal bar convolution layers and more than 3 vertical bar convolution layers, the performance has begun to grow slowly, resulting in the increase in performance being difficult to match the increase in the cost to be invested. Therefore, considering the computational efficiency and computational cost, both the horizontal bar convolution layer and the vertical bar convolution layer are less than or equal to 3.

[0064] After that, the horizontal enhanced feature map is multiplied element by element with the feature of scale H×W×1 obtained by processing it through the spatial attention module and the feature of scale 1×1×C obtained by processing it through the channel attention module (the multiplication result is ), the vertical enhanced feature map is multiplied element by element with the feature of scale H×W×1 obtained by processing it through the spatial attention module and the feature of scale 1×1×C obtained by processing it through the channel attention module (the multiplication result is ), the two groups of multiplication results are fused and then fused with the adjusted feature map P i itself to obtain the processing result S of the bar enhanced convolution module i . The above process can be expressed by the following formula:

[0065]

[0066] In the formula, CA and SA respectively refer to the processing of the channel attention (Channel Attention) module and the spatial attention (Spatial Attention) module, Conv refers to the convolution processing of the 1×1 convolution kernel, which is used to improve the effect of feature fusion, and this convolution processing can be not carried out in other embodiments.

[0067] Considering that although using strip convolution can capture more long-distance spatial context information without adding too much computational complexity, thereby enhancing the feature extraction ability of the network, it will result in a relatively large number of channels in the adjusted feature map output by the Neck part (i.e., the neck network) of the surface defect detection model. Therefore, in order to control the model parameter quantity, in this embodiment, depth convolution is used to reduce the parameter quantity of strip convolution, so at least one of the convolutional kernels in at least one horizontal strip convolutional layer and at least one vertical strip convolutional layer is a depth convolution. Depth convolution can effectively reduce the number of channels, thereby achieving the effect of controlling the number of channels of the feature map input to the detection head. Theoretically, the more depth convolutions are adopted, the better the effect of reducing the number of channels. Therefore, in this embodiment, taking the case of 2 horizontal strip convolutional layers and 2 vertical strip convolutional layers as an example, the adjusted feature map P i Obtain the horizontal enhanced feature map through at least one horizontal strip convolutional layer and at least one vertical strip convolutional layer respectively and the vertical enhanced feature map It is expressed by the formula as follows:

[0068]

[0069] In the formula, DWConv 1×k () represents a horizontal strip convolutional layer with a depthwise separable convolutional kernel, and DWConv k×1 () represents a vertical strip convolutional layer with a depthwise separable convolutional kernel; the above formula represents that the convolutional kernels of the horizontal strip convolutional layer for obtaining the horizontal enhanced feature map and the vertical strip convolutional layer for obtaining the vertical enhanced feature map both adopt depth convolution. In this embodiment, an example of the defect detection result of the surface defect detection constructed by the above construction method is referred to Figure 5 .

[0070] In this embodiment, the channel attention module is specifically:

[0071] CA(F) = sigmoid(Avgpool(F) + Maxpool(F))

[0072] In the formula, CA represents the channel attention module; F represents the feature map to be processed; sigmoid is an activation function used for normalization; Avgpool is average pooling, and Maxpool is max pooling;

[0073] The spatial attention module is specifically:

[0074] SA(F) = sigmoid(Avgpool(F) + Maxpool(F))

[0075] In the formula, SA represents the spatial attention module; F represents the feature map to be processed; sigmoid is an activation function used for normalization; Avgpool is average pooling, and Maxpool is max pooling; in other embodiments, other channel attention modules and spatial attention modules can also be used.

[0076] Embodiment of the robotic arm grasping and classification method

[0077] This embodiment provides a robotic arm grasping and classification method. Referring to Figure 4 , the method specifically includes:

[0078] Obtain the product image to be detected, and input the product image to be detected (i.e., the visual image in Figure 4 ) into the constructed defect detection model to obtain the detection results of the defect position and defect type; send corresponding instructions (i.e., the feedback information in Figure 4 ) to the controller of the robotic arm according to the detection results of the defect position and defect type, so that the controller of the robotic arm controls the robotic arm to perform the grasping and classification actions corresponding to the detection results of the defect position and defect type after receiving the instruction.

[0079] The defect detection model in this embodiment (i.e., the defect detection network guided by position information in Figure 4 , and the content of this part is consistent with the content in the blue background area bounded by the dashed line in Figure 1 ) is obtained by the construction method of the surface defect detection model in the construction method embodiment of the surface defect detection model as described above.

[0080] Since the specific working principle and effect of the surface defect detection model constructed in the robotic arm grasping and classification method in this embodiment have been described in detail in the construction method embodiment of the surface defect detection model above, it will not be elaborated here.

[0081] Embodiment of the computer system

[0082] This embodiment provides a technical solution of a computer system. The computer system includes a processor, and executable program instructions are stored in the processor. The executable program instructions are used to be executed to implement the robotic arm grasping and classification method in the robotic arm grasping and classification method embodiment above.

[0083] Since the specific working principle and effect of the computer system in this embodiment have been described in detail in the robotic arm grasping and classification method embodiment, it will not be elaborated here.

[0084] Embodiment of the computer-readable storage medium

[0085] This embodiment provides a technical solution for a computer-readable storage medium. Computer program instructions are stored in the storage medium, and when the computer program instructions are executed, they implement the robotic arm grasping and classification method in the above-described embodiment of the robotic arm grasping and classification method.

[0086] Since the specific working principle and effects of the computer-readable storage medium in this embodiment have been described in detail in the above-described embodiment of the robotic arm grasping and classification method, they will not be elaborated here.

[0087] It should be understood that the above specific embodiments of the present invention are only for illustrative or explanatory purposes of the principles of the present invention, and do not constitute a limitation on the present invention.

Claims

1. A method for constructing a surface defect detection model, characterized in that: include: 1) Inputting the RGB images and thermal map samples in the training set into the surface defect detection model to be trained, obtaining the initial feature maps of different scales of each layer corresponding to the RGB images and thermal maps through the dual-branch backbone network of the surface defect detection model, and performing feature fusion on the initial feature maps corresponding to the RGB images and thermal maps of the same scale; and then obtaining the adjusted feature maps of each layer according to the fusion results through the neck network of the defect detection model; 2) After processing different adjusted feature maps through the strip enhancement convolution module, they are respectively input into different detection heads of the head network of the defect detection model to obtain the detection results of defect location and classification; after updating the parameters of the defect detection model through the loss function, 1)-2) are repeated until the iteration stops to complete the construction of the defect detection model; The ways to process the adjusted feature map through the strip-enhanced convolution module include: The feature map is adjusted to obtain horizontal and vertical enhanced feature maps through at least one horizontal strip convolution layer and at least one vertical strip convolution layer respectively; The horizontal and vertical strip convolution layers contain convolution kernels of 1×k and k×1 respectively, where k is greater than 1. The horizontal enhanced feature map is fused with its features processed by the attention module, and the vertical enhanced feature map is fused with its features processed by the attention module. The two sets of fusion results are fused and then fused with the adjusted feature map itself.

2. The method for constructing a surface defect detection model according to claim 1, characterized in that: Ways to fuse the horizontal enhanced feature map with its features processed by the attention module include: The horizontal enhanced feature map, the features obtained by processing the horizontal enhanced feature map through the spatial attention module, and the features obtained by processing the horizontal enhanced feature map through the channel attention module are multiplied element by element for fusion.

3. The method for constructing a surface defect detection model according to claim 1, characterized in that: Ways to fuse the vertical enhanced feature map with its features processed by the attention module include: The vertical enhanced feature map, the feature obtained by processing the vertical enhanced feature map through the spatial attention module, and the feature obtained by processing the vertical enhanced feature map through the channel attention module are multiplied element by element for fusion.

4. The method for constructing a surface defect detection model according to any one of claims 1 to 3, characterized in that: Methods for fusing features of the initial feature maps corresponding to the RGB images and heat maps of the same scale include: The result of element-by-element subtraction of the initial feature map corresponding to the RGB image and the heat map of the same scale is processed by the spatial attention module, and then multiplied element-by-element with the initial feature map corresponding to the RGB image and the heat map of the same scale, and then added element-by-element and processed by a convolution layer containing a 1×1 convolution kernel.

5. The method for constructing a surface defect detection model according to any one of claims 1 to 3, characterized in that: The methods of obtaining the initial feature maps of different scales of each layer corresponding to the RGB image and the heat map through the dual-branch backbone network of the surface defect detection model include: The heat map is downsampled layer by layer through the branch corresponding to the heat map in the dual-branch backbone network to obtain initial feature maps of different scales at each layer corresponding to the heat map; the RGB image is downsampled layer by layer through the branch corresponding to the RGB image in the dual-branch backbone network until the downsampling result is a minimum initial feature map whose scale is greater than the upper limit threshold corresponding to the set scale range; after fusing the minimum initial feature map with the initial feature map corresponding to the heat map of the same scale, downsampling layer by layer is continued, and each subsequent downsampling layer is downsampled by the fusion result of the downsampling result of the previous layer and the initial feature map corresponding to the heat map of the same scale, and the initial feature maps of different scales at each layer corresponding to the heat map are obtained according to the results of downsampling layer by layer.

6. The method for constructing a surface defect detection model according to any one of claims 1 to 3, characterized in that: At least one of the convolution kernels in the at least one horizontal strip convolution layer and the at least one vertical strip convolution layer is a depthwise convolution.

7. The method for constructing a surface defect detection model according to any one of claims 1 to 3, characterized in that: The feature fusion of the initial feature maps corresponding to the RGB image and the heat map of the same scale is the feature fusion of the initial feature maps corresponding to the RGB image and the heat map of the same scale within the set scale range; the lower limit threshold corresponding to the set scale range is the scale corresponding to the feature map obtained after 5-layer downsampling of the image input to the defect detection model, and the corresponding upper limit threshold is the scale corresponding to the feature map obtained after 3-layer downsampling of the image input to the defect detection model.

8. The method for constructing a surface defect detection model according to any one of claims 1 to 3, characterized in that: The number of the horizontal strip convolutional layers and the vertical strip convolutional layers is less than or equal to 3.

9. The method for constructing a surface defect detection model according to any one of claims 1 to 3, characterized in that: Both branches in the dual-branch backbone network of the defect detection model adopt the MobileNetV3 network.

10. A robot arm grasping and classification method, characterized in that: include: Obtain an image of a product to be inspected, input the image of the product to be inspected into the constructed defect detection model, and obtain the detection results of the defect location and defect type; Sending corresponding instructions to the controller of the robot arm according to the detection results of the defect position and defect type, so that the controller of the robot arm controls the robot arm to perform grasping and classification actions corresponding to the detection results of the defect position and defect type after receiving the instructions; The defect detection model is constructed by the surface defect detection model construction method according to any one of claims 1 to 9.

11. A computer system comprising a processor, wherein the processor is configured to execute executable program instructions, wherein: The executable program instructions are used to be executed to implement the robot arm grasping classification method according to claim 10.

12. A computer-readable storage medium, wherein computer program instructions are stored in the storage medium, characterized in that: The computer program instructions are used to implement the robot arm grasping classification method as claimed in claim 10 when executed.

Citation Information

Patent Citations

  • Copper plate surface defect detection and automatic classification method based on machine vision and deep learning

    CN113070240A

  • Shielding pedestrian detection method based on low-parameter attention mechanism

    CN115761810A

  • Attention scheme and stripe convolution semantic line detection method based on deep Hough network

    CN116563682A

  • Multi-modal fusion ground stain identification method and system based on attention mechanism

    CN117649579A

  • Weak light target detection method based on infrared and visible light image fusion

    CN118135200A