Part identification method in factory environment based on deep learning
By improving the Backbone, Neck and Head parts of the YOLO11n model, combined with a variety of part image data sets in the factory environment, the problem of low accuracy of part recognition in industrial manufacturing is solved, and the part recognition effect with high robustness and low latency is achieved.
Patent Information
- Application Number
- CN202510356881.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
In the field of industrial manufacturing, part identification has problems such as insufficient generalization capability, high real-time requirements, high model deployment costs, and high noise interference in complex factory environments, resulting in low recognition accuracy.
A improved deep learning algorithm based on the YOLO11n model is constructed, and feature extraction and detection are optimized by introducing CGDown, DySample and DyHead modules in Backbone, Neck and Head parts, combined with a variety of part image data sets in the factory environment.
It improves the accuracy of part recognition, especially in occlusion and reflection, and realizes high robustness and real-time performance of lightweight models, and is suitable for complex industrial scenarios.
Smart Images

Figure CN120259631A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a method for recognizing parts in a factory environment based on deep learning. Background Art
[0002] In the field of industrial manufacturing, there is an urgent need for rapid and accurate recognition of components in automated production lines. Traditional part recognition technologies mainly rely on machine vision systems, which are realized by manually designing features (such as edge contours, textures, or color template matching) combined with classification algorithms (such as support vector machines, random forests). However, such methods have significant limitations in complex factory environments: First, the extraction of manual features is sensitive to changes in light intensity, specular reflection on the part surface, occlusion, and pose diversity, resulting in insufficient generalization ability; Second, when facing the iterative update of part types on the production line, it is necessary to redesign the feature template, significantly increasing the debugging and maintenance costs. In recent years, image recognition technologies based on deep learning (such as convolutional neural networks) have demonstrated superior performance in natural scene object detection tasks through end-to-end feature learning. However, the particularity of industrial scenarios poses higher constraints on the algorithm: 1) The production line has strict requirements for real-time performance. Existing high-precision models (such as ResNet, Faster R-CNN) have a large number of parameters and are difficult to be efficiently deployed on edge computing devices; 2) The cost of obtaining industrial part sample data is high, and under the condition of small samples, it is easy to cause model overfitting; 3) There are metal specular reflections, oil stains interference, and multi-view imaging differences in the factory environment, and the robustness of models trained on existing public datasets has decreased significantly.
[0003] Although existing research has attempted to solve the above problems through transfer learning, data augmentation, or lightweight networks (such as MobileNet), the above optimization paths have the following problems: Lightweight models are difficult to balance accuracy and speed, transfer learning has insufficient adaptation to cross-domain features, and the authenticity of traditional data augmentation strategies for simulating industrial noise is limited. Therefore, the existing technology has the problem of low average accuracy in recognizing parts on the factory production line. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention provides a method for recognizing parts in a factory environment based on deep learning, which improves the accuracy of part recognition for situations such as occlusion and specular reflection through a highly robust and low-latency part recognition model optimized for the complex factory environment.
[0005] The technical solution adopted by the present invention is as follows:
[0006] The present invention provides a method for recognizing parts in a factory environment based on deep learning, including:
[0007] Collect images of various parts in the factory environment and construct an image dataset including different sample groups, which includes occluded samples, reflective samples, and multi-specification mixed samples;
[0008] Construct an improved deep learning algorithm model based on the YOLO11n model:
[0009] In the Backbone part, replace the ordinary convolution modules in the 1st, 3rd, 5th, and 7th layers of the YOLO11n model with CGDown modules;
[0010] In the Neck part, replace the ordinary upsampling modules in the 11th and 14th layers of the YOLO11n model with DySample modules, and concatenate the output of the 11th layer DySample module after replacement with the 6th layer C3k2 module and input it into the 12th layer C3k2 module, and concatenate the output of the 14th layer DySample module after replacement with the 4th layer C3k2 module and input it into the 16th layer C3k2 module; replace the ordinary convolution modules in the 17th and 20th layers with the CGDown modules, and concatenate the output of the 17th layer CGDown module after replacement with the 13th layer C3k2 module and input it into the 19th layer C3k2 module, and concatenate the output of the 20th layer CGDown module after replacement with the 10th layer C2PSA module and input it into the 22nd layer C3k2 module;
[0011] In the Head part, introduce three DyHead detection modules to replace the original ordinary Head detection heads respectively, so as to perform object detection on the feature maps of different sizes output by different layers in the Neck part;
[0012] Use the image dataset to train and validate the improved deep learning algorithm model to obtain an identification model;
[0013] Use the identification model to identify the target object in the image to be measured.
[0014] The further technical solution is:
[0015] For the CGDown module, after adjusting the number of channels of the feature map through a pre-convolution module with a convolution kernel of 3×3 and a stride of 2, it is divided into two parts. The two parts are respectively subjected to dilated convolution and ordinary convolution feature extraction and then spliced; the spliced feature map passes through a normalization layer and the SiLu function, and then is separated into two parts through a post-convolution module with a convolution kernel of 1×1 and a stride of 1. One part is processed through average pooling and two fully connected layers, added to the other part and output.
[0016] In the Backbone part of the front convolution module, each layer of the CGDown module and the C3k2 module connected thereto constitute a stage. After each stage, the number of channels of the model doubles and the feature map is reduced to half of the original.
[0017] The DySample module includes a sampling point generator, a sampling set and grid sampling, for a given upsampling scale factor s and a feature map of size C×H×W The processing process of the DySample module is as follows:
[0018]
[0019] Among them, gridsample represents the grid sampling, The output of the DySample module. in, for The original sampling grid, represents an offset, which is obtained by the sampling point generator;
[0020] The sampling point generator uses a linear layer with input and output channels of C and 2s respectively. 2 , used to generate a size of 2s 2 The offset of Then use the Pixel Shuffle operation to set the offset Reshape into 2×sH×sW, where:
[0021]
[0022] Wherein, sigmoid represents the sigmoid function, linear1 and linear2 are linear functions respectively.
[0023] The DyHead module treats the input as three dimensions and uses a split attention mechanism to pay attention to each independent dimension separately, including: performing scale-aware attention π on the feature level dimension L , performing spatially aware attention π on the spatial position dimension S , performing task-aware attention π in the channel number dimension C , the mathematical expression of the DyHead module as follows:
[0024]
[0025] in:
[0026] Input feature map L, S, and C are the feature level, the product of width W and height H, and the number of channels of the input feature map, respectively;
[0027]
[0028] In the above formula, f(.) is a linear function approximately simulated by a 1x1 convolutional layer; σ() is the hard-sigmoid function;
[0029]
[0030] In the above formula, K is the number of sparse sampling positions; p k +Δp k is the position Δp obtained by self-learning spatial offset k to focus on the discriminative region, Δm k is the self-learning importance scalar at position pk; c is the channel number; ω l,k is the joint representation of the number of feature levels and the number of sparse sampling positions;
[0031]
[0032] In the above formula, the feature slice of the c-th channel, [α 1 , α 2 , β 1 , β 2 T is a hyperparameter used to learn and control the activation threshold.
[0033] The processing process of the DyHead module includes:
[0034] The feature map is first processed by the scale-aware attention π L : Global average pooling, 1×1 convolution, ReLU activation function, Hard-sigmoid activation function are performed in sequence, and then matrix multiplication is performed with itself, and the output is given to the space-aware attention π S ;
[0035] The processing process of the space-aware attention π S includes: performing exponential offset learning and 3×3 convolution, a part of the output passes through the sigmoid activation function, another part is compensated, and then dot multiplication is performed with the input, and the output is given to the task-aware attention π C ;
[0036] The processing process of the task-aware attention π C includes: average pooling, fully connected, ReLU activation function, fully connected, normalization operation, then element-wise addition with (1, 0, 0, 0), and then through the shifted sigmoid function.
[0037] Before the Backbone part of the improved deep learning algorithm model, an image preprocessing module is provided, which performs online image augmentation on the input image using one of Mosic, Mixup, and MedianBlu, and uniformly scales the image to the first size;
[0038] Before the first layer of the CGDown module in the Backbone part, a CBS module as the 0th layer is provided, which is a 3×3 convolutional module with a stride of 2 and is used to downsample the image.
[0039] The three DyHead detection modules respectively take the outputs of the 16th, 19th, and 22nd layers of the YOLO11n model as inputs. Among them, the size of the feature map output by the 16th layer is 80×80×256, the size of the feature map output by the 19th layer is 40×40×512, and the size of the feature map output by the 22nd layer is 20×20×1024.
[0040] The improved deep learning algorithm model uses the SGD optimizer for gradient convergence, and the formula is as follows:
[0041]
[0042] θ t represents the model parameters at time step t, including weights and biases; η t represents the learning rate, which is used to control the step size of parameter update; represents the gradient of the loss function J with respect to the parameter θ t at the sample x (i) ,y (i) ; θ t+1 is the updated model parameter.
[0043] Collect various part images in the factory environment, including images of the target object in occluded and various lighting environments;
[0044] Construct an image dataset including different sample groups, including: performing image screening, classification annotation, then normalizing the center position, width, and height of the target object in the image, and then performing image data augmentation;
[0045] In the occluded samples, the target object is in a state of being partially occluded, fully occluded, or stacked;
[0046] In the reflective samples, the target object has a highly reflective surface of metal or plastic;
[0047] In the multi-specification mixed samples, the size of the target object is 1cm 2 to 10cm 2 .
[0048] The beneficial effects of the present invention are as follows:
[0049] The present invention uses the single-stage network YOLO11 in the deep learning object detection network as the basic model, and simultaneously improves the Backbone, Neck, and Head parts of YOLO11. The dataset consists of various types of parts produced in the factory. By screening and data augmentation of the dataset, the final dataset contains training sets under conditions such as occlusion between parts, reflection of parts, and large differences in part specifications. The improved model is trained through the training set, and the obtained recognition model improves the recognition and detection accuracy of part types.
[0050] The recognition model of the present invention can fuse local features and global context information through the CGDown module with an attention mechanism, enhancing the feature expression ability in complex environments. The feature reduction process is optimized through DySample dynamic upsampling to reduce information loss. Through the DyHead detection module, triple attention mechanisms of scale, space, and task are superimposed to significantly improve the multi-object detection accuracy. Experiments show that the average precision index of the model of the present invention on the self-built factory dataset is increased by 2.7% compared with the original YOLO11n, and the model has prominent lightweight advantages, and the increase in computational complexity is controlled within the applicable range of industrial real-time detection, significantly improving the part detection performance in complex industrial scenarios.
[0051] The present invention realizes an innovative balance between accuracy and efficiency and has strong industrial adaptability. Through modular design, the deployment advantages of the YOLO framework are retained, and the model robustness is strengthened under complex working conditions such as part stacking and uneven illumination. Its low computational overhead characteristics support the rapid migration and application of production line equipment. Compared with traditional models, the recognition model of the present invention breaks through the accuracy bottleneck of multi-part detection at a controllable resource cost, provides a visual detection solution that takes into account both real-time performance and accuracy for intelligent manufacturing scenarios, and has significant engineering application value.
[0052] Other features and advantages of the present invention will be described in the following specification or understood by implementing the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic structural diagram of the improved deep learning algorithm model in the embodiment of the present invention.
[0054] Figure 2 It is a stacked part diagram in a complex environment collected in the embodiment of the present invention.
[0055] Figure 3 It is a schematic network structure diagram of the CGDown module in the embodiment of the present invention.
[0056] Figure 4 This is a schematic diagram of the network structure of the DySample module in an embodiment of the present invention.
[0057] Figure 5 This is a schematic diagram of the network structure of the DyHead detection module in an embodiment of the present invention. Detailed implementation manners
[0058] The following describes the detailed implementation manners of the present invention with reference to the accompanying drawings.
[0059] A method for part recognition in a factory environment based on deep learning in this embodiment includes:
[0060] S1. Construct an image dataset: Collect various part images in the factory environment, perform image screening, classification annotation, then normalize the central position, width, and height of the target object in the image, and then perform image data augmentation to obtain the image dataset.
[0061] Among them, in the deep learning part recognition task in the factory environment, the construction of the dataset is closely combined with the actual working conditions, and complex interference scenarios are mainly covered. As a preferred method, the original data used in this embodiment is various types of stacked parts collected on a simulated industrial assembly line, including situations such as occlusion and complex lighting conditions, as Figure 2 shown. When collecting samples, the lighting diversity (stable light source, dynamic light change, low light environment) of the industrial environment is simulated, and situations such as occlusion are introduced. The captured images are screened, and the parts in these images are divided into 10 categories for annotation. The data format after annotation is: class id ; x center ; y cebter ; width; height. Among them, class id is the target category number, starting from 0; x center , y center is the central position of the target object in the image, normalized to the range [0, 1]; width and height are the width and height of the target object, normalized to the range [0, 1]. After annotation, the dataset is subjected to image data augmentation by translation, brightness jitter, and adding Gaussian noise, and extreme specular reflection and shadow effects are synthesized to enhance the anti-interference ability and improve the model generalization. Finally, occlusion samples (including target objects in a partially occluded, fully occluded, or stacked state), specular reflection samples (including target objects with highly specular surfaces of metal or plastic), and multi-specification mixed samples (including target objects of different sizes from 1 cm 2 - 10 cm 2 ) are retained. Finally, 3040 images of 10 types of parts are obtained, realizing the construction of the image dataset and ensuring the realistic representation ability of the dataset.
[0062] As a preferred embodiment, the image dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1.
[0063] As a specific embodiment, the image data augmentation further includes operations such as adding salt-and-pepper noise, Gaussian noise, translation, adjusting contrast and saturation, and cropping to the collected part images.
[0064] S2. Refer to Figure 1 , and construct an improved deep learning algorithm model based on the YOLO11n model:
[0065] S21. In the Backbone part, replace the ordinary convolution modules in the 1st, 3rd, 5th, and 7th layers of the YOLO11n model with CGDown modules.
[0066] Among them, the CGDown module combines the CGNet idea and is a downsampling module with an attention mechanism, which effectively improves the network's ability to understand features of various scales in complex scenes, enables the network to perform semantic segmentation more accurately, and improves the accuracy of feature extraction.
[0067] As Figure 3 shown, as a specific embodiment, the CGDown module convolves the feature map through a pre-convolution module with a convolution kernel of 3×3 and a stride of 2 to adjust the number of channels, and then divides it into two parts. The two parts are respectively subjected to dilated convolution and ordinary convolution feature extraction and then spliced; the spliced feature map passes through a normalization layer and a SiLu function, and then is separated into two parts through a post-convolution module with a convolution kernel of 1×1 and a stride of 1. One of the parts is processed through average pooling and two fully connected layers, added to the other part and output.
[0068] As a specific embodiment, the improved deep learning algorithm model is provided with an image preprocessing module before the Backbone part, which uses one of Mosic, Mixup, and MedianBlu to perform online image augmentation on the input image and uniformly scales the image to a size of 640×640.
[0069] As a specific embodiment, in the Backbone part, before the 1st layer CGDown module, there is a CBS module as the 0th layer, which is a 3×3 convolution module with a stride of 2 and is used to downsample the image. Subsequently, the input image is subjected to feature extraction through the CGDown module, and the output image size changes to 320×320×128, and then passes through the C3k2 module of the original network.
[0070] As a specific implementation manner, each layer of the CGDown module and the subsequent connected C3k2 module form a stage. After each stage, the number of channels of the model doubles and the feature map is reduced to half of the original. Among them, an SPPF module and a C2PSA module are set at the end of the Backbone part. After passing through the Backbone part, the image size becomes 20×20×1024, and then it is output to the Neck part.
[0071] S22. In the Neck part, the ordinary upsampling modules of the 11th and 14th layers of the YOLO11n model are replaced with DySample modules. The output of the 11th layer DySample module after replacement is concatenated with the output of the 6th layer C3k2 module and then input into the 12th layer C3k2 module. The output of the 14th layer DySample module after replacement is concatenated with the output of the 4th layer C3k2 module and then input into the 16th layer C3k2 module. The ordinary convolution modules of the 17th and 20th layers are replaced with the CGDown module. The output of the 17th layer CGDown module after replacement is concatenated with the output of the 13th layer C3k2 module and then input into the 19th layer C3k2 module. The output of the 20th layer CGDown module after replacement is concatenated with the output of the 10th layer C2PSA module and then input into the 22nd layer C3k2 module.
[0072] Among them, the network layers in the Backbone part are fused to enable it to have the ability to extract deep semantic information and shallow texture information accurate to pixel points at the same time.
[0073] Among them, in the DySample module, a part of the feature image passes through the sampling point generator, is input into the sampling set for processing, and then sent to grid sampling, while another part of the feature map is directly sent to grid sampling. Finally, a feature map with the size of sW×sH×C is obtained. As Figure 4 shown, as a specific implementation manner, the DySample module includes a sampling point generator, a sampling set and grid sampling. For a given upsampling scale factor s and a feature map with the size of C×H×W the processing process of the DySample module is as follows:
[0074]
[0075] Among them, gridsample represents the grid sampling, is the output of the DySample module, and the sampling set Among them, is the original sampling grid of Represents the offset, which is obtained by the sampling point generator;
[0076] The sampling point generator uses a linear layer with the number of input and output channels being C and 2s respectively 2 , for generating the offset of size 2s 2 ×H×W Then, the offset is reshaped to 2×sH×sW through the Pixel Shuffle operation, where:
[0077]
[0078] In the above formula, sigmoid is the sigmoid function, and linear1 and linear2 are linear functions respectively.
[0079] Based on the above formula, it can be achieved that: due to the existence of the normalization layer, the value of a certain output feature is usually in the range of [-1, 1] and centered on 0. Therefore, the moving ranges of local s 2 sampling positions may significantly overlap, and such errors will propagate stage by stage and cause output artifacts. To alleviate this problem, the offset is multiplied by a static range factor of 0.25, so that the moving range of the sampling positions is locally constrained; and to increase the flexibility of the offset, a pointwise dynamic range factor is further generated by linearly projecting the input feature. By using the sigmoid function and a static factor of 0.5, the value of the dynamic range is in the range of [0, 0.5] and centered on 0.25, the same as the static range.
[0080] The DySample module avoids the complex calculations brought by dynamic convolution by learning the positions of sampling points and combining dynamic upsampling. It adopts a simple and efficient method to generate content-aware upsampling results without additional high-resolution feature inputs. This enables the DySample module to reduce the model complexity and computational cost while maintaining high performance.
[0081] S23. In the Head part, three DyHead detection modules are introduced to replace the original ordinary Head detection heads respectively, so as to perform object detection on the feature maps of different sizes output by different layers of the Neck part respectively.
[0082] As a specific implementation manner, the three DyHead detection modules respectively take the outputs of the 16th, 19th, and 22nd layers of the YOLO11n model as inputs, so as to perform object detection on the feature map with a size of 80×80×256 output by the 16th layer, that is, to achieve small object detection, perform object detection on the feature map with a size of 40×40×512 output by the 19th layer, that is, to achieve medium object detection, and perform object detection on the feature map with a size of 20×20×1024 output by the 22nd layer, that is, to achieve large object detection.
[0083] As Figure 5 shown, as a preferred manner, the processing process of the DyHead module includes:
[0084] The feature map first undergoes the scale-aware attention π L processing: global average pooling, 1×1 convolution, ReLU activation function, Hard-sigmoid activation function are performed in sequence, and then matrix multiplication is performed with itself, and the output is given to the spatial-aware attention π S ;
[0085] The processing process of the spatial-aware attention π S includes: exponential offset learning and 3×3 convolution are performed, a part of the output passes through the sigmoid activation function, another part is compensated, and then dot multiplication is performed with the input, and the output is given to the task-aware attention π C ;
[0086] The processing process of the task-aware attention π C includes: average pooling, fully connected, ReLU activation function, fully connected, normalization operation are performed, and then element-wise addition is performed with (1,0,0,0), and then the shifted sigmoid function is passed through.
[0087] As can be seen from the above, the DyHead detection module of this embodiment adds an attention mechanism, which is reflected in three aspects: 1. Scale perception is performed between feature levels; 2. Spatial perception is performed between spatial positions; 3. Task perception is performed between output channels. The DyHead module regards the input as three dimensions (level×space×channel), and adopts a separate attention mechanism to perform attention on each independent dimension respectively, including: performing scale-aware attention π L on the feature level dimension (level-wise), performing spatial-aware attention π S on the spatial position dimension (spatial-wise), performing task-aware attention π C on the channel number dimension (channel-wise), and the mathematical expression of the DyHead module is as follows:
[0088]
[0089] Wherein:
[0090] Input feature map L, S, and C are the feature level, the product of width W and height H, and the number of channels of the input feature map, respectively;
[0091]
[0092] In the above formula, f(.) is a linear function approximately simulated by a 1x1 convolutional layer; σ() is a hard-sigmoid function;
[0093]
[0094] In the above formula, K is the number of sparse sampling positions; p k +Δp k is the position Δp obtained by self-learning spatial offset k to focus on the discriminative region, Δm k is the self-learning importance scalar at position pk; c is the channel number; ω l,k is the joint representation of the number of feature levels and the number of sparse sampling positions;
[0095]
[0096] In the above formula, the feature slice of the c-th channel, [α 1 , α 2 , β 1 , β 2 T is a hyperparameter used to learn and control the activation threshold.
[0097] Among them, the scale-aware attention π L is executed on the level dimension. Feature maps of different levels correspond to different object scales. Adding attention at the level can enhance the scale perception ability of object detection. The spatial-aware attention π S is executed on the spatial dimension. Different spatial positions correspond to the geometric transformation of the object. Adding attention on the spatial dimension can enhance the spatial position perception ability of the object detector. The task-aware attention π C is executed on the channel dimension. Different channels correspond to different tasks. Adding attention on the channel dimension can enhance the perception ability of object detection for different tasks.
[0098] S3. Use the image dataset to train, validate, and test the improved deep learning algorithm model to obtain an identification model.
[0099] Specifically, input the training set into the improved deep learning algorithm model for training. After the images are imported, they will be automatically scaled to a size of (640×640), pass through the convolutional network layer by layer, use the SGD optimizer for gradient convergence, and finally obtain a trained model.
[0100] Among them, the calculation formula for gradient convergence using the SGD optimizer is as follows:
[0101]
[0102] θ t represents the model parameters at time step t, including weights and biases; η t represents the learning rate, which is used to control the step size of parameter update; represents the gradient of the loss function J with respect to the parameter θ t at the sample x (i) ,y (i) ; θ t+1 is the updated model parameter.
[0103] Use the trained model file as the weight file, and use the validation set to independently verify the effect and improve it to obtain the final identification model. Import the test set into the identification model and test the overall performance of the model. The final results show that the average recognition accuracy of the model in the occlusion scenario is increased by 2.7%. The model remains robust in the real factory environment, especially the recall rate in the dynamic scenario of the production line is stable above 97.5%, providing a reliable verification benchmark for industrial implementation.
[0104] As a preferred method, the environmental parameters for constructing the identification model in this embodiment are shown in Table 1 below:
[0105] Table 1 Environmental parameters for constructing the identification model
[0106] device version GPU NVIDIA GeForce RTX 4080 Python 3.10.14 Pytorch 2.2.2 CUDA 12.6
[0107] As a preferred method, the conditional parameters during the model training process in this embodiment are shown in Table 2:
[0108] Table 2 Conditional parameters during model training
[0109]
[0110]
[0111] In this embodiment, two metrics, mean Average Precision (mAP) and the number of parameters of the model (Parameters), are used to evaluate the performance of the recognition model. When calculating mAP, two additional metrics can be derived: Precision and Recall. Precision is used to measure the false detection rate of the model, while Recall is used to measure the missed detection rate of the model. The mathematical formulas for Precision P and Recall R are as follows:
[0112]
[0113] Among them, TP represents true positives; FN represents false negatives; FP represents false positives.
[0114] The mathematical formula for Average Precision is:
[0115]
[0116] AP represents the average accuracy of a single class, which is the area under the Precision-Recall curve. And mAP represents the average accuracy of all classes. The larger the value of mAP, the better the performance of the model in terms of accuracy. The formula for mAP is as follows:
[0117]
[0118] When evaluating the model performance, IOU represents the degree of overlap between the predicted bounding box and the ground truth bounding box. Different IOU thresholds are usually used to determine whether the prediction result is correct. When IOU is greater than or equal to the set threshold, the prediction is considered correct. When IOU is less than the set threshold, the prediction is considered incorrect. This embodiment uses mAP@0.5-0.95 (IOU threshold from 0.5 to 0.95) as the main metric to evaluate the performance of the model in terms of accuracy.
[0119] To verify the superiority of the recognition model constructed in this embodiment, an ablation experiment was conducted on the recognition model constructed in this embodiment, and the results are shown in Table 3:
[0120] Table 3 Ablation experiment results under different conditions
[0121]
[0122] In Table 3, a tick "√" represents setting the corresponding module, and no tick means no corresponding improvement. As shown in Table 3, Experimental Group 8 is the model constructed in this embodiment, and seven comparative experimental groups, namely Group 1 to Group 7, are set. The comparative experimental groups include the original yolo11n model, and one or two parts of the Backbone, Neck, and Head of the original yolo11n model are improved, and the improvement method of each part is the same as that of the corresponding part in this embodiment. It can be seen from Table 3 that the model formed by fusing each improved module, that is, the model of this embodiment, has better effects than the models with only a single module or two modules improved separately, and the recognition effect is better.
[0123] Meanwhile, the recognition model of this embodiment is compared with other YOLO models, and the results are shown in Table 4:
[0124] Table 4 Performance comparison between the recognition model of this embodiment and other YOLO models
[0125]
[0126] It can be seen from Table 4 that compared with the YOLO models of other versions, the model of this embodiment has a greater improvement in effects. With a small increase in the number of parameters, the recognition accuracy is relatively high and the recall rate has an improvement.
[0127] S4. Use the recognition model of this embodiment to recognize the target object in the image to be measured.
[0128] In this embodiment, synchronous improvements are made to the three parts of YOLO11, namely the Backbone, Neck, and Head. In the Backbone part of the network, the CGDown module is introduced, which can extract different features of parts from the feature network. When large-volume parts block small parts on the recognition production line, the CGDown module can better extract the exposed feature parts of small parts, such as fine features like the thread features of bolts, the distance features of buckles, and the thickness features of gaskets. Regarding the reflection problem of parts, the reflection features of parts of different sizes and different categories are also very different, and the CGDown module can effectively extract these feature networks. The DySample module introduced in the Neck part of the network can effectively fuse the features extracted by the Backbone part. For example, it can distinguish or fuse the thread features of the occluded bolt and the feature of the front-end occluding object, making the network itself pay more attention to the features of the parts and effectively distinguish them from the occluding objects. Since the main part of the network pays more attention to small features, the data volume may be large. The DyHead detection module introduced in the Head part of the network avoids complexity and high cost. The DyHead detection module also has an attention mechanism and can effectively identify the complex features fused in the Neck part from multiple dimensions. It is found that the accuracy of the improved network is improved compared with the original model, and mAP@0.5:0.95 is increased from 87.2% to 89.9%, an increase of 2.7%. Therefore, the method of this embodiment significantly improves the recognition effect of parts in cases of occlusion and reflection, etc.
[0129] Those of ordinary skill in the art can understand that the above description is only a preferred embodiment of the present invention and is not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for part recognition in a factory environment based on deep learning, characterized in that, Including: Collecting images of multiple parts in the factory environment to construct an image dataset including different sample groups, which includes occluded samples, reflective samples, and multi-specification mixed samples; Constructing an improved deep learning algorithm model based on the YOLO11n model: In the Backbone part, replace the ordinary convolution modules in the 1st, 3rd, 5th, and 7th layers of the YOLO11n model with CGDown modules; In the Neck part, replace the ordinary upsampling modules in the 11th and 14th layers of the YOLO11n model with DySample modules, and concatenate the output of the 11th layer DySample module after replacement with the 6th layer C3k2 module and input it into the 12th layer C3k2 module, and concatenate the output of the 14th layer DySample module after replacement with the 4th layer C3k2 module and input it into the 16th layer C3k2 module; replace the ordinary convolution modules in the 17th and 20th layers with the CGDown modules, and concatenate the output of the 17th layer CGDown module after replacement with the 13th layer C3k2 module and input it into the 19th layer C3k2 module, and concatenate the output of the 20th layer CGDown module after replacement with the 10th layer C2PSA module and input it into the 22nd layer C3k2 module; In the Head part, introduce three DyHead detection modules to replace the original ordinary Head detection heads respectively to perform object detection on the feature maps of different sizes output by different layers in the Neck part; Training and validating the improved deep learning algorithm model using the image dataset to obtain an identification model; Using the identification model to identify the target object in the image to be measured.
2. The method according to claim 1, wherein The CGDown module divides the feature map into two parts after adjusting the number of channels by convolving the feature map through a pre-convolution module with a convolution kernel of 3×3 and a stride of 2. The two parts are respectively subjected to dilated convolution and ordinary convolution feature extraction and then spliced; the spliced feature map passes through a normalization layer and the SiLu function, and then is separated into two parts through a post-convolution module with a convolution kernel of 1×1 and a stride of 1. One part is processed by average pooling and two fully connected layers, added to the other part and output.
3. The method according to claim 2, wherein For the pre-convolution module, in the Backbone part, each layer of the CGDown module and the C3k2 module connected thereto form a stage. After each stage, the number of channels of the model doubles and the feature map is reduced to half of the original.
4. The method according to claim 1, wherein The DySample module includes a sampling point generator and a sampling set and grid sampling. For a given upsampling scale factor s and a feature map of size C×H×W The processing procedure of the DySample module is as follows Among them, gridsample represents the grid sampling, is the output of the DySample module, and the sampling set Among them, is the original sampling grid of represents the offset, which is obtained by the sampling point generator; The sampling point generator uses a linear layer with the number of input and output channels being C and 2s respectively 2 , for generating the offset with a size of 2s 2 ×H×W Then, the offset is reshaped into 2×sH×sW through the Pixel Shuffle operation, where: In the formula, sigmoid represents the sigmoid function, and linear1 and linear2 are linear functions respectively.
5. The method according to claim 1, characterized in that The DyHead module treats the input as three dimensions and adopts a separate attention mechanism to perform attention on each independent dimension, including: performing scale-aware attention π on the feature level dimension L , performing spatial-aware attention π on the spatial position dimension S , performing task-aware attention π on the number of channels dimension C , the mathematical expression of the DyHead module is as follows: Wherein: Input feature map L, S, and C are the feature level, the product of width W and height H, and the number of channels of the input feature map, respectively; In the above formula, f(.) is an approximation of the linear function by a 1x1 convolutional layer; σ() is the hard-sigmoid function; In the above formula, K is the number of sparse sampling positions; p k +Δp k is the position Δp offset by self-learning in the space k to focus on the distinguishable region, Δm k is the self-learned importance scalar at the position pk; c is the serial number of the channel; ω l,k is the joint representation of the number of feature levels and the number of sparse sampling positions; In the above formula, The feature slice of the c-th channel, [α 1 , α 2 , β 1 , β 2 T are hyperparameters used to learn and control the activation threshold. 6. The method according to claim 5, wherein The processing process of the DyHead module includes: Feature map First, the scale-aware attention π L Processing: First, perform global average pooling, 1×1 convolution, ReLU activation function, and Hard-sigmoid activation function in sequence, then perform matrix multiplication with itself, and output to the spatial-aware attention π S ; The spatial perception attention π S The processing process includes: performing exponential offset learning and 3×3 convolution. A part of the output passes through the sigmoid activation function, another part is compensated, and then a dot product operation is performed with the input, and the output is given to the task perception attention π C ; The task-aware attention π C The processing process includes: average pooling, fully connected, ReLU activation function, fully connected, after normalization operation, element-wise addition with (1, 0, 0, 0), and then passing through the shifted sigmoid function.
7. The method according to claim 1, characterized in that, Before the Backbone part of the improved deep learning algorithm model, an image preprocessing module is provided, which uses one of Mosic, Mixup, and MedianBlu to perform online image augmentation on the input image and uniformly scales the image to the first size; In the Backbone part, before the first layer of the CGDown module, there is a CBS module as the 0th layer, which is a 3×3 convolutional module with a stride of 2 and is used to downsample the image.
8. The method according to claim 1, characterized in that The three DyHead detection modules respectively take the outputs of the 16th, 19th, and 22nd layers of the YOLO11n model as inputs. Among them, the feature map size output by the 16th layer is 80×80×256, the feature map size output by the 19th layer is 40×40×512, and the feature map size output by the 22nd layer is 20×20×1024.
9. The method according to claim 1, characterized in that, The improved deep learning algorithm model uses the SGD optimizer for gradient convergence, and the formula is as follows: θ t denotes the model parameters at time step t, including weights and biases; η t Represents the learning rate, which is used to control the step size of parameter update; Represents the gradient of the loss function J with respect to the parameter θ t at the sample x (i) , y (i) ; θ t+1 are the updated model parameters.
10. The method according to claim 1, characterized in that Collect various part images in the factory environment, including images of the target object under occlusion and various lighting conditions; Construct an image data set including different sample groups, including: performing image screening, classification annotation, then normalizing the center position, width, and height of the target object in the image, and then performing image data augmentation; In the occlusion samples, the target object is in a state of being partially occluded, fully occluded, or stacked; In the specular reflection samples, the target object has a highly specular surface of metal or plastic; The size of the target object in the multi-specification mixed sample is 1 cm 2 to 10 cm 2 .
Citation Information
Cited By
Light and simplified monitoring equipment, method and application for rice crab living environment based on environment sensor and image acquisition device
CN121547556A
Rapid sheet metal part identification method based on binary template matching
CN121724927A