Part Recognition Method Based on Lightweight Improved YOLOv5

By replacing Backbone with PPLC-Net in YOLOv5, and introducing TransformerHead, C3TR module and noise purification module, the problems of large memory usage and high equipment requirements in part recognition are solved, and efficient and accurate part recognition is achieved.

CN116503379BActive Publication Date: 2025-06-17CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310549559.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-16
Publication Date
2025-06-17
Estimated Expiration
2043-05-16

AI Technical Summary

Technical Problem

The existing YOLOv5 algorithm has problems such as redundant network structure, large model parameters, slow training speed and high equipment requirements in part recognition, resulting in large memory usage and high equipment costs.

Method used

By replacing Backbone in YOLOv5 network as PPLC-Net and introducing TransformerHead as the decoupling head, the C3TR module and the noise purification module are added to reduce computing and memory costs and improve the detection accuracy and robustness of the model.

Benefits of technology

A lightweight part recognition model is realized, which reduces memory usage and equipment requirements, improves the accuracy and efficiency of identification and detection, making the model more suitable for deployment on factory robotic arms and other equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116503379B_ABST
    Figure CN116503379B_ABST
Patent Text Reader

Abstract

The present invention relates to a part recognition method based on lightweight improved YOLOv5, belonging to the field of computer vision technology. The method specifically includes: S1: Collect part photos and perform data cleaning and type annotation; S21: Replace the Backbone in the original YOLOv5 network with PPLC-Net, and select the H-Swish activation function; S22: Construct a TransformerHead as the decoupled head of YOLOv5; S23: Add Transformer modules and C3 modules before the detection head and fuse them to form a C3TR module; S24: Introduce a noise purification module; S3: Train the improved YOLOv5 network and use the trained network to identify and detect actual parts to be measured. The present invention can achieve accurate and rapid part recognition and detection, and at the same time improve the problems of large memory occupation and high equipment requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and relates to a part recognition method based on a lightweight improved YOLOv5. Background Art

[0002] With the rapid development of information technology and the manufacturing industry, the rising labor costs have had a huge impact on some labor-intensive manufacturing industries. In the traditional product assembly process, the recognition of parts is generally completed by the naked eye, and the experience and conditions of workers have a great influence on the recognition accuracy of parts. This leads to an increase in the consumption of human resources, and there are situations such as missed inspections and mixed inspections, which affect work efficiency.

[0003] In recent years, with the continuous improvement of technologies such as artificial intelligence and computer vision, intelligent assembly systems use high-precision industrial cameras for image acquisition, automatically recognize the feature information of parts, and perform image processing and analysis to ensure fast and uninterrupted work. It not only helps to improve assembly efficiency but also saves valuable human resources. Currently, the object detection algorithms of convolutional neural networks are mainly divided into two categories: two-stage object detection algorithms and one-stage object detection algorithms. The two-stage object detection algorithm divides the detection task scope into two parts: first, it is necessary to generate candidate regions, which contain the approximate position information of the target; then, the candidate regions are classified to obtain more accurate position information. Representative two-stage object detection algorithms include the RCNN model and many of its improved models. The one-stage object detection algorithm does not need to generate candidate regions and can directly obtain the classification and position information of the target of interest. Typical one-stage object detection algorithms are YOLO and SSD, and these two types of algorithms have their own advantages and disadvantages.

[0004] As a representative of the one-stage object detection algorithm, YOLOv5 has high real-time performance, and its accuracy can also be comparable to that of the two-stage object detection algorithm due to its rich preprocessing methods and effective network architecture. However, its network structure does not have resource constraint conditions, resulting in problems such as redundancy of features and structure in its network. At the same time, the part model has a large number of parameters, slow training speed, and high requirements for equipment, which greatly increases the enterprise cost and also raises the difficulty for enterprises to promote intelligent manufacturing.

[0005] Therefore, there is an urgent need for a lightweight part recognition method to improve the problems of large memory occupancy and high equipment requirements. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a part recognition method based on lightweight improved YOLOv5, which can achieve accurate and rapid part recognition and detection while improving the problems of large memory occupancy and high equipment requirements, making the model more conducive to being carried on devices such as factory robotic arms. It reduces the equipment requirements for intelligent assembly and is more conducive to improving the utilization rate of labor in the assembly process.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A part recognition method based on lightweight improved YOLOv5 specifically includes the following steps:

[0009] S1: Collect part photos, then perform data cleaning, and label different types of pictures using labelImg; select the YOLOv5 network structure for initial model training;

[0010] S2: Improve the YOLOv5 network structure, specifically including:

[0011] S21: Replace the Backbone in the original YOLOv5 network with PPLC-Net, and select the H-Swish activation function;

[0012] S22: Select Transformer to construct TransformerHead as the decoupled head of YOLOv5;

[0013] S23: Add C3TR modules composed of Transformer modules and C3 modules in front of the detection head;

[0014] S24: Introduce a noise purification module;

[0015] S3: Train the improved YOLOv5 network, and use the trained network to recognize and detect actual parts to be measured.

[0016] Furthermore, in step S21, the PPLC-Net network uses deep containers as the basic blocks of the framework, and the components are depthwise separable convolutions. The first layer is convolution + standard normalization + H-Swish function (i.e., conv + bn + hardswish), including 13 layers of depthwise separable convolutions (i.e., DepthSepConv) in the middle. The subsequent GAP is a 7*7 global average pooling (GlobalAverage Pooling). After GAP, there is a point convolution + fully connected layer + H-Swish function (i.e., point conv + FC + hardswish) component, and finally a fully connected layer with an output of 1000 (i.e., FC layer).

[0017] The activation function is replaced by the H-Swish function, and its calculation formula is as follows:

[0018]

[0019] Furthermore, step S22 specifically includes the following steps:

[0020] S221: Add prediction heads to handle the large-scale variance of objects, and integrate Transformer Prediction Heads into YOLOv5; the added prediction heads are generated from images with low-level and high-resolution labels, and are more sensitive to tiny objects; only deployed at the end of the backbone; the Transformer formula is as follows:

[0021]

[0022] Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k represents the dimension of matrices Q and K, which is used to control the range of dot product values.

[0023] S222: Adopt the multi-head attention mechanism, which essentially splits the three parameters Q, K, and V multiple times while the total number of parameters remains unchanged. Each group of split parameters is mapped to different subspaces in the high-dimensional space to calculate the attention weights, and after multiple parallel calculations, the attention information in all subspaces is merged; the formula is as follows:

[0024]

[0025] Among them, W O 、W i Q are the parameter matrices during linear transformation, and W i Q 、W i K 、W i V are the attention weights of the different subspaces where the parameters Q, K, and V are mapped to the high-dimensional space.

[0026] S223: The encoder of the Transformer structure is stacked by N layers of networks, and each layer of network contains an encoder and a decoder structure; among them, the encoder contains two sub-layers: the multi-head attention mechanism and the feed-forward neural network, and each layer in the decoder contains a masked multi-head attention mechanism, an encoder-decoder multi-head attention mechanism, and a feed-forward neural network; the input embedding of the encoder is the input matrix I (t) after the trajectory sequence is processed by block segmentation, and the output embedding of the decoder is the output of the trajectory prediction algorithm The p of the true position of the target at the next moment (t) vector difference;

[0027] S224: Before each module, layer normalization is first performed. After the module, an identity mapping residual branch is connected, and Dropout is connected to promote network convergence and prevent overfitting. The Transformer encoder is composed of multiple alternating multi-head self-attention mechanisms (MSA) and multi-layer perceptron blocks (MLP). The feature calculation process of the l-th layer is as follows:

[0028] z l ' = MSA(LN(z l-1 )) + z l-1 , l = 1, …, L

[0029] z l = MLP(LN(z l ')) + z l ', l = 1, …, L

[0030] Furthermore, in step S23, the C3TR module is formed by fusing the BottleNeck (bottleneck layer) in the original C3 module with the Transformer module. Then it is connected before each detection head.

[0031] Furthermore, in step S24, the introduced noise purification module is implemented by combining two branches in parallel; the advantages of these two branches can complement each other, and the parallel combination can further improve the performance, especially for the performance improvement of the object detection task is significant.

[0032] The upper branch can achieve more refined feature selection, as well as more accurate object localization and detection by combining spatial and channel attention. It is mainly used to select the most relevant spatial positions and suppress the responses of irrelevant positions in order to better locate the object;

[0033] The lower branch is a lightweight attention mechanism. By introducing the "squeeze-and-excitation" mechanism, it can quickly extract the most useful features; specifically, it mainly obtains the importance of the global features by statistically analyzing the global features, and then uses this importance information as weights to re-weight the features of each channel in order to better extract the features;

[0034] The spatial information extracted by the upper branch can effectively remove noise and interference in the image, enabling the detector to focus more on the target area, thereby improving the detection accuracy. At the same time, the channel information extracted by the lower branch can help the detector better understand different colors, textures, etc. in the image, so as to more accurately judge the position and attributes of the target. The advantage of combining the two branches in parallel is that it can make full use of the information extracted by each branch, and thus better understand the image. At the same time, through the parallel combination method, the computing resources can be fully utilized to accelerate the running speed of the detector.

[0035] The beneficial effects of the present invention are as follows: The present invention uses the PP-LCNet module to design a lightweight backbone feature extraction network, which improves problems such as large computational complexity and large number of parameters in the original YOLOv5; by introducing the TransformerHead as the decoupled head of YOLOv5 and the C3TR module, expensive computations and memory costs are reduced; finally, a noise purification module is introduced to improve the detection accuracy and robustness of the model and effectively remove noise in the image.

[0036] The improved algorithm of the present invention effectively realizes accurate and rapid part recognition and detection, improves the problems of large memory occupation and high equipment requirements, making the model more conducive to being carried on devices such as factory robotic arms. While reducing the equipment requirements for intelligent assembly, it is more conducive to improving the utilization rate of labor in the assembly process.

[0037] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0039] Figure 1 is the network structure diagram of the improved YOLOv5 of the present invention;

[0040] Figure 2 is the structure diagram of C3TR;

[0041] Figure 3 is the structure diagram of the noise purification module;

[0042] Figure 4 is the structure diagram of the Transformer model;

[0043] Figure 5This is a method for part recognition based on the lightweight improved YOLOv5 network of the present invention. Detailed implementation mode

[0044] The following uses specific specific examples to illustrate the implementation mode of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation modes. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0045] Please refer to Figures 1 to 5 , the present invention provides a part recognition method based on an improved YOLOv5 network, which specifically includes the following steps:

[0046] Step S1: Collect part photos, then perform data cleaning, and perform labelImg annotation on different types of pictures; select the YOLOv5s network structure for initial model training;

[0047] Step S2: Replace the Backbone in the original YOLOv5 network with PPLC-Net, and select the H-Swish activation function.

[0048] In step S2, the specific operation steps for entering the SE module include:

[0049] Step S21: The PPLC-Net network uses a depth container as the basic block of the framework. This module does not have operations similar to shortcuts. This module no longer requires operations such as concat or other element accumulation, which will greatly improve the running efficiency of the module. In addition, each component of the used module has been deeply optimized by the Intel CPU acceleration library, and the inference speed can exceed other lightweight blocks, including shufflenet-block, etc.

[0050] The components are depthwise separable convolutions. The first layer is conv+bn+hardswish, which contains 13 layers of DepthSepConv in the middle. The subsequent GAP is a 7*7 Global Average Pooling. After GAP, there is a point conv+FC+hardswish component, and finally an FC layer with an output of 1000. Replace the Backbone in the YOLOv5 network with it.

[0051] Step S22: All activation functions use H-Swish, which improves performance with almost no additional increase in inference time. In our module, the output dimension of the network after GAP is very small. To improve the fitting ability of the network, we append a 1×1 conv (equivalent to an FC layer) with a dimension of 1280 after the final GAP layer, which allows for the storage of the model with little increase in inference time.

[0052] The calculation formula of the H-Swish activation function is:

[0053]

[0054] Step S3: Select Transformer to construct TransformerHead as the decoupled head of yolov5. This model has the characteristics of fast forward propagation speed, low structural complexity, and high feature extraction efficiency, and has achieved good results in sequence data processing tasks such as natural language. Specifically, it includes:

[0055] Step S31: Add a prediction head to handle the large-scale variance of objects and integrate Transformer PredictionHeads into YOLOv5. The added prediction head is generated from the image in the low-level and high-resolution rate, and is more sensitive to tiny objects. It is only deployed at the end of the backbone. The Transformer formula is as follows:

[0056]

[0057] Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k represents the dimension of matrices Q and K, which is used to control the range of dot product values.

[0058] Step S32: The essence of the multi-head attention mechanism is to split the three parameters Q, K, and V multiple times while the total number of parameters remains unchanged. Each group of split parameters is mapped to different subspaces in the high-dimensional space to calculate the attention weights, and after multiple parallel calculations, the attention information in all subspaces is merged. The formula is as follows:

[0059]

[0060] Among them, W o 、 are the parameter matrices during linear transformation.

[0061] Step S33: The Transformer model adds structures such as summation, normalization, and multi-layer perceptron on the basis of multi-head attention, such as Figure 4As shown. The encoder of this structure is composed of N layers of networks stacked together, and each layer of network contains an encoder and a decoder structure. Among them, the encoder contains two sub-layers: a multi-head attention mechanism and a feed-forward neural network, while each layer in the decoder contains a masked multi-head attention mechanism, an encoder-decoder multi-head attention mechanism, and a feed-forward neural network. The input embedding of the encoder is the input matrix I after the trajectory sequence is processed by block segmentation (t) , and the output embedding of the decoder is the output of the trajectory prediction algorithm and the p of the true position of the target at the next moment (t) vector difference

[0062] Step S34: First, pass through layer normalization before each module, connect an identity mapping residual branch after the module, and connect Dropout to promote network convergence and prevent overfitting. The Transformer encoder is composed of multiple alternating multi-head self-attention mechanisms (MSA) and multi-layer perceptron blocks (MLP). The feature calculation process of the l-th layer is shown in the following formula

[0063] z l ' = MSA(LN(z l-1 )) + z l-1 , l = 1, …, L

[0064] z l = MLP(LN(z l ')) + z l , l = 1, …, L

[0065] Step S4: Add C3TR modules composed of a combination of Transformer modules and C3 modules before all detection heads

[0066] As Figure 2 shown, the C3TR module is formed by fusing the Bottleneck (bottleneck layer) in the original C3 module with the Transformer module. Then it is connected before each detection head. The input of the C3TR module is the feature map obtained after the original image undergoes convolution, downsampling, and multi-scale feature fusion by the network before the module. The feature map then outputs a sequence of image patches through operations such as segmentation, flattening, and linear transformation. Segmenting the feature map not only makes full use of the convolutional network to filter out a large amount of irrelevant information in the original image, but also takes the extracted feature information as input, which can accelerate network convergence and reduce the training burden on high-resolution images

[0067] Step S5: Introduce a noise purification module

[0068] The noise purification module realizes purification through two branches. The advantages of these two branches can complement each other, and the parallel combination can further improve performance, especially for the performance improvement effect of the object detection task is significant

[0069] The upper branch can achieve more refined feature selection, as well as more accurate target localization and detection by combining spatial and channel attention. It is mainly used to select the most relevant spatial positions and suppress the responses of irrelevant positions in order to better localize the target.

[0070] The lower branch is a lightweight attention mechanism that can quickly extract the most useful features by introducing the "squeeze-and-excitation" mechanism. Specifically, it mainly obtains the importance of global features through statistics on global features, and then uses this importance information as weights to re-weight the features of each channel in order to better extract features.

[0071] Combining them in parallel can further improve performance. For the target detection task, parallel combination can improve the detection accuracy and robustness of the model. After combining the upper and lower branches in parallel, it can effectively remove the noise in the image, and the noise purification module is as shown in Figure 3 As shown. Specifically, the upper branch is responsible for extracting the spatial information of the image, and by weighting the channel attention of the image, the spatial information of the image is strengthened. The lower branch is responsible for extracting the channel information of the image, and by weighting the attention of different channels, the channel information of the image is strengthened. By fusing the information of these two branches, more accurate and robust target detection results can be obtained. Specifically, the spatial information extracted by the upper branch can effectively remove the noise and interference in the image, making the detector pay more attention to the target area, thereby improving the detection accuracy. At the same time, the channel information extracted by the lower branch can help the detector better understand different colors, textures, etc. in the image, so as to more accurately judge the position and attributes of the target. The advantage of combining the two branches in parallel is that it can make full use of the information extracted by each branch, and thus better understand the image. At the same time, through the parallel combination method, the computing resources can be fully utilized to accelerate the running speed of the detector. For the target detection task, the advantages of combining the upper and lower branches in parallel are also very obvious. By strengthening the spatial information and channel information, the detector can more accurately judge the position, size and attributes of the target, thereby improving the detection accuracy. In addition, since the parallel combination can effectively remove the noise, the detector is more robust and can handle more complex and variable scenarios.

[0072] Step S6: Deploy the training weight pretrain.pt file of the improved YOLOv5 on the mobile device, and perform actual recognition and detection, as well as test the running effect of using the embedded device, so that the system can be used in other occasions.

[0073] Comparative experiment:

[0074] In this embodiment, a comparison experiment was set up between the improved algorithm and the original YOLOv5. After 300 epochs of training, a series of detection index data for both the training and testing phases were obtained. Analyzing the results from three performance indicators: accuracy, mAP@0.5, and mAP.5:.95, the results are shown in Table 1 below:

[0075] Table 1

[0076] algorithm P mAP@0.5 mAP.5:.95 original algorithm 0.942 0.901 0.623 improved algorithm of the present invention 0.965 0.941 0.754

[0077] As can be seen from Table 1, the accuracy of the improved model of the present invention is 96.5%; mAP@0.5 is 94.1%; Map.5:.95 is 75.4%. Compared with the original YOLOv5 model, the accuracy has increased by 2.3%; mAP@0.5 has increased by 4%; Map.5:.95 has increased by 13.1%. The experiment shows that the improved model has greatly improved the detection accuracy.

[0078] In this embodiment, a lightweight comparison experiment was set up between the improved algorithm of the present invention and YOLOv5s, YOLOV5m, YOLOV5l, and YOLOV5x. Analyzing the results from three performance indicators: FLOPs, Params, and Size, the results are shown in Table 2 below:

[0079] Table 2

[0080]

[0081]

[0082] As can be seen from Table 2, the improved algorithm of the present invention compared with YOLOv5s which has the smallest depth and the smallest feature map width in YOLOv5. The improved algorithm has reduced FLOPs by 7.8G, only 52.72% of YOLOv5s; reduced Params by 2.9M, only 59.72% of YOLOv5s; and reduced Size by 5M, only 63.90% of YOLOv5s.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A part recognition method based on lightweight improved YOLOv5, characterized in that, The method specifically includes the following steps: S1: Collect part photos, then perform data cleaning, and label different types of pictures; Select the YOLOv5 network structure for initial model training; S2: Improve the YOLOv5 network structure, specifically including: S21: Replace the Backbone in the original YOLOv5 network with PPLC-Net, and select the H-Swish activation function; The PPLC-Net network uses depth containers as the basic blocks of the framework, and the components are depthwise separable convolutions. The first layer is convolution + standard normalization + H-Swish function, including 13 layers of depthwise separable convolutions in the middle, and the subsequent GAP is a 7*7 global average pooling. After GAP, there is a point convolution + fully connected layer + H-Swish function component, and finally a fully connected layer with an output of 1000; S22: Select Transformer to construct TransformerHead as the decoupled head of YOLOv5; S23: Add Transformer modules and C3 modules before the detection head and fuse them to form the C3TR module; S24: Introduce a noise purification module, which is implemented by combining two branches in parallel for purification; The upper branch realizes more refined feature selection by combining spatial and channel attention, is used to select the most relevant spatial positions, and suppresses the responses of irrelevant positions; The lower branch is a lightweight attention mechanism, which quickly extracts the most useful features by introducing the "squeeze-and-excitation" mechanism; Specifically, by statistically analyzing the global features, the importance of the global features is obtained, and then this importance information is used as weights to re-weight the features of each channel; S3: Train the improved YOLOv5 network, and use the trained network to identify and detect actual parts to be measured.

2. The part recognition method according to claim 1, characterized in that, Step S22 specifically includes the following steps: S221: Add prediction heads to handle the large-scale variance of objects, and integrate Transformer Prediction Heads into YOLOv5; S222: Employ a multi-head attention mechanism, which essentially splits the Q , K , V three parameters multiple times while keeping the total number of parameters unchanged. Each set of split parameters is mapped to different subspaces in the high-dimensional space to calculate the attention weights, and after multiple parallel calculations, the attention information in all subspaces is merged; S223: The encoder of the Transformer structure is stacked by N layers of networks, and each layer of network contains an encoder and a decoder structure; among them, the encoder contains two sub-layers, namely the multi-head attention mechanism and the feed-forward neural network, while each layer in the decoder contains a masked multi-head attention mechanism, an encoder-decoder multi-head attention mechanism, and a feed-forward neural network; the input embedding of the encoder is the input matrix after the trajectory sequence is processed by block segmentation , and the output embedding of the decoder is the output of the trajectory prediction algorithm and the vector difference of the true position of the target at the next moment; S224: First pass through layer normalization before each module, followed by an identity mapping residual branch after the module, and Dropout is connected to promote network convergence and prevent overfitting; The Transformer encoder is composed of multiple alternating multi-head self-attention mechanisms and multi-layer perceptron blocks.

3. The part recognition method according to claim 1, characterized in that, In step S222, the expression of the multi-head attention mechanism is: Among them, and are parameter matrices during linear transformation, and and are parameters Q and K and V are attention weights that map to different subspaces in the high-dimensional space.