Disordered part identification method based on improved YOLOv8
By improving the YOLOv8 model and specific variation index classification technology, the problems of inefficiency and low accuracy of traditional part recognition methods in complex environments are solved, and efficient and accurate identification of disordered parts are achieved.
Patent Information
- Application Number
- CN202510090947.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-21
AI Technical Summary
Traditional part recognition methods are inefficient and have low accuracy when dealing with complex scenarios and diverse parts. Especially in industrial environments, the accuracy is reduced and consumes a lot of computing power.
The disordered part recognition method based on improved YOLOv8 is adopted. By acquiring the disordered part images and dividing them into sub-image areas, the effective image texture and contour complexity are determined, and the specific variation index is calculated to divide the specific variation category, and then the improved YOLOv8 model or standard part library is used for identification.
It reduces computing power, improves the efficiency and accuracy of part recognition, and is suitable for complex and changeable industrial environments.
Smart Images

Figure CN119992287A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a disordered parts recognition method based on improved YOLOv8. Background Art
[0002] In the automated production process, accurately and quickly identifying and processing disordered parts is one of the core tasks to improve the automation level of the production line. In recent years, deep learning technology, especially convolutional neural network (CNN) has made significant progress in the field of image recognition. YOLO (You Only Look Once) algorithm, as an advanced real-time object detection technology, can predict the object category and position in the image through a single forward propagation, greatly improving the processing speed.
[0003] For example, Chinese patent publication number: CN117237937A, discloses a method for unordered part recognition based on PointNet++ network. The key steps of the method include local feature extraction, obtaining the center point by adopting the farthest point sampling strategy through the set abstraction layer, and selecting K neighborhood points around each center point to form a local area to obtain richer local features. Subsequently, the PointNet layer is used to capture the relationship between points in these local areas, and this process is repeated twice to deepen feature learning. Next, the Maxpooling layer is used to optimize the unordered point cloud data and redundant features to reduce the complexity. Finally, the MLP network is used to classify the processed features to achieve the point cloud classification task. In order to prevent the model from overfitting, the AdamW optimizer and the loss function of label smoothing and cosine decay are used. Overall, this method significantly improves the accuracy and speed of unordered part classification.
[0004] For example, Chinese patent publication number: CN118366006A discloses a part recognition method and system based on an improved Mask R-CNN instance segmentation network, wherein the recognition method includes: constructing a part recognition model based on an improved Mask R-CNN network; the improved Mask R-CNN network includes: using a Mobilenetv3 network to replace the backbone network of the Mask R-CNN network, and using a bidirectional feature fusion network M-FPN network to replace the FPN network connected to the backbone network. The present invention can accurately and quickly recognize parts in disordered and complex environments, and the model used is light and concise, with small computational complexity and low equipment requirements.
[0005] However, there are still the following problems in the prior art:
[0006] In actual situations, traditional part recognition methods rely on simple sensors and predetermined algorithms, which are inefficient and inaccurate when dealing with complex scenarios and diverse parts. At the same time, although the YOLO algorithm has demonstrated superior performance in multiple application scenarios, its accuracy is reduced and it consumes a lot of computing power when dealing with complex and changing industrial environments, especially when identifying disordered parts on the assembly line. Summary of the invention
[0007] To this end, the present invention provides an unordered parts recognition method based on improved YOLOv8, so as to solve the problem that traditional parts recognition methods rely on simple sensors and predetermined algorithms, which are inefficient and inaccurate when processing complex scenes and diversified parts. At the same time, although the YOLO algorithm has shown superior performance in multiple application scenarios, it has a reduced accuracy and consumes a lot of computing power when processing complex and changeable industrial environments, especially when identifying unordered parts on an assembly line.
[0008] To achieve the above object, the present invention provides a method for identifying disordered parts based on improved YOLOv8, which comprises:
[0009] Acquire a plurality of disordered part images, divide the disordered part images into a plurality of sub-image regions based on part contours, and determine effective image textures and contour complexity of the parts in the sub-image regions;
[0010] Determining a specific variation index of a sub-image region based on the number of effective image textures and contour complexity to classify the sub-image region into specific variation categories;
[0011] According to the specific mutation categories corresponding to the sub-image areas, the disordered parts are identified, including:
[0012] homogenizing the sub-image regions, setting labels for the sub-image regions and dividing data sets, and inputting the data sets into a preconfigured improved YOLOv8 model to identify part categories;
[0013] or, determining a part image within the sub-image area, and comparing it with a standard parts library to identify the part category;
[0014] The improvement process of the improved YOLOv8 model includes adding a ConTainer network module to the YOLOv8 model detection head, determining the HorNet network module as the backbone network of the YOLOv8 model, introducing a lightweight network LightConv module, and optimizing the loss function during model training.
[0015] Furthermore, the process of determining the effective image texture value and contour complexity of the part in the sub-image area includes:
[0016] Determine a number of texture pixels on the part within the sub-image area and the corresponding texture width;
[0017] Performing grayscale relationship analysis based on the texture pixels;
[0018] Determine that the texture whose grayscale is less than or equal to the preset grayscale and whose texture width is within the reference width is a valid image texture;
[0019] Determine a plurality of edge contour segments of the part within the sub-image region;
[0020] The average curvature of the plurality of edge contour segments is determined as contour complexity.
[0021] Furthermore, the process of determining the specific mutation index of the sub-image area includes:
[0022] Determine the ratio of the number of reference effective image textures to the number of effective image textures as the first influencing feature;
[0023] Determining a ratio of the baseline contour complexity to the contour complexity as a second influencing feature;
[0024] The weighted sum of the first influencing feature and the second influencing feature is determined as the specific mutation index.
[0025] Furthermore, the sub-image area is divided into specific abnormality categories, wherein:
[0026] If the specific mutation index is greater than or equal to a preset specific mutation index, the sub-image area is classified as a non-specific mutation category;
[0027] If the specific mutation index is less than a preset specific mutation index, the sub-image area is divided into a specific mutation category.
[0028] Furthermore, the disordered parts are identified according to the specific mutation categories corresponding to the sub-image areas, wherein:
[0029] If the sub-image area is of a non-specific mutation category, then determining the part image in the sub-image area and comparing it with a standard parts library to identify the part category;
[0030] If the sub-image region is a specific mutation category, the sub-image region is homogenized, a label is set for the sub-image region, and a data set is divided, and the data set is input into a preconfigured improved YOLOv8 model to identify the part category.
[0031] Furthermore, the homogenizing the sub-image regions includes:
[0032] Several sub-image areas are placed on the preset regional background image respectively.
[0033] Furthermore, the process of adding a ConTainer network module to the YOLOv8 model detection head includes:
[0034] Add the ConTainer network module to the detection head of the YOLOv8 model;
[0035] Integrate feature sets of different network layers based on the ConTainer network module and transfer the feature sets to the detection head;
[0036] The feature set includes convolutional layers for object classification and bounding box regression.
[0037] Furthermore, the process of introducing the lightweight network LightConv module includes:
[0038] Determine that the LightConv module replaces the original C2f module of the YOLOv8 neck network;
[0039] Performing a preliminary linear transformation on the sub-image area based on the LightConv module;
[0040] Passing the sub-image region to a lightweight convolutional layer;
[0041] Combining the linear layer and the gating signal, the sub-image region after the light convolution is nonlinearized;
[0042] Transfer the nonlinearized sub-image regions to other linear layers.
[0043] Furthermore, the process of optimizing the loss function during model training includes:
[0044] Replace the CIoU function with the EIoU function;
[0045] Based on the EIoU function, the Focal-Loss function is introduced;
[0046] Among them, the EIoU function calculates the difference values of width and height based on the CIoU function to replace the aspect ratio of the CIoU function.
[0047] Furthermore, the process of performing similarity comparison with the standard parts library to identify the part category includes:
[0048] Determining the similarity between the part image in the sub-image area and the part image in the standard part library;
[0049] Determine that the part category in the standard parts library corresponding to the maximum similarity is the part category in the sub-image area;
[0050] Among them, a number of part images are pre-stored in the standard parts library.
[0051] Compared with the prior art, the present invention obtains a number of disordered part images, divides the disordered part images into a number of sub-image areas based on the part contour, determines the effective image texture and contour complexity of the part in the sub-image area; determines the specific variation index of the sub-image area based on the number of effective image textures and contour complexity to divide the specific variation category of the sub-image area; identifies the disordered parts according to the specific variation category corresponding to each of the sub-image areas, including: homogenizing the sub-image areas, setting labels for the sub-image areas and dividing data sets, and inputting the data sets into a preconfigured improved YOLOv8 model to identify the part category; or, determines the part image in the sub-image area, and compares it with a standard part library to identify the part category, thereby reducing computing power and improving the efficiency of part recognition.
[0052] In particular, the present invention determines the specific variation index of the sub-image area based on the effective image texture quantity and contour complexity, and provides data for dividing the specific variation category of the sub-image area. In actual situations, when sorting the factory assembly line, most of the disordered parts are directly identified. Since the disordered parts contain various parts, their characteristics are different. Due to the difference in characteristics, the difficulty of back-end identification will be affected. For example, some parts with complex structures and complex textures are not easy to identify, and some parts are easy to identify. Therefore, if the same method is used to identify some ordered parts, it will lead to a waste of computing power and low efficiency of part recognition. Based on this, the present invention considers calculating the specific variation index of the sub-image area to divide the specific variation category. For the sub-image area of the non-specific variation category, the part category is directly identified according to the standard part library. For the sub-image area of the specific variation category, the YOLOv8 model is improved to identify the part category. In particular, the process of identifying the effective image texture quantity and contour complexity is relatively simple, which can reduce the computing power and improve the accuracy, detection speed and precision of the disordered part recognition.
[0053] In particular, the present invention adds a ConTainer network module to the YOLOv8 model detection head. The ConTainer network module combines the methods of GNN, Transformer and MLP-Mixers, utilizes remote interaction and local convolution operations, effectively integrates and understands images, and improves the adaptability and recognition accuracy of the model for complex scenes without changing the size of the feature vector. Compared with the existing model, ConTainer achieves higher accuracy with fewer parameters and faster convergence speed, realizes more accurate prediction of the existence and position of the target, and improves the accuracy, detection speed and precision of disordered parts recognition.
[0054] In particular, the present invention determines that the HorNet network module is the backbone network of the YOLOv8 model. In the existing models, the backbone networks are mostly C2f network modules. However, in actual operation, the C2f network module will cause the loss of some important features when processing data, especially when processing high-dimensional data. This information loss may affect the model's recognition of subtle features in the image, thereby reducing the accuracy of target detection. Based on this, the present invention replaces the C2f network module with the HorNet network module. The HorNet module combines the local perception ability of traditional CNN and the global perception ability of Transformer, introduces global second-order convergence, enhances the feature extraction ability, can effectively capture global statistical information and subtle spatial relationships in the image, and improves the accuracy, detection speed and precision of disordered parts recognition.
[0055] In particular, the present invention introduces a lightweight network LightConv module to improve the YOLOv8 neck network. The LightConv module has a fixed context window. It determines the importance of elements through a set of weights that do not change with the time step, and realizes effective information processing of disordered parts from linear to nonlinear and then to linear. The LightConv module makes the recognition structure simpler, has fewer parameters, reduces the network complexity, and reduces the model's demand for computing resources while maintaining high-speed processing capabilities, thereby improving the accuracy, detection speed and precision of disordered parts recognition.
[0056] In particular, the present invention optimizes the loss function during YOLOv8 model training. In actual situations, the aspect ratio of the CIoU function is somewhat ambiguous. When processing targets with extreme scale differences, some unstable situations may occur. Based on this, the present invention replaces the CIoU function with the EIoU function and introduces the Focal-Loss function to solve the sample imbalance problem. Through the optimization of the EIoU function, the positioning of the sub-image area is further refined, the error caused by the fuzzy aspect ratio of the CIoU function is reduced, and the accuracy and robustness of the recognition of disordered parts are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 A schematic diagram of the steps of the unordered parts recognition method based on improved YOLOv8 according to an embodiment of the invention;
[0058] Figure 2 A logic block diagram for determining effective image texture of a part within a sub-image region according to an embodiment of the invention;
[0059] Figure 3 A logic block diagram of dividing sub-image areas into specific abnormality categories according to an embodiment of the invention;
[0060] Figure 4 A logic block diagram for identifying disordered parts according to an embodiment of the invention;
[0061] Figure 5 A schematic diagram of an improved YOLOv8 network structure of an embodiment of the invention;
[0062] Figure 6 Schematic diagram of the LightConv network structure of an embodiment of the invention. DETAILED DESCRIPTION
[0063] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0064] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0065] See also Figure 1 , Figure 1 The following is a schematic diagram of the steps of an unordered parts recognition method based on improved YOLOv8 according to an embodiment of the invention. An unordered parts recognition method based on improved YOLOv8 according to the present invention comprises:
[0066] Acquire a plurality of disordered part images, divide the disordered part images into a plurality of sub-image regions based on part contours, and determine effective image textures and contour complexity of the parts in the sub-image regions;
[0067] Determining a specific variation index of a sub-image region based on the number of effective image textures and contour complexity to classify the sub-image region into specific variation categories;
[0068] According to the specific mutation categories corresponding to the sub-image areas, the disordered parts are identified, including:
[0069] homogenizing the sub-image regions, setting labels for the sub-image regions and dividing data sets, and inputting the data sets into a preconfigured improved YOLOv8 model to identify part categories;
[0070] or, determining a part image within the sub-image area, and comparing it with a standard parts library to identify the part category;
[0071] The improvement process of the improved YOLOv8 model includes adding a ConTainer network module to the YOLOv8 model detection head, determining the HorNet network module as the backbone network of the YOLOv8 model, introducing a lightweight network LightConv module, and optimizing the loss function during model training.
[0072] Specifically, there is no limitation on the source of obtaining the plurality of disordered parts images. For example, it can be a real-time shot of a factory assembly line or a historically stored shot of the factory. Of course, both the shot and historically stored shots have been authorized and will not be elaborated here.
[0073] It is understandable that the part contour can be a single contour of a part or a superimposed contour of several parts superimposed on each other, as long as it can be used to divide the sub-image area. Technical personnel in this field can also choose the division method based on actual conditions, which will not be elaborated here.
[0074] Specifically, there is no limitation on the way of dividing the data set. For example, in implementation, 70% of the data is used for training, 15% of the data is used for verification, and 15% of the data is used for testing. Those skilled in the art may also adopt other division methods as long as they can ensure that the model can be trained normally and the training results can be obtained. This will not be elaborated.
[0075] Specifically, there is no limitation on the method of determining the standard parts library. For example, the existing open source parts library can be directly used as the standard parts library, or a number of parts images can be obtained in advance to establish a standard parts library by oneself. Technical personnel in this field can make a choice based on actual conditions, which will not be elaborated here.
[0076] Specifically, the present invention determines that the HorNet network module is the backbone network of the YOLOv8 model. Those skilled in the art can understand that in the existing models, the backbone networks are mostly C2f network modules, but in the actual operation process, the C2f network module will cause the loss of some important features when processing data, especially when processing high-dimensional data. This information loss may affect the model's recognition of subtle features in the image, thereby reducing the accuracy of target detection. Based on this, the present invention replaces the C2f network module with the HorNet network module. The HorNet module is a deep learning module based on recursive gated convolution. It combines the local perception ability of traditional CNN and the global perception ability of Transformer, introduces global second-order convergence, enhances the feature extraction ability, and can effectively capture global statistical information and subtle spatial relationships in the image, thereby improving the accuracy, detection speed and precision of disordered parts recognition.
[0077] See also Figure 2 , Figure 2The following is a logic block diagram of determining the effective image texture of a part in a sub-image region according to an embodiment of the invention. Specifically, the process of determining the effective image texture and contour complexity of a part in a sub-image region includes:
[0078] Determine a number of texture pixels on the part within the sub-image area and the corresponding texture width;
[0079] Performing grayscale relationship analysis based on the texture pixels;
[0080] Determine that the texture whose grayscale is less than or equal to the preset grayscale and whose texture width is within the reference width is a valid image texture;
[0081] Determine a plurality of edge contour segments of the part within the sub-image region;
[0082] The average curvature of the plurality of edge contour segments is determined as contour complexity.
[0083] Specifically, there is no limitation on the method of obtaining several texture pixels on the parts in the sub-image area. It only needs to be able to extract the texture. For example, a simple image segmentation algorithm can be used to directly perform image recognition on the sub-image area to obtain the texture features of the parts in the sub-image area. Other methods can also be used in this field for image recognition and extraction, which will not be repeated here.
[0084] Specifically, the grayscale relationship analysis is performed to determine the effective texture of the parts in the sub-image area. In actual situations, there may be scratches on the parts. If grayscale value analysis is not performed and all textures are identified as effective textures, the number of effective image textures will increase, resulting in deviations in the specific anomaly index and errors in the classification of the specific anomaly category in the sub-image area. Based on this, grayscale analysis of the texture of the parts in the sub-image area can make subsequent feature extraction and recognition more accurate.
[0085] Specifically, the preset grayscale is calculated in advance, a number of historical disordered part images are obtained in advance and the corresponding historical grayscales are obtained, and the average value of the historical grayscales is determined to be the preset grayscale.
[0086] Specifically, the reference width is calculated in advance, a number of historical disordered part images are obtained in advance and the corresponding historical widths are obtained, and the average value of the historical widths is determined as the reference width.
[0087] Specifically, there is no limitation on the division method of the edge contour segments. For example, in implementation, the edge contour segment can be divided into a number of contour segments of equal length. Those skilled in the art can also use other methods for division, which will not be described in detail.
[0088] It can be understood that the process of identifying the effective image texture and contour complexity of the sub-image area is relatively simple, which can effectively reduce computing power and improve computing efficiency.
[0089] Specifically, the process of determining the specific mutation index of the sub-image area includes:
[0090] Determine the ratio of the number of reference effective image textures to the number of effective image textures as the first influencing feature;
[0091] Determining a ratio of the baseline contour complexity to the contour complexity as a second influencing feature;
[0092] The weighted sum of the first influencing feature and the second influencing feature is determined as the specific mutation index.
[0093] Specifically, the number of reference valid image textures is pre-calculated, and the number of valid image textures of several historical sub-image regions is pre-acquired and the average value is calculated, and the average value is determined as the number of reference valid image textures.
[0094] Specifically, the reference contour complexity is obtained by pre-calculation, and the contour complexities of several historical sub-image regions are obtained in advance and the average value is calculated, and the average value is determined as the reference contour complexity.
[0095] Specifically, the sum of the weight coefficients of the first influencing feature and the second influencing feature is 1, the weight coefficient of the first influencing feature is 0.46, and the weight coefficient of the second influencing feature is 0.54.
[0096] See also Figure 3 , Figure 3 The following is a logic diagram of dividing the sub-image area into specific abnormality categories according to an embodiment of the invention. Specifically, the sub-image area is divided into specific abnormality categories, wherein:
[0097] If the specific mutation index is greater than or equal to a preset specific mutation index, the sub-image area is classified as a non-specific mutation category;
[0098] If the specific mutation index is less than a preset specific mutation index, the sub-image area is divided into a specific mutation category.
[0099] Specifically, the preset specific mutation index is selected within the interval [0.65, 0.78].
[0100] Specifically, the present invention determines the specific variation index of the sub-image area based on the effective image texture quantity and contour complexity, and provides data for dividing the specific variation category of the sub-image area. In actual situations, when sorting in a factory assembly line, most of the disordered parts are directly identified. Since the disordered parts contain various types of parts, their characteristics are different. Due to the difference in characteristics, the difficulty of back-end identification will be affected. For example, some parts with complex structures and complex textures are not easy to identify, and some parts are easy to identify. Therefore, if the same method is used to identify some ordered parts, it will lead to a waste of computing power and low efficiency of part recognition. Based on this, the present invention considers calculating the specific variation index of the sub-image area to divide the specific variation category. For the sub-image area of the non-specific variation category, the part category is directly identified according to the standard part library. For the sub-image area of the specific variation category, the YOLOv8 model is improved to identify the part category. In particular, the process of identifying the effective image texture quantity and contour complexity is relatively simple, which can reduce the computing power and improve the accuracy, detection speed and precision of the disordered part recognition.
[0101] See also Figure 4 , Figure 4 The following is a logic block diagram of the invention embodiment for identifying disordered parts. Specifically, the disordered parts are identified according to the specific mutation categories corresponding to the sub-image areas, wherein:
[0102] If the sub-image area is of a non-specific mutation category, then determining the part image in the sub-image area and comparing it with a standard parts library to identify the part category;
[0103] If the sub-image region is a specific mutation category, the sub-image region is homogenized, a label is set for the sub-image region, and a data set is divided, and the data set is input into a preconfigured improved YOLOv8 model to identify the part category.
[0104] See also Figure 5 , Figure 5 The improved YOLOv8 network structure diagram of the embodiment of the invention is shown in FIG. Specifically, the sub-image region homogenization includes:
[0105] Several sub-image areas are placed on the preset regional background image respectively.
[0106] Specifically, there is no limitation on the shape of the preset regional background image. Preferably, for ease of identification, in implementation, a rectangular background image is determined as the regional background image. Those skilled in the art may also make a selection based on actual conditions, which will not be elaborated here.
[0107] Specifically, the process of adding the ConTainer network module to the YOLOv8 model detection head includes:
[0108] Add the ConTainer network module to the detection head of the YOLOv8 model;
[0109] Integrate feature sets of different network layers based on the ConTainer network module and transfer the feature sets to the detection head;
[0110] The feature set includes convolutional layers for object classification and bounding box regression.
[0111] Specifically, the present invention adds a ConTainer network module to the YOLOv8 model detection head. The ConTainer network module is a general building block for multi-head context aggregation. The ConTainer network module is combined with the GNN, Transformer and MLP-Mixers methods, and the long-range interaction and local convolution operations are utilized to effectively integrate and understand the image. Without changing the size of the feature vector, the adaptability and recognition accuracy of the model for complex scenes are improved. Compared with the existing models, ConTainer achieves higher accuracy with fewer parameters and faster convergence speed, realizes more accurate prediction of the existence and position of the target, and improves the accuracy, detection speed and precision of disordered parts recognition.
[0112] See also Figure 6 , Figure 6 The schematic diagram of the LightConv network structure of the embodiment of the invention is shown in FIG. 1 . Specifically, the process of introducing the lightweight network LightConv module includes:
[0113] Determine that the LightConv module replaces the original C2f module of the YOLOv8 neck network;
[0114] Performing a preliminary linear transformation on the sub-image area based on the LightConv module;
[0115] Passing the sub-image region to a lightweight convolutional layer;
[0116] Combining the linear layer and the gating signal, the sub-image region after the light convolution is nonlinearized;
[0117] Transfer the nonlinearized sub-image regions to other linear layers.
[0118] Specifically, there is no limitation on the method of performing the preliminary linear transformation. For example, in implementation, the steps are as follows:
[0119] Project the features of the input sub-image region to expand the original dimension;
[0120] The features after expanding the original dimension are input into the gated linear unit, and the gated signal is activated using a nonlinear activation function;
[0121] Based on the gate signal, multiplying the gate signal with the activation signal to obtain an output value;
[0122] Use the LightConv module to normalize the weights of the features of the sub-image area;
[0123] The output value is fused with the features processed by the LightConv module to obtain the output features;
[0124] Of course, those skilled in the art may also use other methods to perform preliminary linear transformation, which will not be elaborated here.
[0125] It can be understood that through the effective information processing path of linear-nonlinear-linear, the YOLOv8 model can provide accurate capture of high-level features in sequence data while maintaining high computational efficiency.
[0126] Specifically, the present invention introduces a lightweight network LightConv module to improve the YOLOv8 neck network. It can be understood that the LightConv module is a lightweight convolution method, which is mainly used to reduce the computational complexity and number of parameters of traditional convolution while maintaining or improving the performance of the model. It has a fixed context window, and it determines the importance of elements through a set of weights that do not change with the time step, thereby realizing effective information processing of disordered parts from linear to nonlinear and then to linear. The LightConv module makes the recognition structure simpler, with fewer parameters, and reduces the network complexity. While maintaining high-speed processing capabilities, it reduces the model's demand for computing resources and improves the accuracy, detection speed and precision of disordered parts recognition.
[0127] Specifically, the process of optimizing the loss function during model training includes:
[0128] Replace the CIoU function with the EIoU function;
[0129] Based on the EIoU function, the Focal-Loss function is introduced;
[0130] Among them, the EIoU function calculates the difference values of width and height based on the CIoU function to replace the aspect ratio of the CIoU function.
[0131] Specifically, the CIoU function is a loss function for object detection. It improves on the traditional IoU and comprehensively considers factors such as the overlap of bounding boxes, center point distance, and aspect ratio to more accurately measure the similarity between two bounding boxes.
[0132] The EIoU function is an improved IoU loss function, which is used to measure the overlap between the predicted box and the real box in the target detection task. It is optimized based on CIoU, mainly by calculating the difference in width and height between the predicted box and the real box respectively, instead of only considering the aspect ratio;
[0133] The Focal-Loss function is an improved cross entropy loss function used to solve the category imbalance problem in target detection.
[0134] Specifically, the present invention optimizes the loss function during YOLOv8 model training. In actual situations, the aspect ratio of the CIoU function is somewhat ambiguous. When processing targets with extreme scale differences, some unstable situations may occur. Based on this, the present invention replaces the CIoU function with the EIoU function and introduces the Focal-Loss function to solve the sample imbalance problem. Through the optimization of the EIoU function, the positioning of the sub-image area is further refined, the error caused by the fuzzy aspect ratio of the CIoU function is reduced, and the accuracy and robustness of the recognition of disordered parts are improved.
[0135] Specifically, the process of comparing similarity with the standard parts library to identify the part category includes:
[0136] Determining the similarity between the part image in the sub-image area and the part image in the standard part library;
[0137] Determine that the part category in the standard parts library corresponding to the maximum similarity is the part category in the sub-image area;
[0138] Among them, a number of part images are pre-stored in the standard parts library.
[0139] Specifically, there is no limitation on the calculation method of similarity, for example, the structural similarity index method, the convolutional neural network feature comparison method, the feature matching method and the direct pixel difference method. In the implementation, in order to save computing power, the direct pixel difference method is used to determine the similarity of the part images. The specific steps are as follows:
[0140] Resize both images to the same size;
[0141] Determine each pair of corresponding pixels on the image and calculate the difference between the RGB values;
[0142] Evaluate the similarity of the images based on the calculated number of different pixels or the sum of the difference values;
[0143] It is understandable that if the difference is small, the two images are considered to be highly similar, and if the difference is large, the two images are considered to be less similar. It has higher reliability and recognition accuracy for parts with simpler features.
[0144] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0145] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for identifying disordered parts based on improved YOLOv8, characterized in that: include: Acquire a plurality of disordered part images, divide the disordered part images into a plurality of sub-image regions based on part contours, and determine effective image textures and contour complexity of the parts in the sub-image regions; Determining a specific variation index of a sub-image region based on the number of effective image textures and contour complexity to classify the sub-image region into specific variation categories; According to the specific mutation categories corresponding to the sub-image areas, the disordered parts are identified. include, homogenizing the sub-image regions, setting labels for the sub-image regions and dividing data sets, and inputting the data sets into a preconfigured improved YOLOv8 model to identify part categories; or, determining a part image within the sub-image area, and comparing it with a standard parts library to identify the part category; The improvement process of the improved YOLOv8 model includes adding a ConTainer network module to the YOLOv8 model detection head, determining the HorNet network module as the backbone network of the YOLOv8 model, introducing a lightweight network LightConv module, and optimizing the loss function during model training.
2. The unordered parts recognition method based on improved YOLOv8 according to claim 1, characterized in that: The process of determining the effective image texture and contour complexity of the part within the sub-image area includes: Determine a number of texture pixels on the part within the sub-image area and the corresponding texture width; Performing grayscale relationship analysis based on the texture pixels; Determine that the texture whose grayscale is less than or equal to the preset grayscale and whose texture width is within the reference width is a valid image texture; Determine a plurality of edge contour segments of the part within the sub-image region; The average curvature of the plurality of edge contour segments is determined as contour complexity.
3. The unordered parts recognition method based on improved YOLOv8 according to claim 1, characterized in that: The process of determining the specific mutation index of the sub-image area includes: Determine the ratio of the number of reference effective image textures to the number of effective image textures as the first influencing feature; Determining a ratio of the baseline contour complexity to the contour complexity as a second influencing feature; The weighted sum of the first influencing feature and the second influencing feature is determined as the specific mutation index.
4. The unordered parts recognition method based on improved YOLOv8 according to claim 1, characterized in that: The specific abnormality categories of the divided sub-image areas, wherein: If the specific mutation index is greater than or equal to a preset specific mutation index, the sub-image area is classified as a non-specific mutation category; If the specific mutation index is less than a preset specific mutation index, the sub-image area is divided into a specific mutation category.
5. The unordered parts recognition method based on improved YOLOv8 according to claim 1, characterized in that: The disordered parts are identified according to the specific mutation categories corresponding to the sub-image areas, wherein: If the sub-image area is of a non-specific mutation category, then determining the part image in the sub-image area and comparing it with a standard parts library to identify the part category; If the sub-image region is a specific mutation category, the sub-image region is homogenized, a label is set for the sub-image region, and a data set is divided, and the data set is input into a preconfigured improved YOLOv8 model to identify the part category.
6. The unordered parts recognition method based on improved YOLOv8 according to claim 1, characterized in that: The homogenizing the sub-image regions comprises: Several sub-image areas are placed on the preset regional background image respectively.
7. The unordered parts recognition method based on improved YOLOv8 according to claim 1, characterized in that: The process of adding the ConTainer network module to the YOLOv8 model detection head includes: Add the ConTainer network module to the detection head of the YOLOv8 model; Integrate feature sets of different network layers based on the ConTainer network module and transfer the feature sets to the detection head; The feature set includes convolutional layers for object classification and bounding box regression.
8. The unordered parts recognition method based on improved YOLOv8 according to claim 1, characterized in that: The process of introducing the lightweight network LightConv module includes: Determine that the LightConv module replaces the original C2f module of the YOLOv8 neck network; Performing a preliminary linear transformation on the sub-image area based on the LightConv module; Passing the sub-image region to a lightweight convolutional layer; Combining the linear layer and the gating signal, the sub-image region after the light convolution is nonlinearized; Transfer the nonlinearized sub-image regions to other linear layers.
9. The unordered parts recognition method based on improved YOLOv8 according to claim 1, characterized in that: The process of optimizing the loss function during model training includes: Replace the CIoU function with the EIoU function; Based on the EIoU function, the Focal-Loss function is introduced; Among them, the EIoU function calculates the difference values of width and height based on the CIoU function to replace the aspect ratio of the CIoU function.
10. The unordered parts recognition method based on improved YOLOv8 according to claim 1, characterized in that: The process of performing similarity comparison with the standard parts library to identify the part category includes: Determining the similarity between the part image in the sub-image area and the part image in the standard part library; Determine that the part category in the standard parts library corresponding to the maximum similarity is the part category in the sub-image area; Among them, a number of part images are pre-stored in the standard parts library.
Citation Information
Patent Citations
Unordered part identification method based on PointNet + + network
CN117237937A
Part identification method and system based on improved Mask R-CNN instance segmentation network
CN118366006A
Image recognition method, device and equipment and storage medium
CN112580707A
Non-standard part classification and retrieval method based on graph neural network
CN118230049A
Improved YOLOV8-based disordered part detection method and system
CN118608823A