Industrial product defect detection method and system
By using the YOLO object detection model and OpenCV image processing algorithm that accelerates deep learning training in industrial product defect detection, the problem of lightweight and high detection accuracy in the existing technology is solved, and efficient and accurate industrial strip defect detection is achieved, which is suitable for unattended intelligent production.
Patent Information
- Application Number
- CN202510077097.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
Existing industrial product defect detection technology has difficulties in taking into account both model lightweight and high detection accuracy, especially in mobile deployment of industrial equipment.
The YOLO object detection model based on accelerated deep learning training is adopted, combined with the image processing algorithm of OpenCV, the image preprocessing process is designed to improve the accuracy of the detection results.
It has improved the automation level and product quality of industrial strip steel production, achieved unattended intelligent high-quality production, and is suitable for the mobile deployment of industrial equipment.
Smart Images

Figure CN120013886A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial product defect detection, and in particular to an industrial product defect detection method and system. Background Art
[0002] Industrial strip steel is a kind of steel with high strength and corrosion resistance, which is widely used in many fields such as building structure, automobile manufacturing, shipbuilding, rail transportation, etc. The quality of industrial strip steel is directly related to the performance and safety of products in these fields. Therefore, it is very important to ensure the quality of industrial strip steel. Various defects such as cracks, pitting, plaques, inclusions, pressed oxide scale and scratches often occur in the production process of industrial strip steel. These defects will lead to the decline of product quality and may even cause production accidents. In addition, detecting and classifying these defects usually requires a lot of manpower and time, which increases production costs and cycles.
[0003] With the continuous improvement of industrial automation level, the industry usually uses industrial product surface defect detection algorithms to detect various defects that often occur in the production process. However, the current industrial defect detection algorithms mainly rely on methods based on image processing. Traditional image processing methods such as threshold segmentation, edge detection and region growing methods use image processing technology to distinguish the defective parts from the non-defective parts in the image, separate the defective parts from the image, and then process and analyze them. Although these traditional methods have achieved certain application effects in industrial detection, these methods are usually limited by the selection of features, the influence of factors such as illumination and noise, and often require manual definition of rules and features, resulting in relatively low robustness and generalization ability of the algorithm. The emergence of deep learning, through the use of deep neural network structure and end-to-end learning methods, can automatically learn features and achieve better performance and robustness in industrial detection.
[0004] At present, the mainstream algorithms for surface defect detection of industrial products based on deep learning can be roughly divided into two categories: single-stage and double-stage. The double-stage method generates candidate regions through a candidate box generator and then classifies the candidate regions. Representative algorithms include the RCNN series. The single-stage method uses a convolutional neural network for end-to-end detection, directly generates multiple bounding boxes on the image, and performs positioning and classification. Representative algorithms include the YOLO series and SSD. Compared with the double-stage method, the single-stage method has faster detection efficiency and is more suitable for industrial defect detection scenarios. Therefore, many researchers have studied the single-stage algorithm. For example, Li et al. proposed a deep learning model based on a multi-scale feature extraction module. Although it can effectively enhance the feature extraction capability, it is not suitable for deployment on hardware devices with low computing power due to the large number of parameters. In order to solve the problem of network feature misalignment, Yu et al. proposed a dense feature pyramid network (AD-FPN) to refine the scale difference and perform effective alignment, thereby alleviating the problem of feature misalignment in the FPN-based method, but the large increase in the number of parameters limits practical applications. Ma et al. proposed the STYOLO model, which makes full use of high-quality samples to adjust data distribution and optimize the training process through adaptive sampling and dynamic label allocation, but affects the detection rate. Xie et al. proposed a solution to the problem that it is difficult to achieve both high efficiency and high precision. First, the YOLO model is simplified by deep separable convolution and rate set connection. Secondly, the feature pyramid network is improved to enhance the correlation of multi-scale detection target position. However, these improvements introduce additional parameters and computational complexity, which have a certain impact on the detection speed. Zhao et al. designed a DFPN network to improve the neck detection effect. The feature representation is enhanced through a dual feature pyramid structure to achieve rich information extraction and make full use of low-level features. However, when dealing with dense defects, missed detection is prone to occur. Although the above methods have achieved certain results in improving defect detection accuracy, there is still a problem that the model lightweight and high detection accuracy cannot be taken into account at the same time, and they are not applicable to mobile deployment in industrial equipment.
[0005] Therefore, how to provide a new type of industrial defect detection technology to improve the automation level and product quality of industrial strip steel production and realize unmanned intelligent high-quality production is a problem that needs to be solved urgently. Summary of the invention
[0006] The embodiments of the present invention provide an industrial product defect detection method and system to solve the above technical problems in the prior art.
[0007] In order to have a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended to be a general review, nor is it intended to identify key / important components or to delineate the scope of protection of these embodiments. Its only purpose is to present some concepts in a simple form as a preface to the detailed description that follows.
[0008] According to a first aspect of an embodiment of the present invention, a method for detecting defects in an industrial product is provided.
[0009] In one embodiment, the industrial product defect detection method includes:
[0010] Real-time collection of image data of industrial products to be inspected;
[0011] Performing image preprocessing on the image data, and inputting the image data after image preprocessing into a preconfigured defect target detection model;
[0012] Using the defect target detection model to perform defect target detection on the image data after image preprocessing, and output the defect detection result of the industrial product;
[0013] Among them, the defect target detection model is a YOLO target detection model based on accelerated deep learning training.
[0014] In one embodiment, performing image preprocessing on the image data includes:
[0015] Read the color image data collected by the camera in real time based on the OpenCV computer vision library;
[0016] Convert the color image data into grayscale image data, and use the OSTU algorithm to perform binarization processing on the grayscale image data to divide the grayscale image data into foreground image data and background image data;
[0017] Based on the foreground image data and the background image data, using the Canny edge detection algorithm to extract defect edge information;
[0018] The connected domain analysis is performed on the extracted defect edge information to obtain the connected domain contour, and the obtained connected domain contour is drawn on the original image and the defect information is marked.
[0019] In one embodiment, pre-configuring a defect object detection model includes:
[0020] Acquire a historical industrial product defect image dataset, and perform data preprocessing on the historical industrial product defect image dataset to obtain a defect sample dataset;
[0021] Dividing the defect sample data set into a training set, a validation set and a test set;
[0022] Based on the training set, validation set, and test set, the automatic mixed precision algorithm is used to accelerate the deep learning training of the YOLO target detection model to obtain the defect target detection model.
[0023] In one embodiment, data preprocessing is performed on the historical industrial product defect image dataset to obtain a defect sample dataset including:
[0024] Performing data cleaning on the historical industrial product defect image dataset, and using a labeling tool to label and / or correct the bounding box and category label of each defect area in the historical industrial product defect image dataset after the data cleaning;
[0025] Based on the annotated and / or corrected historical industrial product defect image dataset, the historical industrial product defect image dataset is screened according to defect elements to ensure that the historical industrial product defect image samples of each defect category are evenly distributed;
[0026] The screened historical industrial product defect image dataset is subjected to data enhancement and format conversion to obtain a defect sample dataset.
[0027] In one embodiment, the defect elements include: defect type, defect location and defect area size; the data enhancement processing includes data rotation processing, data flipping processing and data scaling processing.
[0028] According to a second aspect of an embodiment of the present invention, an industrial product defect detection system is provided.
[0029] In one embodiment, the industrial product defect detection system comprises:
[0030] A data acquisition module, used for real-time acquisition of image data of industrial products to be inspected;
[0031] A data processing module, used for performing image preprocessing on the image data, and inputting the image data after image preprocessing into a preconfigured defect target detection model;
[0032] A defect detection module, used to perform defect target detection on the image data after image preprocessing using the defect target detection model, and output defect detection results of industrial products;
[0033] Among them, the defect target detection model is a YOLO target detection model based on accelerated deep learning training.
[0034] In one embodiment, when the data processing module performs image preprocessing on the image data, the color image data collected by the camera in real time is read based on the OpenCV computer vision library; the color image data is converted into grayscale image data, and the grayscale image data is binarized using the OSTU algorithm to divide the grayscale image data into foreground image data and background image data; based on the foreground image data and the background image data, the Canny edge detection algorithm is used to extract defect edge information; a connected domain analysis is performed on the extracted defect edge information to obtain a connected domain contour, and the obtained connected domain contour is drawn on the original image and the defect information is marked.
[0035] In one embodiment, the industrial product defect detection system also includes: a model configuration module, which is used to pre-configure a defect target detection model, and the model configuration module obtains a historical industrial product defect image dataset when pre-configuring the defect target detection model, and performs data preprocessing on the historical industrial product defect image dataset to obtain a defect sample dataset; divides the defect sample dataset into a training set, a validation set, and a test set; based on the training set, the validation set, and the test set, uses an automatic mixed precision algorithm to accelerate deep learning training of the YOLO target detection model to obtain a defect target detection model.
[0036] In one embodiment, when the model configuration module performs data preprocessing on the historical industrial product defect image dataset to obtain a defect sample dataset, the model configuration module performs data cleaning on the historical industrial product defect image dataset, and uses a labeling tool to label and / or correct the boundary box and category label of each defect area in the historical industrial product defect image dataset after data cleaning; based on the labeled and / or corrected historical industrial product defect image dataset, the historical industrial product defect image dataset is filtered according to defect elements to ensure that the distribution of historical industrial product defect image samples of each defect category is balanced; the filtered historical industrial product defect image dataset is subjected to data enhancement processing and format conversion processing to obtain a defect sample dataset.
[0037] In one embodiment, the defect elements include: defect type, defect location and defect area size; the data enhancement processing includes data rotation processing, data flipping processing and data scaling processing.
[0038] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:
[0039] The present invention is based on the image processing algorithm of OpenCV2 and the YOLO deep learning target detection algorithm, combined with the industrial product defect dataset, fully considering the possible high temperature environment, incomplete detection area, light interference and other practical factors in the production environment, designs the image preprocessing process, and uses a highly robust defect model to improve the accuracy of the results, improve the automation level and product quality of industrial strip steel production, and lay the foundation for unmanned intelligent high-quality production.
[0040] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0042] Figure 1 is a flow chart of a method for detecting defects in industrial products according to an exemplary embodiment;
[0043] Figure 2 is a structural block diagram of an industrial product defect detection system according to an exemplary embodiment;
[0044] Figure 3 The figure is a schematic diagram showing the structure of a computer device according to an exemplary embodiment. DETAILED DESCRIPTION
[0045] The following description and accompanying drawings fully illustrate the specific embodiments of this article so that those skilled in the art can practice them. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. The scope of the embodiments of this article includes the entire scope of the claims, as well as all available equivalents of the claims. Herein, the terms "first", "second", etc. are only used to distinguish one element from another, without requiring or implying any actual relationship or order between these elements. In fact, the first element can also be called the second element, and vice versa. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the structure, device or equipment including a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also include elements inherent to such structure, device or equipment. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the structure, device or equipment including the elements. Each embodiment is described in a progressive manner herein, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the embodiments can be referred to each other.
[0046] The terms "longitudinal", "lateral", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like used herein to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing this article and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In the description herein, unless otherwise specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a mechanical connection or an electrical connection, it can also be the internal connection of two elements, it can be a direct connection, or it can be an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0047] As used herein, the term "plurality" means two or more than two, unless otherwise specified.
[0048] In this document, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.
[0049] In this article, the term "and / or" is a description of the association relationship between objects, indicating that three relationships may exist. For example, A and / or B means: A or B, or, A and B.
[0050] It should be understood that, although the various steps in the flow chart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0051] Each module in the device or system of the present application can be implemented in whole or in part by software, hardware, or a combination thereof. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to the above modules.
[0052] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0053] Figure 1 An embodiment of an industrial product defect detection method of the present invention is shown.
[0054] In this optional embodiment, the industrial product defect detection method includes:
[0055] Step S101, collecting image data of the industrial product to be inspected in real time;
[0056] Step S102, performing image preprocessing on the image data, and inputting the image data after image preprocessing into a pre-configured defect target detection model; wherein the defect target detection model is a YOLO target detection model based on accelerated deep learning training;
[0057] Step S103, using the defect target detection model to perform defect target detection on the image data after image preprocessing, and outputting the defect detection result of the industrial product.
[0058] Figure 2 An embodiment of an industrial product defect detection system of the present invention is shown.
[0059] In this optional embodiment, the industrial product defect detection system includes:
[0060] The data acquisition module 201 is used to acquire image data of the industrial product to be inspected in real time;
[0061] The data processing module 202 is used to perform image preprocessing on the image data, and input the image data after image preprocessing into a pre-configured defect target detection model; wherein the defect target detection model is a YOLO target detection model based on accelerated deep learning training;
[0062] The defect detection module 203 is used to perform defect target detection on the image data after image preprocessing by using the defect target detection model, and output the defect detection result of the industrial product.
[0063] In the above optional embodiment, when performing image preprocessing on the image data, the color image data collected by the camera in real time is read based on the OpenCV computer vision library; the color image data is converted into grayscale image data, and the grayscale image data is binarized using the OSTU algorithm to divide the grayscale image data into foreground image data and background image data; based on the foreground image data and background image data, the Canny edge detection algorithm is used to extract the defect edge information; the extracted defect edge information is subjected to a connected domain analysis to obtain a connected domain contour, and the obtained connected domain contour is drawn on the original image and the defect information is marked.
[0064] Specifically, the cv2.drawContours() function of the OpenCV computer vision library is used to draw contours and the cv2.putText() function is used to add text annotations. The processed image is displayed in the window, the cv2.imshow() function is used to display the image, and the cv2.waitKey(1) function is used to wait for key events to realize real-time processing of the video stream. After the processing is completed, the camera resources are released and all OpenCV windows are closed.
[0065] As for the OSTU algorithm, the OTSU algorithm (Otsu's method), also known as the Otsu method or the maximum inter-class variance method, is a threshold selection technology widely used in the field of image processing, mainly used for binary segmentation of grayscale images. Since the background and defects on the steel strip have significant grayscale differences, the present invention mainly uses this method to obtain the defective area of the steel strip. The algorithm is divided into 5 main steps, as follows:
[0066] Step 1: Calculate the image histogram. Calculate the histogram of the pixel value distribution of the input grayscale image. The histogram reflects the number of pixels at each grayscale level in the image and is the basis for subsequent calculations.
[0067] Suppose there is a grayscale image of size M×N, where the grayscale value of each pixel ranges from 0 to L-1 (for example, for an 8-bit image, L=256). The histogram is an array H of length L, where H(i) represents the number of times grayscale level i appears in the image.
[0068]
[0069] Where: H(i) is the number of pixels at gray level i; I(x,y) is the gray value of the image at position (x,y); δ(a,b) is the Kronecker delta function, when a=b, δ(a,b)=1; otherwise, δ(a,b)=0; M and N are the height and width of the image respectively; i is the gray level from 0 to L-1.
[0070] The above formula means that every pixel of the entire image is traversed, and if the grayscale value of a pixel is equal to i, H(i) increases by 1. Finally, each element H(i) in the array HH represents the number of times the grayscale level i appears in the entire image.
[0071] Step 2: Traverse all possible thresholds: For each possible threshold T in the image, divide all pixels into two categories: one is pixels with grayscale values less than or equal to T, which are considered background; the other is pixels with grayscale values greater than T, which are considered foreground. Based on these divisions, the number of pixels (or pixel ratio), mean μ0 and μ1, and between-class variance of the two categories of pixels can be calculated
[0072] In the OTSU algorithm, the process of traversing all possible thresholds is actually trying the grayscale from the minimum to the maximum one by one. Specifically, if the grayscale range of an image is from 0 to 255 (i.e., an 8-bit grayscale image), then the possible threshold T is traversed from 0 to 255.
[0073] At each step of the traversal, the current threshold T is used as a decision point to divide all pixels in the image into two categories: background (all pixels with grayscale values less than or equal to T) and foreground (all pixels with grayscale values greater than T).
[0074] The number of pixels in each category (or the proportion of total pixels), denoted as w0(T) and w1(T)
[0075]
[0076] The average gray value (mean) of each category is denoted as μ0 (mean of background) and μ1 (mean of foreground).
[0077]
[0078] Among them, p(i) is the probability of gray value i appearing, which is obtained by normalizing the grayscale histogram of the image, w0(T) and w1(T) are the ratio of background and foreground pixels to the total number of pixels, μ0 and μ1 are the average grayscale values of these two categories, and L is the maximum grayscale level of the image plus 1.
[0079] Step 3: Calculate the inter-class variance. The inter-class variance is a statistic that measures the difference between two types of pixels (foreground and background), and is defined as the square of the difference between the mean values of the two types of pixels multiplied by the ratio of the total number of pixels in the two types. The formula is as follows:
[0080]
[0081] Among them, n0 and n1 are the number of background and foreground pixels respectively, and μ0 and μ1 are the means of the two types of pixels respectively.
[0082] Step 4: Find the maximum inter-class variance. Traverse all possible thresholds T and calculate the inter-class variance corresponding to each threshold. Select the threshold T that maximizes the inter-class variance. %pt as the optimal threshold.
[0083] Step 5: Binarize the image. Use the optimal threshold T found %pt Binarize the original image: set the pixel value to be less than or equal to the threshold value T %pt The pixels with values greater than the threshold T are set as the background color (such as black), and the pixels with values greater than the threshold T are set as %pt The pixels are set as the foreground color (such as white). In this way, the image is effectively divided into foreground and background.
[0084] Connected Component Analysis (CCA) is an important method in image processing, which is used to identify and mark the area composed of pixels with the same pixel value and adjacent positions in the image. In order to further obtain these characteristics in the steel strip, the present invention performs connected domain analysis based on the binary image obtained by the above method. The main steps are as follows:
[0085] Step 1: Seed point selection: Select an unvisited foreground pixel from the binary image as the starting point (seed point).
[0086] Step 2: Region growing: Starting from the seed point, all adjacent pixels with the same value are added to the current connected domain along the connectivity rule.
[0087] Step 3: Label assignment: Assign a unique label to the current connected domain and record its attributes (such as area, centroid, etc.).
[0088] Step 4: Traverse the image. Continue to search for unvisited foreground pixels and repeat the above process until all foreground pixels are assigned to their respective connected domains.
[0089] Step 5: Eliminate small and large patches. Industrial defects in steel strips are generally continuous and have a large area, but there are closed boundaries. Therefore, setting an area threshold can eliminate a lot of interference. This article has tried to optimize the area threshold many times, and finally set patches with an area less than 50 pixels and an area greater than 10,000 pixels as noise.
[0090] After the connected domain analysis, most of the point noise and flake noise in the steel strip will be eliminated. However, in the denoising result, there are long strips of noise, which mainly come from the boundary area of the background, and further processing is still needed for this noise. (In the field of image processing, connected domain analysis is an effective method to remove isolated noise points, but for long strips or linear noise, other technologies need to be used for further optimization. Effective processing can be performed through frequency domain filtering technology. The specific steps include: first, perform a fast Fourier transform (FFT) on the original image to convert it from the spatial domain to the frequency domain; then analyze the spectrum to determine the specific frequency range corresponding to the long strip noise; then, design and apply a bandstop filter, which is specifically used to attenuate these specific frequency components, thereby effectively suppressing the long strip noise; after filtering, the image is converted back to the spatial domain through the inverse fast Fourier transform (IFFT); finally, check the effect of the processed image, and if necessary, adjust the filter parameters appropriately to obtain a better denoising effect. This method can specifically remove long strip noise in the background boundary area and other locations to improve image quality.)
[0091] In the above optional embodiment, pre-configuring the defect target detection model includes: obtaining a historical industrial product defect image dataset, and performing data preprocessing on the historical industrial product defect image dataset to obtain a defect sample dataset;
[0092] The defect sample data set is divided into a training set, a validation set and a test set. Based on the training set, the validation set and the test set, the YOLO target detection model is accelerated by using an automatic mixed precision algorithm to perform deep learning training to obtain a defect target detection model.
[0093] When the data preprocessing is performed on the historical industrial product defect image dataset to obtain the defect sample dataset, the historical industrial product defect image dataset is subjected to data cleaning processing, and the boundary box and category label of each defect area of the historical industrial product defect image dataset after data cleaning processing are annotated and / or corrected using annotation tools; based on the annotated and / or corrected historical industrial product defect image dataset, the historical industrial product defect image dataset is screened according to defect elements to ensure that the historical industrial product defect image samples of each defect category are evenly distributed; the screened historical industrial product defect image dataset is subjected to data enhancement processing and format conversion processing to obtain the defect sample dataset. Wherein, the defect elements include: defect type, defect location and defect area size; the data enhancement processing includes data rotation processing, data flipping processing and data scaling processing.
[0094] In specific applications, first of all, data cleaning includes removing invalid or low-quality images, such as blurred, underexposed or overexposed images, which can be automatically detected or manually checked and deleted using image quality assessment (IQA) tools. Secondly, in order to ensure that each defect area has the correct bounding box and category label, it is also necessary to check and correct the annotation errors, which can be annotated and corrected using annotation tools such as LabelMe, LabelImg, etc. In addition, in order to ensure that the data set covers various situations, data screening can also involve selecting the most representative samples, ensuring that the distribution of samples of each category is balanced through statistical analysis methods, and screening according to factors such as defect type, location, size, etc. In addition, data enhancement techniques (such as rotation, flipping, scaling, etc.) can be used to increase the diversity and representativeness of the data set. Finally, format conversion ensures that the data set meets the input requirements of the target detection model, such as converting the annotated data into YOLO format or COCO format, each image corresponds to a .txt file, containing the category ID and normalized bounding box coordinates of the defect, and scripts can be written to automate this process, read the original annotation file (such as XML or JSON format), parse the annotation information and generate a file in the target format.
[0095] When testing the defect target detection model, install the necessary dependent libraries, such as OpenCV, NumPy, and UltralyticsYOLOv8, and configure the model's hyperparameters and dataset path. Use the training set to train the model, regularly evaluate the performance on the validation set, monitor the loss curve and accuracy changes, and use techniques such as early stopping and learning rate decay to optimize the training. Evaluate the model performance on an independent test set, and calculate indicators such as Precision, Recall, F1Score, and mAP. To further optimize the model, you can use transfer learning and ensemble learning, and apply NMS technology to improve the detection quality. Finally, deploy the trained model to the actual application scenario, write the inference code, monitor the model performance for a long time, and continuously optimize the model based on feedback. Through these steps, the defect target detection model can be effectively trained and optimized to ensure its high accuracy and stability in practical applications.
[0096] In specific applications, most of the images in the industrial product defect dataset have a lot of noise, but the placement posture is basically fixed, and the defect target detection model is suitable for the current situation. Since the denoising algorithm itself will bring certain uncertainties, and the YOLOv8n model itself requires certain noise samples to enhance the robustness of the model, the original noise can be retained during the training process.
[0097] The .json file in the training set provides the upper left coordinates, lower right coordinates and label values of the industrial defect area in each image. This information is converted into the relative position, width and length of the center point of the marking box as the input of the model training to eliminate the influence of different sizes caused by different images. The format of the newly converted label file is txt format. These data are normalized into the COCO dataset format for model input, and some image samples are divided into the validation set to measure the model effect. The results of some target boxes before and after the conversion are shown in Table 1:
[0098] Table 1 is the conversion of JSON file to TXT file
[0099]
[0100] As for the establishment of defect target detection model, based on the requirements of "fast and accurate", the YOLOv8 Nano model is established as the industrial defect target detection model. The YOLOv8 Nano model is a lightweight variant of the YOLOv8 series model. As a newer structure of the You Only Look Once model series, YOLOv8 abandons the traditional anchor frame-based design and uses the Anchor-Free method instead. This change helps to reduce computing resource consumption and has the potential to improve detection accuracy and robustness. It can eliminate various interferences in industrial defect target recognition and has better performance in image recognition with different aspect ratios. YOLOv8 Nano reduces the structure of the original model as much as possible, sacrificing some data accuracy in exchange for extremely high image reasoning efficiency and training speed, reducing the time cost of model training, and is also more friendly to lower-performance devices.
[0101] The present invention builds the YOLOv8n model based on the PyTorch and Ultralytics framework, which contains 22 layers of structural blocks, a total of 225 layers, a total number of parameters of 3011043, and a required computing power of 8.2FLOPS. The detailed list of the model structure is shown in Table 2:
[0102] Table 2 is the YOLOv8n model structure table
[0103]
[0104]
[0105] The training of the YOLOv8 Nano model can be completed on a personal computer with an AMD Ryzen 76800H CPU and an NVIDIA GeForce RTX 3060 Laptop GPU GDDR6@6GB (192bits). The deep learning software environment involved is CUDA11.8+cuDNN8.9.2+Torch2.2.2, the iteration cycle is set to 100 times, and the learning rate is initialized to 0.0001 according to the sample size and dynamically adjusted during the training process.
[0106] Automatic Mixed Precision (AMP) is a technology for accelerating deep learning training. It automatically converts part of the calculation from single precision (float32) to half precision (float16) during the training process, thereby reducing memory usage and improving the calculation speed, especially on hardware that supports half-precision operations (such as NVIDIA GPU). In addition to the traditional deep learning hyperparameter settings, the present invention adopts this method to accelerate the training process of the industrial defect segmentation model. During the training process, some data containers are randomly assigned to float16 type during the forward and backward propagation of YOLOv8n. Although a certain degree of precision loss is caused, the video memory occupancy rate and iteration efficiency are significantly reduced, from the original average 4.57GB video memory to 2.21GB, and the iteration efficiency is increased from the original 1.71it / s to 3.11it / s. It only takes 10 seconds to complete the training cycle. Although there are certain fluctuations in accuracy during the training process, the final accuracy is less affected.
[0107] In the pre-trained model provided by the Ultralytics framework, the learning rate is set to a dynamic learning rate, and the learning rate is dynamically reduced by 20% when the loss rate drops to different thresholds. The optimizer selects SGD (stochastic gradient descent) to increase the model training speed, the batch size is set to 1, and the epoch is set to 100. During the training process, the accuracy, recall rate, mAP95 and other indicators are steadily improved, and finally converge to more than 95%.
[0108] The main accuracy indicators involved in YOLOv8n training include Box Loss, Classification loss, Distribution Focal Loss, Precision, and Recall. The detailed description of each indicator is as follows:
[0109] Box Loss is a key component in object detection algorithms, especially in one-stage detectors such as YOLOv8. Box Loss is used to measure the difference between the bounding box predicted by the model and the actual annotated bounding box, with the goal of training the model to accurately predict the location of the target object. Box Loss includes multiple sub-losses, such as IoU loss, and the calculation formula is as follows:
[0110]
[0111] Among them, Area of Intersection represents the intersection area, and Area of Union represents the combined area.
[0112] Classification loss is used to measure the difference between the model's prediction of the input data and the true category label. It is part of the loss function that needs to be minimized during model training. In industrial defect detection, it mainly represents the gap between the defect area and the background area. It is quantified by binary cross entropy (BCE Loss). The formula is as follows:
[0113]
[0114] Where N is the number of candidate boxes; y i is the actual label of the i-th candidate box (0 or 1); is the model prediction score of the i-th candidate box, which is also between 0 and 1 after being processed by the sigmoid function.
[0115] Distribution Focal Loss is a loss function used to optimize bounding box prediction. It combines the idea of Focal Loss and aims to reduce the impact of easy-to-classify samples (usually easy-to-distinguish positive and negative samples) on the loss function during training, so as to pay more attention to samples that are difficult to classify or have large bounding box position offsets. The formula is as follows:
[0116] DFL(p i ,t i )=-α i ·(1-p i ) 3 IoU loss (t i )
[0117] Among them, p i is the confidence (or IoU prediction value) of the model's prediction of the i-th sample bounding box, reflecting the degree of overlap between the model's prediction box and the true box; t i is the true bounding box of the i-th sample; α i is the class weight of the sample (can be a fixed value or related to IoU), which is used to adjust the loss contribution of positive and negative samples or samples of different difficulty; γ is the focus parameter, which controls the intensity of the focus effect and is usually set to a value greater than 0; IoU loss (t i ) is a bounding box regression loss calculated based on IoU or its improved version (such as CIOU), which measures the difference between the predicted box and the true box position.
[0118] Precision and Recall are traditional accuracy verification indicators, and the calculation formula is as follows:
[0119]
[0120]
[0121] Among them, TP is a true positive (predicted as positive and actually positive), FP is a false positive (predicted as positive but actually negative), and FN is a false negative (predicted as negative but actually positive).
[0122] Figure 3 An embodiment of a computer device of the present invention is shown. The computer device may be a server, and the computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store static information and dynamic information data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the steps in the above method embodiment are implemented.
[0123] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0124] In addition, the present invention also provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiment when executing the computer program.
[0125] In addition, the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.
[0126] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0127] The present invention is not limited to the structures which have been described above and shown in the drawings, and various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A method for detecting defects in industrial products, characterized in that: include: Real-time collection of image data of industrial products to be inspected; Performing image preprocessing on the image data, and inputting the image data after image preprocessing into a preconfigured defect target detection model; Using the defect target detection model to perform defect target detection on the image data after image preprocessing, and output the defect detection result of the industrial product; Among them, the defect target detection model is a YOLO target detection model based on accelerated deep learning training.
2. The industrial product defect detection method according to claim 1, characterized in that: Performing image preprocessing on the image data includes: Read the color image data collected by the camera in real time based on the OpenCV computer vision library; Converting the color image data into grayscale image data, and performing binarization processing on the grayscale image data using the OSTU algorithm to divide the grayscale image data into foreground image data and background image data; Based on the foreground image data and the background image data, using the Canny edge detection algorithm to extract defect edge information; The connected domain analysis is performed on the extracted defect edge information to obtain the connected domain contour, and the obtained connected domain contour is drawn on the original image and the defect information is marked.
3. The industrial product defect detection method according to claim 1, characterized in that: Pre-configured defect object detection models include: Acquire a historical industrial product defect image dataset, and perform data preprocessing on the historical industrial product defect image dataset to obtain a defect sample dataset; Dividing the defect sample data set into a training set, a validation set and a test set; Based on the training set, validation set, and test set, the automatic mixed precision algorithm is used to accelerate the deep learning training of the YOLO target detection model to obtain the defect target detection model.
4. The industrial product defect detection method according to claim 3, characterized in that: The defect sample dataset obtained by performing data preprocessing on the historical industrial product defect image dataset includes: Performing data cleaning on the historical industrial product defect image dataset, and using a labeling tool to label and / or correct the bounding box and category label of each defect area in the historical industrial product defect image dataset after the data cleaning; Based on the annotated and / or corrected historical industrial product defect image dataset, the historical industrial product defect image dataset is screened according to defect elements to ensure that the historical industrial product defect image samples of each defect category are evenly distributed; The screened historical industrial product defect image dataset is subjected to data enhancement and format conversion to obtain a defect sample dataset.
5. The industrial product defect detection method according to claim 4, characterized in that: The defect elements include: defect type, defect location and defect area size; the data enhancement processing includes data rotation processing, data flipping processing and data scaling processing.
6. An industrial product defect detection system, characterized in that: include: A data acquisition module, used for real-time acquisition of image data of industrial products to be inspected; A data processing module, used for performing image preprocessing on the image data, and inputting the image data after image preprocessing into a preconfigured defect target detection model; A defect detection module, used to perform defect target detection on the image data after image preprocessing using the defect target detection model, and output defect detection results of industrial products; Among them, the defect target detection model is a YOLO target detection model based on accelerated deep learning training.
7. The industrial product defect detection system according to claim 6, characterized in that: When performing image preprocessing on the image data, the data processing module reads the color image data collected by the camera in real time based on the OpenCV computer vision library; converts the color image data into grayscale image data, and uses the OSTU algorithm to binarize the grayscale image data, and divides the grayscale image data into foreground image data and background image data; based on the foreground image data and the background image data, uses the Canny edge detection algorithm to extract defect edge information; performs connected domain analysis on the extracted defect edge information to obtain a connected domain outline, and draws the obtained connected domain outline on the original image and marks the defect information.
8. The industrial product defect detection system according to claim 6, characterized in that: Also includes: A model configuration module is used to pre-configure a defect target detection model, and when pre-configuring the defect target detection model, the model configuration module obtains a historical industrial product defect image data set, and performs data preprocessing on the historical industrial product defect image data set to obtain a defect sample data set; divides the defect sample data set into a training set, a validation set, and a test set; based on the training set, the validation set, and the test set, uses an automatic mixed precision algorithm to accelerate deep learning training of the YOLO target detection model to obtain a defect target detection model.
9. The industrial product defect detection system according to claim 8, characterized in that: When the model configuration module performs data preprocessing on the historical industrial product defect image dataset to obtain a defect sample dataset, the model configuration module performs data cleaning on the historical industrial product defect image dataset, and uses a labeling tool to label and / or correct the boundary box and category label of each defect area in the historical industrial product defect image dataset after data cleaning; based on the labeled and / or corrected historical industrial product defect image dataset, the historical industrial product defect image dataset is screened according to defect elements to ensure that the distribution of historical industrial product defect image samples of each defect category is balanced; and data enhancement processing and format conversion processing are performed on the screened historical industrial product defect image dataset to obtain a defect sample dataset.
10. The industrial product defect detection system according to claim 9, characterized in that: The defect elements include: defect type, defect location and defect area size; the data enhancement processing includes data rotation processing, data flipping processing and data scaling processing.