A die casting product defect detection method based on an improved YOLOX model
By improving the YOLOX model and combining it with the ShuffleNetV2-plus network and ECA attention mechanism, the problems of incomplete type identification and low accuracy in die-cast product defect detection are solved, and efficient automated detection of die-cast product defects is achieved.
Patent Information
- Application Number
- CN202310330040.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-03-30
AI Technical Summary
Existing technologies for defect detection in die-cast products suffer from insufficient comprehensiveness in defect identification and low accuracy, hindering the automation, intelligentization, and greening of die-cast product production.
An improved YOLOX model, combined with ShuffleNetV2-plus network and ECA attention mechanism, is adopted. Through dataset expansion and image processing, the model's ability to identify defects in die-cast products, including watermarks, blistering, shrinkage cavities, etc., is improved.
It improves the accuracy and efficiency of defect detection in die-cast products, reduces the cost of manual labeling, and supports the automation and intelligentization of die-cast product production.
Smart Images

Figure CN116468679B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the field of die casting product defect detection, and is applied to a die casting product production process, in particular to a die casting product defect detection method based on an improved YOLOX model. BACKGROUND
[0002] Metal die casting products are widely used, have a large production quantity, and have many product types, and play a vital role in the industry. The die casting products have excellent characteristics and are widely used in many part fields such as automobiles, aerospace, internal combustion engines and the like. In the face of a huge market in China, the application of the die casting products also has a wide prospect and demand, and how to build an automatic, intelligent and green die casting production workshop is a higher requirement proposed in the new era.
[0003] In the production process of the die casting product, due to improper process control or due to a mold and the like, the die casting product may have defects, and manual identification and screening are required to ensure product quality. Due to high temperature and large noise in the die casting workshop, and high work pressure, the body and mind of the detection personnel will be greatly affected, and manual identification and screening are low in efficiency and accuracy, thereby hindering the automation, intelligence and green process of the die casting product production, and therefore, how to improve the automatic detection and efficiency of the die casting product defects is crucial.
[0004] A die casting defect detection method based on deep learning improves the YOLACT algorithm capable of identifying defects and performing semantic segmentation (Research on Casting Defect Detection Method Based on Deep Learning_Peng Lei), and in the experiment, there are 67 kinds of sub-defects and 2727 defect images, and the defect recognition rate is improved from 62.0% to 65.8%, and the detection rate is also optimized. Mery et al. generate two kinds of aluminum die casting product defect sets through a three-dimensional ellipsoid model and a GAN model, the defect set is composed of a large number of normal products and a small number of defective products, and the results show that the three-dimensional ellipsoid model is more effective than the GAN model, and the mAP reaches 71.02%, and the research has the phenomena that the defect types of the die casting product are few and the recognition accuracy is not high. At present, there are still some problems to be solved in the research on the die casting product defect visual detection, for example, the defect type recognition of the actual production die casting product is not enough, so that the actual production situation cannot be reflected, or the average accuracy is not high although the model recognizes many defect types. SUMMARY
[0005] To solve the above problems, the application develops an improved YOLOX model to improve the effectiveness of the YOLOX model in actual production. The application reduces the loss of channel features of the original attention mechanism on the basis of the ShuffleNetV2-plus-YOLOX model, thereby better improving the recognition of small target defects in the die casting product defect picture, and thus improving the overall detection and recognition effect. According to the principle of geometric transformation and image processing, the application expands the data set and automatically generates a data labeling program module, thereby saving the collection cost of the data set and accelerating the automation process of model training. The application mainly recognizes and detects the defects such as water lines, blisters, shrinkage, discoloration, mechanical strain, deformation, cracks, flash, multiple flesh and mucosal strain of the die casting product.
[0006] The application is implemented by at least one of the following technical solutions.
[0007] A die casting product defect detection method based on an improved YOLOX model, comprising the following steps:
[0008] Collecting picture data of the die casting product and pre-processing the picture data;
[0009] Constructing a die casting product defect detection model, wherein the die casting product defect detection model adopts a ShuffleNetV2-plus structure as a backbone network to improve the defect detection capability, and adopts an ECA attention mechanism to reduce the information loss problem of the dimension reduction channel;
[0010] Training the die casting product defect detection model and detecting using the die casting product defect detection model, taking mAP and FPS on the same computer as the final evaluation result;
[0011] Using the trained die casting product defect detection model to detect and classify the die casting product defects.
[0012] Further, the pre-processing comprises the following steps:
[0013] Step S1, dividing the picture data into a training set, a validation set and a test set;
[0014] Step S2, enhancing the obtained data set by geometric transformation methods such as folding, rotating, scaling and translating, and generating new data labels;
[0015] Step S3, using image processing operations such as adding salt and pepper noise, Gaussian noise and grayscale transformation on the picture data to simulate the actual production environment, and further expanding the data set using step S2.
[0016] Further, the die casting product defect detection model adopts an improved YOLOX model, including an improved backbone network, a feature fusion part and a decoupling head.
[0017] The improved backbone network includes an improved ShuffleNetV2-plus network, which separates and extracts features from a detection data set.
[0018] The feature fusion part includes a feature pyramid network structure and a pyramid attention structure, which can pass information from different directions and improve the ability of the model to fuse features.
[0019] The decoupling head classifies and analyzes the extracted features and improves the convergence speed.
[0020] Further, the improved ShuffleNetV2-plus network introduces an ECA attention mechanism instead of the original SE attention mechanism in the ShuffleNetV2-plus structure and Xception structure.
[0021] Further, the Xception structure separates the channel convolution and the point-wise convolution, and then performs channel shuffling at the end. The order of the channel convolution and the point-wise convolution can be exchanged.
[0022] Further, the loss function of the improved YOLOX model includes a positioning loss, a confidence prediction loss and a prediction loss.
[0023] Further, the positioning loss uses an IOU loss function; the confidence prediction loss and the prediction loss use a BCEWithLogitsLoss loss function.
[0024] The BCEWithLogitsLoss function is a combination of the BCE loss function and the Sigmoid activation function, which can calculate the loss of multi-classification problems.
[0025] Further, the decoupling head of the improved YOLOX model adopts the strategies of Anchor-free, Multi-positives and SIMOTA.
[0026] The Anchor-free strategy takes the center of the die casting product defect sample as a positive sample and specifies the FPN level for each sample.
[0027] The Multi-positives strategy takes the fixed size area where the center is located as a positive sample; the SIMOTA strategy distributes multiple defect target boxes to the same number of positive samples.
[0028] Furthermore, the image data collected in step S1 is processed by the improved YOLOX model and then uniformly sized into RGB images of 640×640 pixels.
[0029] Furthermore, the die-casting defect detection model is used to identify defects in die-casting products including watermarks, blistering, shrinkage cavities, discoloration, mechanical scratches, deformation, cracks, flash, excess material, and mucosal scratches.
[0030] Compared with existing technologies, the beneficial effects of the present invention are as follows:
[0031] This invention addresses the problem of insufficient variety of defects in die-casting parts. It collects various die-casting defect data in real-world scenarios and automatically expands the dataset through geometric transformation and image processing. This solves the overfitting problem caused by insufficient datasets and the high labor intensity and tediousness of image data annotation, saving labor costs during training. Furthermore, it improves the deep learning YOLOX model by replacing the YOLOX backbone network with ShuffleNetV2-plus and introducing an ECA attention mechanism into the traditional ShuffleNetV2-plus structure. This enhances the model's sensitivity to channel features and improves the overall detection accuracy, providing feasible support for the deep integration of artificial intelligence technology with traditional manufacturing. Attached Figure Description
[0032] Figure 1 Here is a flowchart of a defect detection method for die-cast products based on an improved YOLOX model, as an example.
[0033] Figure 2 This is a diagram of the improved ShuffleNet module structure from the embodiment.
[0034] Figure 3 This is a structural diagram of the improved Xception module in the embodiment. Detailed Implementation
[0035] To enable those skilled in the art to better understand the present invention, preferred embodiments are illustrated in the accompanying drawings. The drawings supplement the textual description with graphics, allowing for a direct and visual understanding of each technical feature and the overall technical solution of the invention; however, they should not be construed as limiting the scope of protection of the invention. Clearly, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0036] like Figure 1 The method for detecting defects in die-cast products based on the improved YOLOX model, as shown, includes the following steps:
[0037] Step S1, collect picture data of die casting product defects through a camera, or obtain picture data of various defects of die casting products through a web crawler to form a data set;
[0038] Step S2, divide the data set into a training set, a validation set and a test set according to a ratio of 81:9:10;
[0039] Step S3, perform data augmentation on the labeled data set through geometric transformation methods such as folding, rotating, scaling and translating, and simultaneously generate new augmented data labels;
[0040] Step S4, simulate the actual production environment by using image processing operations such as adding salt and pepper noise, Gaussian noise and grayscale transformation on the picture data, and further expand the data set by repeating step S3;
[0041] Step S5, construct a die casting product defect detection model (improved YOLOX model): use ShuffleNetV2-plus network as the backbone network of the model, instead of the original backbone network DarkNet53, to improve the defect detection capability, and use ECA (Efficient Channel Attention) attention mechanism instead of SE (Squeeze-and-Excitation) attention mechanism in ShuffleNetV2-plus structure, thereby forming ShuffleNetV2-plus-YOLOX model, to reduce the information loss problem of dimension reduction channel;
[0042] Step S6, train the die casting product defect detection model;
[0043] Step S7, use the trained die casting product defect detection model to detect the test set, and use mAP (mean Average Precision) and FPS (Frames Per Second) on the same computer as the final evaluation result;
[0044] Step S8, use the trained die casting product defect detection model to detect and classify die casting product defects.
[0045] The model construction, model training and acquisition method related to the die casting product defect detection method of the present application will be described in detail below.
[0046] First, collect die casting product defect picture data, take pictures of defective die casting product parts through a camera (the size of the obtained original pictures may be different), and perform data augmentation on the obtained picture data to obtain the overall die casting product defect data set.
[0047] As a preferred embodiment, the specific processing of data enhancement is as follows:
[0048] (1) The acquired 42 original pictures are manually labeled using LabelImage labeling software, and saved as XML files with the same name as the picture files in the general PASCAL VOC format, which records the necessary information required in defect detection management, such as the category of defects and the coordinate values of the two endpoints of the minimum bounding rectangle of the defects in the picture, etc.
[0049] (2) Add salt and pepper noise, Gaussian noise, and grayscale transformation to each original picture, and generate XML files with the same name as the original picture files, with the picture file named in English as "original picture name_operation".
[0050] (3) Fold each picture horizontally and vertically. At the same time, convert the original picture annotation rectangle coordinates according to the horizontal and vertical folding formulas, and generate new XML files, with the naming format of "original picture name_[Horizontally]" and "original picture name_[Vertically]" respectively.
[0051] (4) Rotate each picture with 0° as the starting angle and 350° as the final rotation angle, with a step of 10°, to expand the data set in this way; at the same time, the rectangle coordinates in the original picture annotation file are also rotated and transformed, and new XML files are generated, with the naming format of "original picture name_rotation angle".
[0052] (5) Perform equal scale processing on the length and width of each picture, such as enlarging by 2 times, etc., and convert the original picture annotation rectangle coordinates according to the equal scale formula, and generate new XML files, with the naming format of "original picture name_[2]" etc.
[0053] (6) Perform unequal scale processing on the length and width of each picture, scaling to a specified pixel size, such as 200x200, 300x300, 250x200, 200x250, 250x300, 300x250, 350x300, 300x350, 350x400, 400x350. At the same time, convert the original picture annotation rectangle coordinates according to the unequal scale formula, and generate new XML files, with the naming format of "original picture name_[200x200]" etc.
[0054] (7) For each picture, 4 direction translations are performed, such as right-down translation (50px, 50px), right-up translation (50px, -50px), left-down translation (-50px, 50px), and left-up translation (-50px, -50px); at the same time, the defect rectangular frame coordinates marked on the original picture are converted according to the translation formula, and a new XML file is generated, and the XML file naming format is like "original picture name_[50, 50]", "original picture name_[50, -50]", "original picture name_[-50, 50]", and "original picture name_[-50, -50]".
[0055] As a preferred embodiment, the improved YOLOX model is as follows:
[0056] The improved YOLOX model comprises an improved backbone network, a feature fusion network and a decoupling head; the improved backbone network mainly extracts features of the die casting product defect dataset; the feature fusion network fuses features in different directions using a feature pyramid network structure (FPN) and a pyramid attention structure (PAN) to improve the feature fusion capability of the model; and the decoupling head classifies and regresses the extracted features.
[0057] As shown in Figure 2 , Figure 3 , the improved backbone network is mainly composed of an improved ShuffleNetV2-plus network and an Xception basic function block, and is subdivided into four different size structure units, namely ShuffleNet3x3, ShuffleNet5x5, ShuffleNet7x7 and Xception, and the corresponding array indexes of the four units are 0, 1, 2 and 3. The improved backbone network as a whole contains four stages. In the four stages, there are 4, 4, 8 and 4 structure units (these units are different arrangements of ShuffleNet3x3, ShuffleNet5x5, ShuffleNet7x7 and Xception four units, and one of the four modules can be repeatedly selected), and the four units (which can be repeatedly selected) are arranged according to experience to form a structure composed of 20 units described above, which are numbered 4, 4, 8 and 4 units from front to back, which are divided into four different stages), a total of 20 structure units. In the first stage, the stride (note: step length) is 2; in the remaining stages, the stride is 1; and the RE LU activation function is used in the previous stage, and the HS activation function is used in the remaining stages, the ECA attention mechanism is not enabled in the first two stages, and the attention mechanism is enabled in the last two stages.
[0058] The loss function of the improved YOLOX model includes a positioning loss, a confidence prediction loss and a prediction loss; the positioning loss uses an IOU (Intersection over Union) loss function; the confidence prediction loss and the prediction loss use a BCEWithLogitsLoss (Binary Cross Entropy With Logits Loss) loss function.
[0059] The BCEWithLogitsLoss function is a combination of a BCE (Binary Cross Entropy Loss) loss function and a Sigmoid activation function, and can perform loss calculation for a multi-classification problem.
[0060] In the feature fusion network, a Feature Pyramid Networks (FPN) uses an up-sampling method to fuse semantic features transmitted from top to bottom with lower layer features; a Pyramid Attention Network (PAN) uses a down-sampling method to fuse position features transmitted from bottom to top with upper layer features; the feature fusion network of the embodiment combines the two methods to enhance the feature fusion capability of the model.
[0061] The decoupling head of the improved YOLOX model adopts the strategies of Anchor-free (without prior frame), Multi-positives (multiple positive samples) and SIMOTA. Anchor-free takes the center of the die casting product defect sample as a positive sample, and specifies the FPN level for each sample; the Multi-positives strategy divides a region centered on the picture data as a positive sample; SIMOTA distributes multiple defect target frames to the same number of positive samples. Using the Anchor-free method, the prediction position can be reduced from 3 to 1, the prediction parameter value is reduced, and the prediction speed is accelerated, but at the same time, it also leads to the single selection of positive samples, which may cause the number of positive samples to be less than that of negative samples. Therefore, the embodiment also uses the Multi-positives method to expand the selection of the center point from a single point to a 3x3 neighborhood grid near the center point, and divides the region into positive samples, which makes up for the shortcomings of the Anchor-free method and improves the accuracy of model detection.
[0062] The optimal transport assignment algorithm (SIMOTA) is an algorithm that converts the label assignment process into an optimal transport problem. The simplified optimal transport assignment (SIMOTA) algorithm replaces the original Sinkhorn-Knopp algorithm in OTA with a TOP-K strategy that dynamically assigns K samples (K is set to 10 in this embodiment, but is written as K for general versatility, and the K samples are automatically assigned by the SIMOTA algorithm based on the maximum overlap between the predicted box and the GT) that are adaptive to the GT (ground truth, classification accuracy of the training set of supervised learning). The training time of the model is reduced to 75% of the original, and the detection accuracy of the improved YOLOX model is improved.
[0063] The improved YOLOX model processes the results by first reducing the dimension of the original 512-channel feature map to 256 through a 1x1 convolution, and then forming the regression and classification branches. The regression and classification calculations are performed through two 3x3 convolutions in parallel branches, and an additional IOU branch is added to the regression branch. Finally, the results are fused and summarized. The three prediction results obtained by each feature layer are as follows:
[0064] (1) Reg_output(h, w, 4): The IOU loss between the predicted box and the real box is calculated by feature point extraction, so as to predict the position information of the target box. The four parameters are x, y, w, and h, where x and y are the center coordinates of the predicted box, and w and h are the width and height of the predicted box. In this example, the size of this result is h x w x 4.
[0065] (2) Obj_output(h, w, 1): The cross-entropy loss is calculated by the positive and negative samples (the positive and negative samples are internal operations of the model, when the model prediction box and GT box overlap is greater than the set threshold, this example is 0.5) and the prediction result of whether the feature point contains an object, and is used to determine whether there is a corresponding defect in the predicted box. In this example, the size of this result is h x w x 1.
[0066] (3) Cls_output(h, w, num_classes): The cross-entropy is calculated by the class of the real box and the class of the feature point, and the defect class in the box is determined. In this example, the size of this result is h x w x 21.
[0067] After the fusion of the above three prediction results, the result obtained by each feature layer is Output(h, w, 4+1+num_classses), wherein 4+1+num_classses represents the dimension of the vector obtained after the final fusion. The first four parameters of the vector obtained through the final fusion are the position information of the target box; the fifth parameter judges whether the target defect exists in the target box; and the num_classes parameter represents the number of target defect categories contained in the target box.
[0068] As another preferred embodiment, the basic structure of the ShuffleNetV2-Plus network is as follows:
[0069] The ShuffleNetV2-Plus network is composed of two structures: ShuffleNetV2 and Xception. The ShuffleNetV2 is composed of two structure units, as shown in FIG. 2. Figure 2 When the stride is 1, the structure unit only has the right main branch. The image features first pass through channel splitting, and then pass through two groups of 1x1 point convolution and 3x3 channel convolution. The convolved feature maps are filtered through the ECANet attention mechanism and then mixed using channel shuffle to mix the results obtained by the two convolutions, as shown in FIG. 2(a). Figure 2 When the stride is 2, the basic structure unit adds a left branch. The left branch is composed of a 3x3 channel convolution and a 1x1 point convolution. Except that it does not contain channel separation, the remaining part is the same as the structure unit when the stride is 1, as shown in FIG. 2(b). Figure 2
[0070] The Xception is also composed of two structure units, as shown in FIG. 3. Figure 3 When the stride is 1, the structure unit only has the right main branch. The image features first pass through channel splitting, and then pass through three groups of 3x3 channel convolution and 1x1 point convolution. The convolved features are filtered through the ECANet attention mechanism and then mixed using channel shuffle to mix the results obtained by the two convolutions, as shown in FIG. 3(a). Figure 3 When the stride is 2, the basic structure unit adds a left branch. The left branch is composed of a group of 3x3 channel convolution and 1x1 point convolution. Except that it does not contain channel separation, the remaining part is the same as the structure unit when the stride is 1, as shown in FIG. 3(b). Figure 3
[0071] The preferred embodiments of the application disclosed above are only to facilitate the elucidation of the application. The preferred embodiments do not describe all the details of the application and limit the application to the specific embodiments described. Obviously, many modifications and variations can be made in light of the teachings above. The description is chosen and described in order to best explain the principles of the application and its practical application to thereby enable others skilled in the art to best utilize the application and get the best results from the application. The application is only limited by the claims and their full scope and equivalents.
Claims
1. A die casting product defect detection method based on an improved YOLOX model, characterized by, The method comprises the following steps: Collecting picture data of the die-casting product, and preprocessing the picture data; A die-casting product defect detection model is constructed, which uses a ShuffleNetV2-plus structure as a backbone network to improve the defect detection capability, and uses an ECA attention mechanism to reduce the information loss problem of dimension reduction channels; The die-casting product defect detection model uses an improved YOLOX model, which includes an improved backbone network, a feature fusion part and a decoupling head; The improved backbone network includes an improved ShuffleNetV2-plus network, which separates and extracts features from the detection data set; The feature fusion part includes a feature pyramid network structure and a pyramid attention structure, which can transmit information from different directions and improve the ability of the model to fuse features; The decoupling head classifies and analyzes the extracted features and improves the convergence speed; The improved ShuffleNetV2-plus network introduces an ECA attention mechanism into the ShuffleNetV2-plus structure and the Xception structure instead of the original SE attention mechanism; The Xception structure separates the channel convolution and the point-wise convolution, and then mixes the channels at the end, and the order of the channel convolution and the point-wise convolution can be exchanged; The decoupling head of the improved YOLOX model uses the Anchor-free, Multi-positives and SIMOTA strategies; The Anchor-free strategy takes the center of the die-casting product defect sample as a positive sample, and specifies the FPN level for each sample; The Multi-positives strategy takes the fixed size region where the center is located as a positive sample; the SIMOTA strategy distributes multiple defect target boxes to the same number of positive samples; The die-casting product defect detection model is trained and detected using the die-casting product defect detection model, and the mAP and the FPS on the same computer are used as the final evaluation results; The trained die-casting product defect detection model is used to detect and classify die-casting product defects.
2. The die casting product defect detection method based on the improved YOLOX model according to claim 1, wherein, The preprocessing comprises the following steps: Step S1, dividing the picture data into a training set, a validation set and a test set; Step S2, enhancing the data set with annotations by using folding, rotating, scaling and translating geometric transformation methods, and generating new data annotations; Step S3, using image processing operations such as adding salt and pepper noise, Gaussian noise and grayscale transformation on the picture data to simulate the actual production environment, and further expanding the data set using step S2.
3. The die casting product defect detection method based on the improved YOLOX model according to claim 1, characterized in that, The loss function of the improved YOLOX model includes a positioning loss, a confidence prediction loss and a prediction loss.
4. The die casting product defect detection method based on the improved YOLOX model according to claim 3, characterized in that, The positioning loss uses an IOU loss function; the confidence prediction loss and the prediction loss use a BCEWithLogitsLoss loss function; The BCEWithLogitsLoss function is a combination of the BCE loss function and the Sigmoid activation function, which can calculate the loss of multi-classification problems.
5. The die casting product defect detection method based on the improved YOLOX model according to claim 2, characterized in that, The picture data collected in step S1 is processed by the improved YOLOX model and unified into an RGB image with a size of 640*640 pixels.
6. The die casting product defect detection method based on the improved YOLOX model according to claim 1, wherein, The die casting product defect detection model is used to identify defects of the die casting product, including water marks, blistering, shrinkage, discoloration, mechanical tearing, deformation, cracking, flash, thick meat, and sticking film tearing.
Citation Information
Patent Citations
Ancient character and font recognition method based on improved YOLO v3
CN111126404A
Magnetic resonance image analysis method for brain
CN115359012A