Penicillin bottle body defect detection method based on improved YOLOv8

The improved YOLOv8 model with SSPA and MMFF modules addresses the challenge of low precision in detecting glass bottle body defects by enhancing feature extraction and fusion, resulting in higher accuracy and efficiency for vial defect detection.

CN120318607AActive Publication Date: 2025-07-15XIANGTAN UNIV

Patent Information

Application Number
CN202510804742.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The existing YOLOv8 model has low accuracy in detecting defects in glass bottles of Cylin bottles, making it difficult to effectively detect complex defects such as scratches, dirt and abnormal loading.

Method used

The spatially separable pooling attention module SSPA and feature enhancement module MMFF are introduced into the backbone of the YOLOv8 model. Through improved feature extraction and fusion methods, the capture ability of subtle and complex defect features is enhanced.

Benefits of technology

The accuracy of detection of defects in the body of the cillin bottle is improved, especially for detection of dirt and volume abnormalities in complex backgrounds, achieving higher detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318607A_ABST
    Figure CN120318607A_ABST
Patent Text Reader

Abstract

The invention discloses an improved YOLOv8-based penicillin bottle body defect detection method, and belongs to the technical field of computer vision, and the method comprises the following steps: obtaining and preprocessing penicillin bottle body defect data to obtain a training set and a test set; establishing a defect target detection model of the improved YOLOv8; training a defect target detection model by adopting the training set, optimizing a loss function, and updating a model weight parameter until the loss function is converged; and testing the defect target detection model by adopting the test set. According to the invention, a space separable pooling attention module SSPA is introduced, so that the perception capability of the network on subtle features such as bottle body scratches can be enhanced; a feature enhancement module MMFF is introduced, convolution kernels of different sizes are adopted to extract defect features of different sizes, such as smudginess, scratches and loading quantity anomaly, under a complex background, and the feature fusion effect is improved by module stacking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a detection method for defects on the body of ampoules based on improved YOLOv8. Background Art

[0002] With the rapid development of the medical field, the requirements for defect detection of medical products are increasing day by day, and defect detection methods based on computer vision are gradually occupying the mainstream position. Traditional object detection algorithms are based on manually designed feature extractors and machine learning algorithms. By using artificially defined features and combining classifiers to detect and classify objects, this process requires a large amount of professional knowledge and experience, and has a high complexity. In contrast, object detection methods based on deep learning perform better in production applications. Deep learning neural networks learn object features from raw images through end-to-end training, complete detection tasks in the corresponding background, reduce manual intervention, and improve the detection accuracy and efficiency in complex environments.

[0003] As a glass pharmaceutical container, ampoules are common medical packaging products. The appearance integrity of ampoules and the internal conditions of the bottle body directly or indirectly affect the quality and safety of the pharmaceuticals in the ampoules, which is related to the life and health of patients. Ensuring the product quality and safety of pharmaceuticals is of utmost importance. During the process of filling pharmaceuticals into the bottles, various problems may occur, such as foreign matter residues, bottle body breakage, incomplete sealing, and pharmaceutical re-dissolution. These pose challenges to the high-efficiency and accurate quality inspection of ampoules.

[0004] The production line inspection process for pharmaceuticals filled in ampoules is divided into four parts: plastic top cover detection, aluminum cap detection, top of the pharmaceutical detection, and glass bottle body detection. Timely and accurately detecting and screening defective products during this process can effectively reduce the waste of production resources and lower economic costs. The practical application of the YOLO series of models can achieve a balance between detection speed and detection accuracy, and is more suitable for real-time detection on the ampoule production line. The YOLOv8-n basic model can achieve high-precision detection of the plastic top cover, aluminum cap, and top of the pharmaceutical of the ampoule. However, the defect conditions of the glass bottle body of the ampoule are complex: the scratches on the bottle body are not obvious; the sizes of the dirt are different and the distribution is discrete; the defect features of abnormal filling amounts have irregular shapes. The existence of these phenomena increases the difficulty of detection, resulting in low overall detection accuracy of the basic network for the glass bottle body. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a detection method for defects on the body of ampoules based on improved YOLOv8 with a simple algorithm and high detection accuracy.

[0006] The technical solution of the present invention to solve the above technical problems is: a detection method for defects on the body of ampoules based on improved YOLOv8, comprising the following steps: S1: Obtain defect data of the ampoule body and perform preprocessing to obtain a training set and a test set; S2: Establish a defect target detection model based on improved YOLOv8; Taking YOLOv8 as the basic framework, introduce the spatial separable pooling attention module SSPA in the shallow layer of the backbone network Backbone; introduce the feature enhancement module MMFF in the neck Neck, replace the C2f module, use convolutional kernels of different sizes to extract features of defects and stack them to improve the feature fusion effect; the detection head Head outputs defect targets; S3: Use the training set to train the defect target detection model, optimize the loss function, and update the model weight parameters until the loss function converges; S4: Use the test set to test the defect target detection model.

[0007] In the above detection method for defects on the body of ampoules based on improved YOLOv8, the specific process of step S1 is as follows: S11: The ampoule enters the image acquisition area through a conveyor belt. During the process of the ampoule being sent for inspection at a constant speed, a sensor is triggered, and an industrial camera samples the ampoule images from multiple angles; S12: Perform target area positioning, image segmentation, defect annotation, and data augmentation on the collected picture samples to complete image preprocessing, and divide the training set and the test set according to a ratio.

[0008] In the above detection method for defects on the body of ampoules based on improved YOLOv8, in step S12, the specific processes of target area positioning, image segmentation, and data augmentation are as follows: Target area positioning: Perform binarization on the collected original picture to display the outline of the bottle body with a white background; after calculating the centroid of the white background area, find the positions of the bottle wall and the bottle bottom along the gray level changes in the horizontal and vertical directions; make a rectangular frame that is centrosymmetric about the centroid to locate the target area; Image segmentation: Fine-tune and scale the rectangular frame, segment the image and remove the background area to obtain the area of the bottle body to be inspected, which is convenient for detecting defects existing in the detection area; batch process to obtain the samples to be inspected; Data augmentation: For the processed samples to be inspected, perform data augmentation operations to expand the data volume.

[0009] The above-mentioned method for detecting defects on the vial body based on improved YOLOv8. In step S2, the spatial separable pooling attention module SSPA includes two parts: attention weight calculation and information aggregation. Among them, the attention weight calculation consists of three parts: local feature extraction, global feature generation, and spatial separation features; information aggregation involves merging the inter-group information of different feature groups after adjusting the attention weights of multiple groups of features. If the given input feature , , represents the real number field, respectively represent the three dimensions of the number of channels, height, and width. Divide into groups, denoted as , where represents the th group of features in , / / represents integer division operation; each group of features undergoes feature extraction and fusion through three parts of calculations, and finally the features of each group are aggregated.

[0010] In the above-mentioned method for detecting defects on the vial body based on improved YOLOv8, the process of attention weight calculation in step S2 includes: Local feature extraction: For each group of input features, deploy , , -sized pooling windows for global average pooling in the horizontal direction, and deploy , , -sized pooling windows for global average pooling in the vertical direction; Deploying the pooling window in a single direction embeds the spatial position information into the channel dimension; Different-sized pooling windows balance the short-distance dependence relationships when establishing long-distance dependence relationships in the discrete area, and capture multi-scale features under different distance dependence relationships, obtaining two local feature maps as follows: ; ; Among them, , respectively represent the splicing operations along the horizontal and vertical directions, respectively represent the horizontal and vertical coordinates of the pixel, is the size of the global average pooling window, , ; Global feature generation: Use a standard 1×1 convolution to process and , respectively obtaining feature maps and the feature map , and perform a matrix multiplication operation to generate a global feature map : ; ; ; wherein, represents a 1×1 convolution operation, represents a matrix multiplication operation; Apply a 1×1 standard convolution to the global feature map for information fusion, and pass through the Sigmoid function to obtain the global feature attention weight : Spatial separation feature: Upsample and to the original feature map size to obtain the feature map and the feature map , and perform a transpose between the dimension and the dimension; ; ; wherein, represents a transpose operation, represents an upsampling operation; and are concatenated in the dimension direction to obtain the feature map of , and then pass through a 3×3 standard convolution for feature extraction to obtain the feature enhancement map , realizing local feature recombination across spatial dimensions and enabling complementary transmission of feature information in the horizontal and vertical directions; ; wherein, represents a 3×3 standard convolution operation, is the merging function, represents the 3rd dimension; Pass through a 1×1 standard convolution and generate a new weight by the Sigmoid function; Separate the new weight and both along the H dimension to obtain the local feature weight in the horizontal direction, the local feature map , Local feature weight in the vertical direction , Local feature map in the vertical direction ; ; ; ; ; Among them, is the separation operation, is the Sigmoid function; Multiply the local feature map and the local feature weight correspondingly, and perform independent separation and adjustment of the features in different directions in the spatial dimension. After passing through the sigmoid function, generate the attention weight ; ; Among them, represents the element-wise multiplication operation.

[0011] In the above-mentioned detection method for the defects of the ampoule bottle body based on the improved YOLOv8, in step S2, the process of information aggregation is as follows: For the already generated and , multiply with , in sequence, perform spatial weighting to obtain the enhanced feature map as follows: ; Among them, the enhanced feature map ; Similarly, process each group of features in groups. After obtaining groups of enhanced features, perform information aggregation, and restore the size to to obtain the output ; ; Among them, is the batch normalization function; is the merging function, which is used to merge the g groups of features in the channel dimension.

[0012] In the above-mentioned detection method for the defects of the ampoule bottle body based on the improved YOLOv8, in step S2, the feature enhancement module MMFF uses convolution kernels of different sizes to extract features from the defects and stack them to improve the feature fusion effect, specifically as follows: The input of MMFF is , and the output is , , MMFF captures spatial feature information through two parallel branches and performs feature fusion. The first branch contains a convolutional layer Conv1×1 - batch normalization layer BN - activation function layer SiLU. The convolutional layer Conv1×1 - batch normalization layer BN - activation function layer SiLU is denoted as CBS, which performs weighted summation on all channels at the same position to fuse channel features. The input of the first branch generates an output ; The second branch contains a multi - scale depth - separable convolution module MDSConv and N re - parameterized convolution modules RepConv - Block. It uses convolution kernels of different sizes to capture defect features, helping the network focus on local details and context information. At the same time, stacking multiple RepConv - Blocks enhances the representation ability of features. The input of the second branch generates an output ; and After completing the element - wise addition operation, the final output is obtained through an additional CBS, and the number of channels remains unchanged.

[0013] In the above - mentioned method for detecting defects on the vial body based on improved YOLOv8, in step S2, RepConv - Block uses a multi - branch convolutional layer during the training phase, and then combines multiple computational modules into one during the inference phase, re - parameterizing the parameters of the branches onto the main branch; RepConv - Block extracts and enhances target features by repeatedly applying convolutional layers and activation functions; MDSConv includes depth - wise convolutions of 3×3, 5×5, 7×7 and a point - wise convolution of 1×1. The input feature map undergoes parallel multi - scale depth - wise convolutions, and grouped convolutions are performed in the channel dimension to extract spatial features of different sizes; after the multi - branch element - wise addition operation, feature aggregation is achieved by point - wise convolution.

[0014] In the above - mentioned method for detecting defects on the vial body based on improved YOLOv8, in step S3, the training set images are input into the defect target detection model. After feature extraction by the backbone network Backbone and feature fusion by the neck Neck, the detection head Head outputs the prediction boxes of defect targets; The error function between the target prediction result and the ground truth label is composed of the bounding box regression loss , the object class loss and the confidence loss weighted as follows: ; where is the weighting coefficient of the bounding box regression loss, is the weighting coefficient of the object class loss, is the weighting coefficient for confidence loss; Bounding box regression loss is used to measure the position and shape of the predicted box and the ground truth box, as shown in the following formula: ; ; ; Where: represents the intersection over union, is the predicted box, is the ground truth box, is the center point of the predicted box, is the center point of the ground truth box; is and the Euclidean distance; is the diagonal length of the smallest rectangle enclosing the predicted box and the ground truth; is the difference in the width-to-height ratio between the predicted box and the ground truth box; is the adjustment factor; The target class loss function is as follows: ; Where, is the total number of classes; is the class index; is the true class label; is the predicted class probability; Confidence loss is as follows: ; Where, is the true confidence, is the confidence value predicted by the network; Optimization , the model parameters are iteratively updated to make the model parameters converge to the minimum value and the model performance reach the optimal; After the training is completed, select the weight with the highest accuracy to obtain the optimal model for detecting the defects of the ampoule bottle body.

[0015] For the above-mentioned method for detecting the defects of the ampoule bottle body based on the improved YOLOv8, in step S4, the performance of the defect target detection model for detecting the defects of the ampoule bottle body is evaluated through evaluation metrics, and the evaluation metrics include , , Precision , Recall , represents the average precision of a single class; is the average value of all classes ; The calculation is as follows: ; ; ; ; Among them, represents the number of correctly detected targets, which is a true positive; represents the number of misdetected targets, which is a false positive; represents the number of real targets that are not detected, which is a false negative; represents precision , represents recall ; represents that the recall value is , represents the precision when the recall is , is the total number of categories; is the category index, represents the th value of the category.

[0016] The beneficial effects of the present invention are as follows: The present invention can solve the problems of difficult detection and low precision of partial defects of ampoule bottles, specifically manifested as: 1. Introduce the spatially separable pooling attention module SSPA, which can enhance the network's perception ability for subtle features such as bottle scratches; 2. Introduce the feature enhancement module MMFF, use convolutional kernels of different sizes to extract defect features of different sizes such as dirt, scratches, and abnormal filling volume under complex backgrounds, and stack the modules to improve the feature fusion effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is the overall flow chart of the present invention.

[0018] Figure 2 is the overall framework diagram of the defect target detection model of the present invention.

[0019] Figure 3 is the ampoule bottle defect sample diagram of the present invention, where the left, middle, and right diagrams are defect sample diagrams of dirt, scratch, and abnormal filling volume respectively.

[0020] Figure 4 is the working principle diagram of the spatially separable pooling attention module SSPA of the present invention.

[0021] Figure 5Schematic diagram of the feature enhancement module MMFF of the present invention.

[0022] Figure 6 Schematic diagram of the RepConv-Block of the present invention.

[0023] Figure 7 Schematic diagram of the MDSConv of the present invention. Detailed implementation manners

[0024] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0025] As Figure 1 shown, a detection method for defects on the body of ampoules based on improved YOLOv8 includes the following steps: S1: Obtain the defect data of the ampoule body and perform preprocessing to obtain a training set and a test set.

[0026] The specific process of step S1 is as follows: S11: The ampoule enters the image acquisition area through the conveyor belt. During the process of the ampoule being sent for inspection at a constant speed, the sensor is triggered, and the industrial camera samples the ampoule images from multiple angles, and 1241 original images are collected at a resolution of 2448×2048; S12: Perform target area localization, image segmentation, defect annotation, and data augmentation on the collected picture samples to complete image preprocessing, and divide the training set and the test set according to a ratio.

[0027] The specific processes of target area localization, image segmentation, and data augmentation are as follows: Target area localization: Perform binarization on the collected original pictures to display the bottle body contour with a white background; after calculating the centroid of the white background area, find the positions of the bottle wall and the bottle bottom along the gray-scale changes in the horizontal and vertical directions; make a rectangular frame that is centrosymmetric about the centroid to locate the target area; Image segmentation: In the actual production line, the position of the ampoule may be slightly offset. Therefore, the rectangular frame is fine-tuned and scaled, the image is segmented, and the background area is removed to obtain the area of the bottle body to be inspected, which is convenient for detecting possible dirt, scratches, and abnormal filling; batch processing is performed to obtain the samples to be inspected; Data augmentation: For the processed sample pictures, data augmentation operations are performed, and the data volume is expanded to 3723 through shear transformation Shear (±10° horizontally, ±10° vertically), adding noise Noise (1.05%), and hue transformation Hue (between -15° and +15°).

[0028] As Figure 3As shown, the defect categories of the vial body include three types, namely stain, scratch, and load anomaly.

[0029] S2: Establish a defect target detection model that improves YOLOv8, such as Figure 2 shown; Taking YOLOv8 as the basic framework, introduce the Spatial Separable Pooling Attention Module (SSPA) in the shallow layer of the backbone network Backbone to enhance the network's perception ability of subtle features such as scratches on the vial body; introduce the Feature Enhancement Module (MMFF) in the neck Neck, replace the C2f module, and use convolutional kernels of different sizes to extract features of defects such as stains, scratches, and load anomalies and stack them to improve the feature fusion effect; the detection head Head outputs defect targets.

[0030] Backbone contains the CBS (Conv-BN-SiLU, Convolutional Layer-Batch Normalization Layer-SiLU Activation Function Layer) module, the Feature Enhancement Module C2f, the Spatial Separable Pooling Attention Module SSPA, and the Spatial Pyramid Pooling Module SPPF. Backbone receives the original image, and the CBS module and the C2f module are stacked alternately to gradually extract high-order semantic features from the image; feature enhancement is achieved through SSPA; finally, the SPPF module aggregates multi-scale context information in the deep network to generate a multi-scale feature map.

[0031] The CBS module performs preliminary feature extraction through a standard 3×3 convolutional block, and then passes through the batch normalization layer BN and the SiLU activation function to accelerate training convergence and improve computational efficiency.

[0032] The C2f module is divided into three parts: segmentation, processing, and splicing, to retain shallow details and deep semantics. The C2f module receives the feature map output from the CBS module as input, and evenly divides the input along the channel dimension into two parts, and the two parts are connected by a residual connection; the first part is directly passed, and the second part is processed through the stacked Feature Fusion Bottleneck structure Bottleneck; the processed feature maps of the two parts are concatenated along the channel dimension and then passed through a 1×1 convolutional module to adjust the number of channels to obtain the final output feature map. The Bottleneck structure therein contains a 1×1 convolutional module, a 3×3 convolutional module, and a residual connection. The 1×1 convolutional module compresses the number of channels of the feature map to reduce the amount of computation; the 3×3 convolutional module performs spatial feature extraction; the residual connection adds the input feature to the extracted feature to avoid gradient disappearance.

[0033] such as Figure 4As shown, in the step S2, the spatial separable pooling attention module SSPA includes two parts: attention weight calculation and information aggregation. Among them, the attention weight calculation consists of three parts: local feature extraction, global feature generation, and spatial separable features; information aggregation involves merging the inter-group information of different feature groups after adjusting the attention weights for multiple groups of features. If the given input feature , , represents the real number field, represent the three dimensions of the number of channels, height, and width respectively. Divide into groups, denoted as , where represents the th group of features in , / / represents the integer division operation; each group of features undergoes feature extraction and fusion through three parts of calculations. During this process, the channel dimension remains unchanged while reducing the computational complexity. Finally, the features of each group are aggregated to enhance the network's ability to capture various defect features such as dirt, scratches, and abnormal filling amounts.

[0034] The process of attention weight calculation includes: Local feature extraction: For each group of input features, deploy pooling windows of sizes , , in the horizontal direction for global average pooling, and deploy pooling windows of sizes , , in the vertical direction for global average pooling; Deploying the pooling window in a single direction embeds the spatial location information into the channel dimension; Different sizes of pooling windows balance the short-distance dependence relationships when establishing long-distance dependence relationships in discrete regions, enabling the capture of multi-scale features under different distance dependence relationships, resulting in two local feature maps as follows: ; ; Among them, , represent the concatenation operations along the horizontal and vertical directions respectively, represent the horizontal and vertical coordinates of the pixel respectively, is the size of the global average pooling window, , ; Global feature generation: Use a standard 1×1 convolution to process and , respectively obtaining feature maps and the feature map , and perform a matrix multiplication operation to generate a global feature map : ; ; ; wherein, represents a 1×1 convolution operation, represents a matrix multiplication operation; and two feature maps, each pixel representing local refined features in the current direction. The two perform a matrix multiplication operation to perform cross-dimensional information interaction on local features in different spatial dimensions, and better establish the correlation of global feature information.

[0035] Apply a 1×1 standard convolution to the global feature map for information fusion, and obtain the global feature attention weight through the sigmoid function , the global feature attention weight is expressed as: ; Spatial separation features: Upsample and to the original feature map size to obtain the feature map and the feature map , respectively, and perform on dimension and transpose between dimensions; ; ; wherein, represents a transpose operation, represents an upsampling operation; and are concatenated in the dimension direction to obtain the feature map , and then feature extraction is performed through a 3×3 standard convolution to obtain a feature enhancement map , realizing the recombination of local features across spatial dimensions, and enabling the complementary transmission of feature information in the horizontal and vertical directions; ; wherein, represents a 3×3 standard convolution operation, is a merging function, represents the third dimension; Pass through a 1×1 standard convolution, and generate new weights by the sigmoid function; the new weights and both perform a separation operation along the H dimension to obtain the local feature weights in the horizontal direction and the local feature map in the horizontal direction and the local feature weights in the vertical direction and the local feature map in the vertical direction ; ; ; ; ; Among them, is the separation operation, is the Sigmoid function; Multiply the local feature map and the local feature weights correspondingly, perform independent separation adjustment of the features in different directions in the spatial dimension, and generate attention weights through the sigmoid function ; ; Among them, represents element-wise multiplication operation.

[0036] The process of information aggregation is as follows: For the already generated and , multiply with , in turn for spatial weighting to obtain the enhanced feature map as follows: ; Among them, the enhanced feature map ; Similarly, process each group of features in groups. After obtaining groups of enhanced features, perform information aggregation and restore the size to to obtain the output ; ; Among them, is the batch normalization function; is the merging function, which is used to merge g groups of features in the channel dimension.

[0037] The SPPF module is located at the end of the backbone network and is mainly used for multi-scale feature fusion. The SPPF module first compresses the number of channels of the input feature map using 1×1 convolution; then performs max pooling through three cascaded pooling windows of size 5×5; concatenates the original feature map and the results of the three poolings along the channel dimension, and restores the number of channels through 1×1 convolution.

[0038] The Neck adopts an FPN-PAN bidirectional structure, which includes an upsampling module , a 3×3 convolution module, a C2f module, a Concat splicing module, and a feature enhancement module MMFF. The Neck receives the output of the backbone network as input. First, it passes through the upsampling module in the Feature Pyramid Network (FPN) to enlarge the size of the feature map, and the Concat splicing module performs feature splicing and enhances the features through MMFF; secondly, it passes through the 3×3 convolution module of the Path Aggregation Network (PAN) to reduce the size of the feature map, and the Concat splicing module performs feature splicing and the C2f module fuses features of different scales to enhance semantic and spatial information.

[0039] FPN structure: a top-down transfer path. Upsamples the deep feature map and transfers it to the shallow feature map, performs feature map splicing in the channel dimension, and enhances the features through MMFF.

[0040] The feature enhancement module MMFF uses convolution kernels of different sizes to extract features of defects and stacks them to improve the feature fusion effect, as follows: As Figure 5 shown, the input of MMFF is , and the output is , , MMFF captures spatial feature information through two parallel branches and performs feature fusion. The first branch includes a convolutional layer Conv1×1 - batch normalization layer BN - activation function layer SiLU. The convolutional layer Conv1×1 - batch normalization layer BN - activation function layer SiLU is denoted as CBS, which performs weighted summation on all channels at the same position to fuse channel features. The input of the first branch is to generate the output ; the second branch includes a multi-scale depthwise separable convolution module MDSConv and N reparameterized convolution modules RepConv - Block, which use convolution kernels of different sizes to capture defect features, help the network focus on local details and larger - scale context information, and at the same time stack multiple RepConv - Blocks to enhance the feature representation ability and extract more complex and high - level features. The input of the second branch is to generate the output ; and After completing the element-wise addition operation, the final output is obtained through an additional CBS. The number of channels remains unchanged. CBS consists of a 1×1 standard convolution, BN, and the activation function SiLU.

[0041] As Figure 6 shown, the RepConv-Block uses a multi-branch convolutional layer during the training phase, and then combines multiple computational modules into one during the inference phase, reparameterizing the parameters of the branches onto the main branch, thereby reducing the computational amount and memory consumption and improving the inference speed. The RepConv-Block extracts and enhances the target features by repeatedly applying convolutional layers and activation functions.

[0042] As Figure 7 shown, MDSConv includes depthwise convolutions of 3×3, 5×5, and 7×7 and a pointwise convolution of 1×1. The input feature map undergoes parallel multi-scale depthwise convolutions, and grouped convolutions are performed in the channel dimension to extract spatial features of different sizes, reducing the computational amount; the residual connection helps the information flow of the deep network, avoids the vanishing gradient, and enhances the feature expression ability; after the multi-branch undergoes element-wise addition operation, feature aggregation is achieved by pointwise convolution.

[0043] PAN structure: The bottom-up transmission path. The shallow feature map is downsampled by a 3×3 convolution module with stride = 2 to transmit shallow information to the deep feature map, and they are concatenated in the channel dimension and feature fusion is performed by the C2f module. The Neck applies a bidirectional path structure, combines feature maps of different levels, avoids the problem of missing spatial details in deep features and insufficient semantic information in shallow features, and enhances the effect of multi-scale feature fusion.

[0044] Head: It contains a decoupled head and an Anchor-Free mechanism. The decoupled head separates the classification task and the regression task into two independent branches, avoiding interference between tasks and improving the classification accuracy and the stability of regression localization. The Anchor-Free mechanism is adopted, which consists of three detection layers. Feature maps of different scales are used to detect target objects of different sizes. Each detection layer outputs corresponding vectors through three branches of classification, confidence, and regression, and they are screened through non-maximum suppression (NMS). Finally, the predicted bounding box and category of the target in the original image are generated and marked. The Head is mainly responsible for converting the multi-scale feature maps output by the Neck into target detection results. The Head outputs a vector that contains the class probability, object confidence, and bounding box information of the target.

[0045] S3: The defect target detection model is trained using the training set, the loss function is optimized, and the model weight parameters are updated until the loss function converges.

[0046] In the present invention, since the dataset of ampoule bottle body defects is custom - processed and constructed without pre - trained weights, the network needs to be trained from scratch. The Stochastic Gradient Descent (SGD) optimizer is used for 100 epochs, the initial learning rate is 0.01, the momentum is set to 0.937, and the weight decay is 0.0005. The warm - up method is used for training, the warm - up period is 3 epochs, the input image size is set to 640×640, and the batch size is 16.

[0047] In step S3, the training set images are input into the defect target detection model. After feature extraction by the backbone network Backbone and feature fusion by the neck Neck, the prediction boxes of the defect targets are output by the detection head Head. The error function between the target prediction result and the true label is composed of the bounding box regression loss , the target class loss and the confidence loss weighted as follows: ; where is the weighting coefficient of the bounding box regression loss, is the weighting coefficient of the target class loss, is the weighting coefficient of the confidence loss; The bounding box regression loss is used to measure the position and shape of the prediction box and the true box, as follows: ; ; ; where: represents the intersection - over - union ratio, is the prediction box, is the true box, is the center point of the prediction box, is the center point of the true box; is and 's Euclidean distance; is the diagonal length of the smallest rectangle enclosing the prediction box and the true one; is the difference in the width - height ratio between the prediction box and the true box; is the adjustment factor; The target class loss function is as follows: ; where is the total number of classes; is the class index; is the true class label; is the predicted class probability; Confidence loss is as follows: ; where, is the true confidence, is the confidence value predicted by the network; Optimize , iteratively update the model parameters to converge the model parameters to the minimum value and achieve the optimal model performance; After 100 rounds of training are completed, select the weight with the highest accuracy to obtain the optimal model for detecting defects on the vial body.

[0048] S4: Use the test set to test the defect target detection model.

[0049] Evaluate the performance of the defect target detection model for detecting vial body defects through evaluation metrics. The evaluation metrics include , , Precision , Recall , represents the average precision of a single class; is the average value of all classes ; The calculation is as follows: ; ; ; ; where, represents the number of correctly detected targets, which is a true positive; represents the number of misdetected targets, which is a false positive; represents the number of true targets that are not detected, which is a false negative; represents Precision , represents Recall ; represents that the recall value is , represents that the recall is when the precision is is the total number of classes; is the class index, represents the th class value.

[0050] The present invention has been trained and tested, and at the same time, the existing advanced object detection models are compared with the present invention. For various types of defects in the dataset and the overall @50 As shown in Table 1, Stain, Scratch, and Load represent the defect categories of dirt, scratch, and abnormal filling volume respectively; @50 is a specific form of and the predicted bounding boxes in the calculation

[0051]

[0052] Among them, Size represents the size of the input image.

[0053] It can be seen from Table 1 that the detection accuracy of the present invention for various types of defects on the vial body has been improved compared with the basic network, and the overall detection effect is also better than that of the existing advanced object detection network.

Claims

1. An inspection method for the defects of ampoule bottle bodies based on improved YOLOv8, characterized in that, The steps include the following: S1: Obtain the defect data of the vial body and perform preprocessing to obtain a training set and a test set; S2: Establish a defect target detection model by improving YOLOv8; Taking YOLOv8 as the basic framework, introduce the Spatial Separable Pooling Attention Module (SSPA) into the shallow layer of the backbone network Backbone; introduce the Feature Enhancement Module (MMFF) into the Neck, replace the C2f module, use convolutional kernels of different sizes to extract features of the defects and stack them to enhance the feature fusion effect; the detection head Head outputs the defect targets; S3: Use the training set to train the defect target detection model, optimize the loss function, and update the model weight parameters until the loss function converges; S4: Use the test set to test the defect target detection model.

2. The vial body defect detection method based on the improved YOLOv8 according to claim 1, characterized in that, The specific process of step S1 is as follows: S11: The vial enters the image acquisition area through the conveyor belt. During the process of the vial being sent for inspection at a constant speed, the sensor is triggered, and the industrial camera samples the vial images from multiple angles; S12: Perform target area localization, image segmentation, defect annotation, and data augmentation on the collected picture samples to complete image preprocessing, and divide the training set and the test set according to a certain proportion.

3. The method for detecting defects on the body of ampoules based on improved YOLOv8 according to claim 2, characterized in that, In step S12, the specific processes of target area localization, image segmentation, and data augmentation are as follows: Target area localization: Perform binary operation on the collected original picture to display the vial body contour with a white background; after calculating the centroid of the white background area, find the positions of the bottle wall and the bottle bottom along the gray-scale changes in the horizontal and vertical directions; make a rectangle box that is centrosymmetric about the centroid to locate the target area; Image segmentation: Fine-tune and scale the rectangle box, segment the image and remove the background area to obtain the area of the vial to be inspected, which is convenient for detecting the defects existing in the detection area; batch process to obtain the samples to be inspected; Data augmentation: For the processed samples to be inspected, perform data augmentation operations to expand the data volume.

4. The method for detecting defects on the vial body based on the improved YOLOv8 according to claim 1, wherein In step S2, the Spatial Separable Pooling Attention Module (SSPA) includes two parts: attention weight calculation and information aggregation. Among them, the attention weight calculation consists of three parts: local feature extraction, global feature generation, and spatial separable features; information aggregation includes the inter-group information merging of different feature groups after the multi-group features are adjusted by the attention weights; If the given input features , , represent the real number field, represent the number of channels, height, and width dimensions respectively, and divide into groups, denoted as , where represents the th group of features in , / / represents the integer division operation; Each group of features undergoes three parts of calculation for feature extraction and fusion, and finally the features of each group are aggregated.

5. The method for detecting defects on the vial body based on the improved YOLOv8 according to claim 4, wherein, In step S2, the process of attention weight calculation includes: Local feature extraction: For each set of input features, deploy , , -sized pooling windows for global average pooling in the horizontal direction, and deploy , , -sized pooling windows for global average pooling in the vertical direction; The deployment of the pooling window in a single direction embeds the spatial position information into the channel dimension; Pooling windows of different sizes balance the short-distance dependencies when establishing long-distance dependencies in discrete regions, enabling the capture of multi-scale features under different distance dependencies, resulting in two local feature maps as follows: ; ; Among them, and respectively represent the splicing operations along the horizontal and vertical directions, respectively represent the horizontal and vertical coordinates of the pixel, is the size of the global average pooling window, , ; Global feature generation: Process using a standard 1×1 convolution and , respectively obtaining feature maps and feature map , and Perform a matrix multiplication operation between them to generate the global feature map : ; ; ; Among them, represents a 1×1 convolution operation, represents matrix multiplication operation; For the global feature map Apply a standard 1×1 convolution for information fusion and pass through the Sigmoid function to obtain the global feature attention weights : Spatial separation feature: Upsample and to the size of the original feature map to obtain feature maps and feature map respectively, and perform a transpose between the dimensions and the dimensions of ; ; ; Among them, represents a transpose operation, represents an upsampling operation; And In Concatenate in the dimensional direction to obtain The feature map of, and then perform feature extraction through a standard 3×3 convolution to obtain the feature enhancement map , realizing local feature recombination across spatial dimensions and enabling complementary transmission of feature information in the horizontal and vertical directions; ; Among them, represents a 3×3 standard convolution operation, is the merging function, represents the third dimension; After a 1×1 standard convolution, new weights are generated by the Sigmoid function; the new weights and are both subjected to a separation operation along the H dimension to obtain the local feature weights in the horizontal direction , the local feature map in the horizontal direction , the local feature weights in the vertical direction , and the local feature map in the vertical direction ; ; ; ; ; Among them, is a separation operation, is a Sigmoid function; Multiply the local feature map and the local feature weights correspondingly, independently separate and adjust the features in different directions in the spatial dimension, and generate attention weights through the sigmoid function ; ; Among them, represents an element-wise multiplication operation.

6. The method for detecting defects on the body of ampoules based on improved YOLOv8 according to claim 5, wherein, In step S2, the process of information aggregation is: For what has been generated and , multiply with , sequentially to perform spatial weighting to obtain an enhanced feature map as follows: ; Among them, the enhanced feature map ; Similarly, process each group of features in the group to obtain a group of enhanced features, then perform information aggregation and restore the size to ; ; Among them, is the batch normalization function; is the merging function, which is used to merge g groups of features in the channel dimension.

7. The vial body defect detection method based on the improved YOLOv8 according to claim 6, wherein, In step S2, the Feature Enhancement Module (MMFF) uses convolutional kernels of different sizes to extract features of the defects and stack them to enhance the feature fusion effect, specifically as follows: The MMFF input is , and the output is , . MMFF captures spatial feature information through two parallel branches and performs feature fusion. The first branch contains a convolutional layer Conv1×1 - batch normalization layer BN - activation function layer SiLU. The convolutional layer Conv1×1 - batch normalization layer BN - activation function layer SiLU is denoted as CBS, which performs weighted summation on all channels at the same position to fuse channel features. The input of the first branch is to generate the output ; The second branch contains a multi - scale depth - separable convolution module MDSConv and N re - parameterized convolution modules RepConv - Block, which use convolution kernels of different sizes to capture defect features, helping the network focus on local details and context information. At the same time, stacking multiple RepConv - Blocks enhances the representation ability of features. The input of the second branch is to generate the output ; and After completing the element - wise addition operation, the final output is obtained through an additional CBS, and the number of channels remains unchanged.

8. The method for detecting defects on the vial body based on the improved YOLOv8 according to claim 7, characterized in that, In the step S2, the RepConv-Block uses a multi-branch convolutional layer during the training phase, and then combines multiple computational modules into one during the inference phase, reparameterizing the parameters of the branches onto the main branch; the RepConv-Block extracts and enhances the target features by repeatedly applying convolutional layers and activation functions; the MDSConv includes depthwise convolutions of 3×3, 5×5, and 7×7 and a pointwise convolution of 1×1. The input feature map undergoes parallel multi-scale depthwise convolutions, and grouped convolutions are performed in the channel dimension to extract spatial features of different sizes; after the multi-branches undergo element-wise addition operations, feature aggregation is achieved by pointwise convolution.

9. The method for detecting defects on the body of ampoules based on improved YOLOv8 according to claim 7, wherein, In the step S3, the training set images are input into the defect target detection model. After feature extraction by the backbone network Backbone and feature fusion by the neck Neck, the detection head Head outputs the prediction boxes of the defect targets. Error function between the target prediction result and the true label composed by bounding box regression loss , target class loss and confidence loss weighted as follows: ; Among them, is the weighted coefficient of the bounding box regression loss, is the weighted coefficient of the target class loss, is the weighted coefficient of the confidence loss; Bounding box regression loss Used to measure the position and shape of the predicted box and the ground truth box, as shown in the following formula: ; ; ; Wherein: represents the intersection over union, is the predicted bounding box, is the ground truth bounding box, is the center point of the predicted bounding box, is the center point of the ground truth bounding box; is and 's Euclidean distance; is the diagonal length of the minimum rectangle enclosing the predicted and ground truth bounding boxes; is the difference in aspect ratio between the predicted and ground truth bounding boxes; is the adjustment factor; Target class loss function As follows: ; wherein, is the total number of categories; is the category index; is the true category label; is the predicted category probability; Confidence loss As shown in the following formula: ; Among them, is the true confidence level, is the confidence value predicted by the network; Optimization , the model parameters are iteratively updated to converge the model parameters to the minimum value, and the model performance reaches the optimal state; After the training is completed, the weights with the highest accuracy are selected to obtain the optimal model for detecting defects on the vial body.

10. The vial body defect detection method based on improved YOLOv8 according to claim 1, characterized in that In the step S4, the performance of the defect target detection model for detecting the defects of the vial body is evaluated by evaluation indicators, and the evaluation indicators include , , precision , recall , represents the average precision of a single class; is the average value of all classes , and the calculation is as follows: ; ; ; ; Among them, represents the number of correctly detected targets, which is the true positive; represents the number of misdetected targets, which is the false positive; represents the number of real targets not detected, which is the false negative; represents the precision , represents the recall rate ; represents that the recall rate takes the value of , represents that the recall rate is when the precision is is the total number of categories; is the category index, represents the th value of the category.

Citation Information

Patent Citations

  • YOLOv5m-based transparent plastic bottle surface defect detection method

    CN116452565A

  • Group behavior identification method and device based on separation time-space relationship, and medium

    CN118262404A

  • High body seriola quinqueradiata detection method based on YOLOv8 network structure

    CN118279935A

Cited By

  • MSFE-YOLO-based poultry hatching egg shell defect detection method and system

    CN120558976A

  • Oil seal defect intelligent detection method based on improved YOLOv12

    CN120707566A

  • Improved yolov12-based intelligent detection method for oil seal defects

    CN120707566B

  • Bottle body packaging defect detection system based on machine vision

    CN121073965A

  • Container damage detection method based on improved UPAD-YOLO

    CN121074053A