Hydrogen cylinder outer surface defect detection method based on improved YOLOX
By improving the YOLOX model, combining CSPNet and SE modules, PANet integrates the feature pyramid, solving the problem of low accuracy of surface defect detection of hydrogen cylinders by traditional detection methods, and achieving more efficient and reliable defect detection.
Patent Information
- Application Number
- CN202510078001.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional detection methods have low accuracy in detecting surface defects of hydrogen cylinders, and rely on professional equipment and technicians, making it difficult to meet modern needs.
Using the hydrogen cylinder external surface defect detection method based on improved YOLOX, image features are extracted using CSPNet in the backbone part of YOLOX, combined with the SE module to enhance the feature importance weight, the Neck part uses PANet to integrate the feature pyramid and perform target position and category prediction in the Head part.
It effectively improves the accuracy of target detection, improves the efficiency and reliability of surface defect detection of hydrogen cylinders, and reduces manual intervention and costs.
Smart Images

Figure CN119941695A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of hydrogen cylinder outer surface defect detection, and in particular to a hydrogen cylinder outer surface defect detection method based on improved YOLOX. Background Art
[0002] As the global demand for environmental protection and clean energy increases, hydrogen energy, as a clean energy with great potential, has attracted widespread attention. As an important equipment for hydrogen energy storage and transportation, the safety of automotive hydrogen cylinders is of vital importance. During use, automotive hydrogen cylinders may have internal and external surface defects such as cracks, corrosion, and wear. If they are not discovered in time, they may cause hydrogen leakage or explosion, posing a threat to the safety of personnel and property. Therefore, it is crucial to accurately and quickly detect defects in hydrogen cylinders. Traditional detection methods such as visual inspection, ultrasonic inspection, and eddy current inspection have problems such as low efficiency, insufficient accuracy, and high requirements for personnel experience. Visual inspection is greatly affected by the state of workers and has low detection efficiency; ultrasonic and eddy current inspections require professional equipment and technicians, and have limited detection effects on minor defects. With the popularization of hydrogen-powered vehicles, traditional detection methods can no longer meet modern needs, and more advanced and efficient detection technologies are urgently needed. Deep learning technology provides a new solution. By analyzing images of the surface of hydrogen cylinders, it can automatically identify and detect various defects and improve detection accuracy and efficiency. The deep learning model has good adaptability and scalability, can handle different types of hydrogen cylinders, and continuously improves performance with data accumulation and model improvement. Combining deep learning with existing inspection equipment can realize automated inspection, reduce manual intervention, reduce costs, and improve inspection reliability and consistency. Therefore, it is of great significance to study intelligent and efficient hydrogen cylinder defect detection algorithms. Summary of the invention
[0003] Aiming at the problem that traditional detection methods have low detection accuracy for surface defects of hydrogen cylinders, the present invention proposes a method for detecting outer surface defects of hydrogen cylinders based on improved YOLOX, which effectively improves the accuracy of target detection.
[0004] The present invention adopts the following technical scheme: a method for detecting defects on the outer surface of a hydrogen cylinder based on an improved YOLOX, comprising the following steps: the backbone part of YOLOX uses CSPNet as the backbone network to extract image features; the SE module can enhance the features important to the current task and suppress irrelevant or minor features by weighting each channel; the Neck part is responsible for integrating the feature pyramid and adopts PANet to enhance the transmission of information; the output part of the Head is used to predict the position and category of the target; the improvement of YOLOX is that the SE module adds a global maximum pooling branch on the basis of global average pooling, and the combination of these two pooling operations enables the model to simultaneously obtain the average global information and the most significant local information; the pooled output is mapped to the corresponding weight through a 1×1 convolution, and then the standard fully connected layer in the SE module is replaced by a 1×1 convolution operation.
[0005] The above-mentioned hydrogen cylinder outer surface defect detection method based on improved YOLOX first uses the data enhancement technology Mosaic or MixUp to enhance the input hydrogen cylinder image data to improve the robustness of the model.
[0006] The above-mentioned method for detecting defects on the outer surface of hydrogen cylinders based on improved YOLOX, Mosaic is a method of splicing four pictures into a new image, which can increase the diversity during training while retaining the contextual information between different targets; MixUp is a method of generating a new image by linearly interpolating two images and their labels; this method can effectively smooth the decision boundary and reduce overfitting.
[0007] The above-mentioned hydrogen cylinder outer surface defect detection method based on improved YOLOX, CSPNet effectively alleviates the gradient vanishing problem and enhances the diversity of features by separating feature maps and connecting them across stages; it includes convolutional layers, batch normalization layers and activation functions to improve the nonlinear expression ability of the network.
[0008] In the above-mentioned method for detecting defects on the outer surface of hydrogen cylinders based on improved YOLOX, after the backbone part extracts features at different network layers, the neck part fuses these features to ensure that the model can utilize both high-level semantic information and low-level detail information. PANet improves the fluidity of features through the path aggregation mechanism, so that information can be better transmitted from low layers to high layers, and vice versa.
[0009] In the above-mentioned method for detecting defects on the outer surface of hydrogen cylinders based on improved YOLOX, the hydrogen cylinder image data is obtained by photographing the surface defects of the hydrogen cylinder using an industrial camera; and the labels are annotated through labelimg, and the surface defects of the hydrogen cylinder are divided into wear, gouges, scratches, and native defects; a total of 950 images containing four types of defects are collected, and the data set is expanded to 8100 through random area cropping, mirror inversion, brightness transformation and image rotation operations; the number of training sets, test sets, and validation sets are 6561, 810, and 729, respectively.
[0010] The method of the present invention first uses new data enhancement technologies, such as Mosaic and MixUp, to improve the robustness of the model; the backbone part uses CSPNet as the backbone network to extract image features; the SE module combines the advantages of global average pooling and global maximum pooling, so that the model can capture richer and more expressive descriptors; the Neck part is responsible for integrating the feature pyramid, using PANet to enhance the transmission of information, and finally outputs the predicted position and category of the target; the present invention effectively improves the accuracy of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 This is a structural diagram of the network model of the present invention.
[0012] Figure 2 This is a structural diagram of the improved SE attention of the present invention.
[0013] Figure 3 It is the detection result diagram of the present invention. DETAILED DESCRIPTION
[0014] A method for detecting defects on the outer surface of a hydrogen cylinder based on improved YOLOX includes the following steps: first, new data enhancement techniques, such as Mosaic and MixUp, are used to improve the robustness of the model; the backbone part uses CSPNet (Cross Stage Partial Network) as the backbone network to extract image features; the SE module can enhance the features that are important to the current task and suppress irrelevant or minor features by weighting each channel; the Neck part is responsible for integrating the feature pyramid and adopts PANet (Path Aggregation Network) to enhance the transmission of information; the output part of the Head is used to predict the location and category of the target.
[0015] At the input of this network, the MixUp method is used to achieve data enhancement: first, one image is padded left and right to scale it to the standard input size (640×640); then the second image is padded up and down and scaled to the same size; then, the two images are weighted superimposed according to a certain fusion coefficient. During the training process, the MixUp method is enabled by default, and the fusion coefficient is set between 0.5 and 1.5 to generate the final MixUp image. The Mosaic method is also used to achieve data enhancement. The Mosaic method selects four images and performs geometric transformations on each image, such as superposition, cropping, and scaling, and then splices them according to four regions, resets the label file coordinates, and generates a new enhanced image. In the YOLOX algorithm, the Mosaic method is enabled by default, and the coefficient is set between 0.1 and 2.0 to enrich the input data set. However, the Mosaic operation is not enabled throughout the process. In the last 15 cycles of training, this function is turned off by default to effectively prevent the trained model from deviating in real-world scene detection.
[0016] YOLOX uses CSPDarknet-53 as its backbone network for feature extraction, solving the gradient vanishing and exploding problems. CSPDarknet-53 is an improved version based on the Darknet-53 backbone network. CSPDarknet draws on the design concept of CSPNet, which is designed to solve the problem of gradient information duplication in traditional convolutional neural networks, reduce network computing bottlenecks and improve gradient performance. Ultimately, CSPNet effectively reduces the model size by reducing the number of model parameters and floating-point operations by half without reducing the inference speed and accuracy.
[0017] The SE module is used to fuse different types of global information and enhance the feature extraction ability and robustness of the target detection model. A global maximum pooling branch is added to the Squeeze part to combine the advantages of global average pooling and global maximum pooling, so that the model can capture richer and more expressive descriptors. This combination enables the network to better adapt to the size, shape and texture changes of the target. The improved method actually introduces a richer feature selection strategy, which not only pays attention to the average information of the overall features, but also attaches importance to the most significant local information, thereby improving the discrimination ability of the model. In addition, the output pooling layer is mapped into corresponding weights through 1×1 convolution, so that the model can adaptively adjust the weights of global average pooling and global maximum pooling according to the task and scene during training. This adaptive learning strategy helps to improve the generalization ability of the model and perform better when dealing with diverse target detection tasks. Finally, the standard fully connected layer in the SE unit is replaced by a 1×1 convolution operation, which not only effectively reduces the number of parameters of the model, but also avoids overfitting, thereby improving the adaptability of the model.
[0018] The Neck part of YOLOX is mainly responsible for feature fusion and transmission in order to detect targets at different scales. PANet, as its Neck structure, improves the detection ability of small objects, thereby improving detection accuracy; the feature extraction and fusion network of the YOLOX algorithm inherits the network structure of YOLOv5 and adopts a bottom-up feature pyramid network with position enhancement. PAFPN is a fusion structure of the feature pyramid network and the path aggregation network. The traditional FPN structure transfers the feature map to PAN through a top-down upsampling process. The PAN structure uses bottom-up downsampling to make the top-level features contain stronger position information. Finally, fusion is performed, and good detection effects are achieved for targets of different scales.
[0019] The prediction end of the algorithm of the present invention adopts a decoupled head structure. For the regression task in target detection, the shape contour features of the target, that is, the underlying position information, have a greater impact on it; while the classification task needs to be more sensitive to the feature differences between samples and needs to pay attention to the high-level semantic features of the image. The decoupled head structure of YOLOX connects two branches after a 1×1 convolution, and the 1×1 convolution adjusts the feature map channel. The two parallel branches complete the classification and regression of the target respectively. In the classification branch, a prediction score of the target box category is generated for each target, thereby realizing N binary classification judgments; in the regression branch, first determine whether it is the foreground or the background, and then perform positioning prediction based on the coordinate information. After completing the classification and regression feature predictions, the feature layers of different sizes are predicted using the decoupled head structure respectively, and finally the information of the three feature layers is integrated to obtain the final prediction result.
[0020] The present invention will be further described below in conjunction with the accompanying drawings: Reference Figure 1 In this embodiment, a hydrogen cylinder surface defect detection algorithm based on improved YOLOX includes the following steps: Step 1: First, use Mosaic and Mixup to perform data enhancement on the input image; The specific implementation method of Mosaic is: (1) Randomly select four images from the dataset.
[0021] (2) Scale these four images to the same size according to a certain ratio.
[0022] (3) Stitch them together into a new image.
[0023] (4) Update the annotation information to ensure that the location of each target is correct.
[0024] The specific implementation method of Mixup is: (1) Randomly select two images and their corresponding labels from the dataset.
[0025] (2) Use a weight factor (usually between 0 and 1) to perform a weighted combination of the two images and labels.
[0026] (3) The generated new images and their labels are used for training.
[0027] Step 2: The enhanced data is extracted through the backbone, first through the Focus module: (1) Split the input image: Split the input image into four parts.
[0028] (2) Merge features: Merge these parts and concatenate them in the channel dimension.
[0029] (3) Convolution operation: Convolution is performed on the merged features to extract high-level features.
[0030] Then, the deep features of the input image are extracted through multiple feature extraction layers constructed by convolutional layers, batch normalization layers, activation functions, and the CSPNet backbone network.
[0031] Step 3: The three layers of features extracted enter the SE module respectively to enhance the features important to the current task and suppress irrelevant or minor features.
[0032] Step 4: The next step is to enter the Neck part to fuse these features to ensure that the model can utilize both high-level semantic information and low-level detail information. Specifically, it includes: (1) Feature upsampling: Upsample the low-level feature map to the same spatial size as the high-level feature map.
[0033] (2) Feature splicing: The upsampled low-level feature map is spliced with the high-level feature map to enhance the feature representation.
[0034] (3) Convolution processing: Perform convolution operation on the concatenated feature maps to generate fused feature maps, which are then used in the final detection head.
[0035] Step 5: After the three features of different scales are fused, they enter the output head for prediction. Each output layer will contain a classifier to predict the category of the target. In addition to category prediction, each output layer will also have a regression part to predict the position and size of the bounding box. Through regression, these outputs can generate the final detection box.
[0036] Step 6: Model training parameter setting: The model training uses a Windows server with an NVIDIA RTX A4000 graphics card with a video memory of 16 GB, and the test software is PyCharm2021.1. The number of training sets, validation sets, and test sets are 6561, 810, and 729 respectively, the image size is 640×640, the total number of training epochs is 300, the learning rate is initialized to 0.001, the optimization function selects the Adam optimizer, and the network framework is pytorch2.1.
[0037] Step 7: Model training process: Combine classification loss and regression loss to ensure that the model optimizes both the accuracy of object detection and the precision of bounding boxes, and then update the network parameters through gradient back propagation. Use the test set to test the trained network model, and use evaluation indicators to evaluate the test results.
[0038] Step 8: Image detection process: Input the image to be detected into the trained model for detection. The detection results are as follows: Figure 3 As shown, the green frame detection result is wear, the purple frame detection result is gouge, and the cyan frame detection result is scratch.
Claims
1. A method for detecting defects on the outer surface of a hydrogen cylinder based on improved YOLOX, characterized in that: The following steps are involved: The backbone part of YOLOX uses CSPNet as the backbone network to extract image features of image data; the SE module can enhance the features that are important to the current task and suppress irrelevant or minor features by weighting each channel; the Neck part is responsible for integrating the feature pyramid and uses PANet to enhance the transmission of information; the output part of the Head is used to predict the location and category of the target; the improvement of YOLOX lies in that the SE module adds a global maximum pooling branch on the basis of global average pooling. The combination of these two pooling operations enables the model to simultaneously obtain the average global information and the most significant local information; the pooled output is mapped to the corresponding weight through a 1×1 convolution, and then the standard fully connected layer in the SE module is replaced by a 1×1 convolution operation.
2. According to claim 1, a method for detecting outer surface defects of hydrogen cylinders based on improved YOLOX is characterized in that: First, use the data enhancement technology Mosaic or MixUp to enhance the input hydrogen cylinder image data to improve the robustness of the model.
3. The method for detecting outer surface defects of hydrogen cylinders based on improved YOLOX according to claim 2 is characterized in that: Mosaic is a method that concatenates four images into a new image, which increases diversity during training while retaining contextual information between different targets; MixUp is a method that generates a new image by linearly interpolating two images and their labels; this method can effectively smooth decision boundaries and reduce overfitting.
4. A method for detecting outer surface defects of hydrogen cylinders based on improved YOLOX according to claim 1, 2 or 3, characterized in that: CSPNet effectively alleviates the problem of gradient vanishing and enhances the diversity of features by separating feature maps and connecting them across stages; it includes convolutional layers, batch normalization layers, and activation functions to improve the nonlinear expression ability of the network.
5. A method for detecting outer surface defects of hydrogen cylinders based on improved YOLOX according to claim 1, 2 or 3, characterized in that: After the backbone extracts features at different network layers, the neck fuses these features to ensure that the model can utilize both high-level semantic information and low-level detail information. PANet improves the fluidity of features through the path aggregation mechanism, allowing information to be better transmitted from low layers to high layers, and vice versa.
6. A method for detecting outer surface defects of hydrogen cylinders based on improved YOLOX according to claim 1, 2 or 3, characterized in that: The hydrogen cylinder image data is obtained by photographing the surface defects of hydrogen cylinders using an industrial camera; and labeled through labelimg. The surface defects of hydrogen cylinders are divided into wear, gouges, scratches, and native defects; a total of 950 images containing four types of defects are collected, and the data set is expanded to 8100 through random area cropping, mirror inversion, brightness transformation and image rotation operations; the number of training sets, test sets, and validation sets are 6561, 810, and 729, respectively.