Battery CT Image Defect Detection Method Based on Improved YOLOv8

By improving the YOLOv8 network model, combining two-dimensional Gabor filtering and dual attention mechanism, the problem of low detection accuracy of battery CT image defects is solved, automated positioning and qualitative judgment are realized, detection accuracy is improved and model parameters are reduced.

CN118587158BActive Publication Date: 2025-06-17ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410621569.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-06-17
Estimated Expiration
2044-05-20

AI Technical Summary

Technical Problem

The prior art has problems with low judgment accuracy in positioning and qualitative judgment of battery CT images, which cannot meet the needs of automated detection in industrial scenarios.

Method used

The improved YOLOv8 network model is adopted, combined with two-dimensional Gabor filtering and dual attention mechanism, and the accuracy of battery CT image defect detection is improved through feature extraction, feature fusion and lightweight convolution shuffling modules.

Benefits of technology

It realizes automatic positioning and qualitative judgment of internal defects of the battery, improves detection accuracy, and reduces the number of parameters of the model, and has a certain degree of portability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118587158B_ABST
    Figure CN118587158B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting defects in battery CT images based on improved YOLOv8. The method includes: obtaining defect images inside the battery using a CT scanning device, and dividing them into a training set, a validation set, and a test set; performing feature extraction and feature fusion on the battery CT images through two-dimensional Gabor filtering; combining a dual attention mechanism and replacing ordinary convolutional layers with a lightweight convolutional shuffle module to obtain an improved network model; inputting the training set into the network model for training to obtain a trained object detection model; and performing performance evaluation on the model through the test set to obtain the detection results of defects inside the battery. By improving the YOLOv8 network and combining the Gabor feature extraction and fusion algorithm, the present invention can effectively improve the detection accuracy of battery internal defect CT images, reduce the size of model parameters, and improve the portability of the model while accurately detecting defects inside the battery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image detection, and particularly relates to a method for detecting defects in battery CT images based on improved YOLOv8. Background Art

[0002] A lithium-ion battery is a device that converts and stores energy based on the principle of electrochemical reactions, mainly composed of four parts: a positive electrode, a negative electrode, an electrolyte, and a separator. Internal structural defects in the battery often occur at the densely stacked electrode sheets, and common internal defects include electrode fracture, internal impurities, electrode wrinkles, etc. These defects may lead to a decline in battery performance, a reduction in capacity, deterioration of charge and discharge performance, and even pose safety problems. Currently, the methods for detecting defects in new energy vehicle batteries mainly include manual inspection, physical property testing, and non-destructive testing, etc.

[0003] As one of the most common non-destructive testing methods, CT detection can obtain the internal structure image of the battery and can perform three-dimensional reconstruction of the battery, thereby accurately detecting the internal defects of the battery. Compared with manual dissection testing and physical property testing, CT detection uses the attenuation law of X-rays in materials for non-contact measurement and will not cause damage to the object under test; compared with ultrasonic testing, CT scanning can directly detect and locate the internal structural defects of the battery, with advantages such as high precision and fast speed.

[0004] CT scanning is based on the weakening and absorption characteristics of radiation in the object to be detected. It uses an X-ray source to emit X-rays with a certain energy and intensity, and scans the object to be detected from multiple directions by rotating the X-ray source at a small angle. During the process of penetrating the object to be detected, the X-rays will attenuate due to absorption and scattering effects. By capturing the intensity of the X-rays after penetrating the object through a detector, a series of two-dimensional images can be formed. Different materials inside the object to be detected in the two-dimensional images will present different gray values. The components with a larger density absorb X-rays more strongly, so they present higher gray values in the image.

[0005] However, due to the special structure of the battery and the complexity of defect formation, there are still some problems and challenges in the current CT detection methods. Currently, for the qualitative discrimination method of battery CT images, it is still limited to manual judgment after obtaining the images. Even when assisted by machine learning, most methods are still based on manual features plus classifiers. The processing of CT images only goes as far as size measurement and gray value discrimination, and it is impossible to achieve automatic analysis of internal defects.

[0006] In summary, for the detection of internal defects in batteries, CT detection is a non-destructive detection method that combines accuracy and real-time performance and can intuitively reflect the internal defects of batteries. However, for battery defect CT images, there is still a lack of a suitable method for localization and qualitative discrimination, and the judgment accuracy of existing methods is relatively low, unable to meet the automated detection requirements in industrial scenarios. Therefore, how to provide an automated method for the localization and qualitative discrimination of battery internal defect CT images, so as to improve the accuracy of battery internal defect detection, is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for detecting battery CT image defects based on improved YOLOv8 in view of the deficiencies of the prior art. The present invention extracts and fuses features of battery CT images through two-dimensional Gabor filtering; combined with a dual attention mechanism, a lightweight convolutional shuffle module is used to replace the ordinary convolutional layer, which can effectively improve the defect detection accuracy and the size of the detection model of battery CT images; the present invention can automatically detect the defect position, automatically discriminate the defect type, and accurately distinguish battery defects from electrode backgrounds.

[0008] The purpose of the present invention is achieved through the following technical solutions: A method for detecting battery CT image defects based on improved YOLOv8 includes the following steps:

[0009] (1) Use a CT scanning device to obtain CT images containing internal defects of the battery to construct an image data set, and divide it into a training set, a validation set, and a test set;

[0010] (2) Extract features of the CT images in the image data set through a two-dimensional Gabor filter, and fuse the extracted features to obtain a fused feature map corresponding to the current CT image;

[0011] (3) Improve the YOLOv8 network model by combining a dual attention mechanism module and a lightweight convolutional shuffle module to obtain an improved YOLOv8 network model;

[0012] (4) Input the training set after feature extraction and feature fusion by the two-dimensional Gabor filter into the improved YOLOv8 network model for training, and adjust the parameters of the improved YOLOv8 network model according to the CIoU loss function to obtain a trained YOLOv8 network model;

[0013] (5) Input the test set after feature extraction and feature fusion by the two-dimensional Gabor filter into the trained YOLOv8 network model, and output the types of defect images corresponding to each CT image in the test set and a txt file containing the center coordinates of the predicted box of the defect image and the length and width of the rectangular box to obtain the defect detection results of each CT image in the test set.

[0014] Further, the internal defects of the battery include electrode fracture, electrode wrinkling, internal foreign matters, and burrs.

[0015] Further, the step (2) includes the following sub-steps:

[0016] (2.1) Adjust the CT image f(x, y) in the image dataset to a grayscale image, where x and y represent the abscissa and ordinate of the CT image, respectively;

[0017] (2.2) Construct a two-dimensional Gabor filter; wherein, the two-dimensional Gabor filter includes two Gabor filter kernels with a size of K, a wavelength of λ, a phase offset of a shape aspect ratio of γ, and directions θ of parallel stripes of 0° and 90° respectively. The standard deviation of the two-dimensional Gabor filter on the time domain x-axis is σ;

[0018] (2.3) Transform the coordinates of the CT image f(x, y) adjusted to a grayscale image through a coordinate transformation equation to obtain the CT image f′(x′, y′) with adjusted coordinates, where x′ and y′ represent the abscissa and ordinate of the CT image with adjusted coordinates, respectively;

[0019] (2.4) Input the CT image f′(x′, y′) with adjusted coordinates into the two-dimensional Gabor filter, and sequentially perform step-by-step feature extraction on the CT image with adjusted coordinates using the 0° Gabor filter kernel and the 90° Gabor filter kernel to obtain a first image feature and a second image feature;

[0020] (2.5) Perform feature fusion on the first image feature and the second image feature through feature stitching to obtain a fused feature map corresponding to the current CT image.

[0021] Further, the expression of the coordinate transformation equation is:

[0022] x′ = xcosθ + ysinθ (1)

[0023] y′ = -xsinθ + ycosθ (2)

[0024] where θ represents the direction of the parallel stripes of the Gabor filter kernel;

[0025] The calculation formula of the two-dimensional Gabor filter is:

[0026]

[0027] where the coordinates x and y are transformed to obtain the coordinates x′ and y′ after coordinate transformation, λ is the wavelength of the filter, Let \(\varphi\) be the phase shift, \(\gamma\) be the shape aspect ratio, \(\sigma\) be the standard deviation of the two-dimensional Gabor filter on the x-axis in the time domain, and \(\theta\) be the direction of the parallel stripes.

[0028] Furthermore, step (3) includes the following sub-steps:

[0029] (3.1) After the last C2f layer in the backbone network of the YOLOv8 network model and before the SPPF module, add a dual attention mechanism module to obtain an improved backbone network;

[0030] (3.2) Replace all CBS modules in the feature fusion module of the YOLOv8 network model with lightweight convolutional shuffle modules to obtain an improved feature fusion module;

[0031] (3.3) Based on the improved backbone network, the improved feature fusion module, and the detection head of the YOLOv8 network model, construct an improved YOLOv8 network model.

[0032] Furthermore, the dual attention mechanism module includes a channel attention module and a spatial attention module. The dual attention mechanism module takes the input feature map \(F\in\mathbb{R}^{H\times W\times C}\), C×H×W passes it through the channel attention module to obtain the channel attention map \(M_c\in\mathbb{R}^{C}\), c \(\in\mathbb{R}^{C}\), C×1×1 then element-wise multiplies the channel attention map \(M_c\in\mathbb{R}^{C}\), c \(\in\mathbb{R}^{C}\), C×1×1 with the input feature map \(F\in\mathbb{R}^{H\times W\times C}\), C×H×W to obtain the feature map after channel processing; then inputs the feature map after channel processing into the spatial attention module to obtain the spatial attention map \(M_s\in\mathbb{R}^{H\times W}\), s \(\in\mathbb{R}^{H\times W}\), 1×H×W then element-wise multiplies the spatial attention map \(M_s\in\mathbb{R}^{H\times W}\), s \(\in\mathbb{R}^{H\times W}\), 1×H×W with the feature map after channel processing to obtain the reconstructed feature map; where \(H\) is the height, \(W\) is the width, and \(C\) is the number of channels.

[0033] Furthermore, the channel attention module includes an average pooling layer, a max pooling layer, a shared network generated by a multi-layer perceptron, and a sigmoid activation function. The input of the channel attention module is the feature map \(F\in\mathbb{R}^{H\times W\times C}\). First, it passes through the average pooling layer and the max pooling layer respectively to obtain two feature maps, then inputs the two feature maps into the shared network, and element-wise sums the two features output by the shared network to obtain a combined feature vector, and then performs an activation operation through the sigmoid activation function to obtain the channel attention map \(M_c\in\mathbb{R}^{C}\), C×H×W \(\in\mathbb{R}^{C}\); c \(\in\mathbb{R}^{C}\), C×1×1 ;

[0034] The spatial attention module includes an average pooling layer, a max pooling layer, a convolutional layer, and a sigmoid activation function. The input of the spatial attention module is the feature map after channel processing. First, it passes through the average pooling layer and the max pooling layer respectively to obtain two feature maps. Then, the two feature maps are concatenated in the channel dimension. The concatenated feature map is convolved through the convolutional layer, and finally, it undergoes an activation operation through the sigmoid activation function to obtain the spatial attention map M s ∈R 1×H×W 。

[0035] Further, the input of the lightweight convolutional shuffle module first undergoes downsampling through the convolutional layer, and then uses the depth convolutional layer to perform a convolutional operation on a single channel and then apply it to multiple channels. Secondly, the output result of the convolutional layer and the output result of the depth convolutional layer are concatenated to obtain the feature map after convolutional concatenation, and then the shuffle operation is performed to mix the features output by the dense convolutional operation into the output of the lightweight convolutional shuffle module with a small computational amount.

[0036] Further, the improved backbone network includes two CBS modules, a C2f module, a CBS module, a C2f module, a CBS module, a C2f module, a CBS module, a C2f module, a dual attention mechanism module, and an SPPF module connected in sequence; the improved feature fusion module includes an Unsample module, a C2f module, an Unsample module, a C2f module, a lightweight convolutional shuffle module, a C2f module, a lightweight convolutional shuffle module, and a C2f module.

[0037] Further, the CIoU loss function is calculated by the following formula:

[0038]

[0039] where CIoU Loss represents the value of the CIoU loss function, IoU is the intersection over union of the localization box, b is the center point of the predicted box, b gt is the center point of the ground truth box, ρ 2 (b, b gt ) represents the Euclidean distance between the center point b and the center point b gt , c represents the diagonal length of the smallest bounding rectangle that can contain both the predicted box and the ground truth box, α represents the weight of the localization box, and υ is the coefficient for measuring the consistency of the width-to-height ratio of the localization box;

[0040] The formula for calculating the intersection over union of the localization box is:

[0041]

[0042] Among them, A represents the ground truth box, B represents the predicted box, and the intersection over union (IoU) of the localization box is used to characterize the localization accuracy of the YOLOv8 network model;

[0043] The calculation formula for the coefficient measuring the consistency of the aspect ratio of the localization box is:

[0044]

[0045] where w / h is the aspect ratio of the predicted box, and w gt / h gt is the aspect ratio of the ground truth box;

[0046] The calculation formula for the weight of the localization box is:

[0047]

[0048] where IoU is the intersection over union of the ground truth box A and the predicted box B.

[0049] Compared with the prior art, the beneficial effects of the present invention are:

[0050] (1) The present invention first uses Gabor filtering to extract features from the CT images of battery internal defects, and then fuses the 0° features and 90° features obtained by feature extraction, enhancing the contrast between the defects and the electrode background and improving the detection accuracy of the model for CT images of battery internal defects; since there is no large difference in gray value between the electrode wrinkles and the electrode background, the present invention can effectively distinguish the electrode wrinkles from the electrode background through Gabor feature extraction, and the detection accuracy for the defect of electrode wrinkles is improved most significantly.

[0051] (2) The present invention adds the CBAM module to the YOLOv8 network model. The CBAM module obtains the channel correlation and spatial correlation, strengthens the features with strong correlation, and suppresses the features with weak and unimportant correlation, improving the detection accuracy of the model.

[0052] (3) In the industrial scenario, the detection of CT images of battery internal defects needs to be transplanted to specific devices for real-time detection, and the portability of the model is crucial; the present invention replaces the ordinary convolutional layer with GSConv, reducing the number of parameters while ensuring the accuracy, thereby reducing the memory occupied by the model and having a certain portability. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flowchart of the method for detecting battery CT image defects based on the improved YOLOv8 of the present invention;

[0054] Figure 2It is a flowchart of two-dimensional Gabor filter feature extraction and feature fusion in the embodiments of the present invention;

[0055] Figure 3 It is a battery defect map after two-dimensional Gabor filter feature extraction and feature fusion processing in the embodiments of the present invention;

[0056] Figure 4 It is a structural diagram of the improved YOLOv8 network model constructed in the embodiments of the present invention;

[0057] Figure 5 It is a structural diagram of the constructed dual attention mechanism module in the embodiments of the present invention;

[0058] Figure 6 It is a structural diagram of the constructed channel attention module in the embodiments of the present invention;

[0059] Figure 7 It is a structural diagram of the constructed spatial attention module in the embodiments of the present invention;

[0060] Figure 8 It is a structural diagram of the constructed lightweight convolution shuffle module in the embodiments of the present invention;

[0061] Figure 9 It is a detection effect diagram of four types of internal battery defects in the embodiments of the present invention;

[0062] Figure 10 It is a detection accuracy transformation curve graph of four types of internal battery defects when IoU takes 0.5 in the embodiments of the present invention. Detailed implementation manners

[0063] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application.

[0064] The terms used in the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0065] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination". Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus that comprises the element.

[0066] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0067] In this embodiment, the selected battery model is a ternary lithium battery, with a battery voltage of 3.7V and a capacity of 8000mAh. The resolution of the selected CT scanner is 2μm. Experiments are carried out using the deep learning framework Pytorch, with the operating system being Ubuntu20.04, the programming language being Python 3.8.10, the CPU being 7vCPU Intel(R)Xeon(R)CPU E5-2680 v4@2.40GHz, and the GPU being TITAN Xp with a GPU video memory of 12GB.

[0068] See Figure 1 , the battery CT image defect detection method based on improved YOLOv8 of the present invention specifically includes the following steps:

[0069] (1) Use a CT scanning device to obtain CT images containing internal defects of the battery to construct an image dataset, and divide it into a training set, a validation set, and a test set.

[0070] Furthermore, the internal defects of the battery include but are not limited to common internal defects of the battery such as electrode fracture, electrode fold, internal foreign matter, and burrs. The internal defects of the battery are concentrated in the electrode sheet area. For this area, a CT scanning device is used to obtain CT images containing internal defects of the battery to construct a dataset.

[0071] (2) As Figure 2As shown, the CT images in the image dataset are subjected to feature extraction through a two-dimensional Gabor filter, and the extracted features are fused to obtain a fused feature map corresponding to the current CT image.

[0072] (2.1) Adjust the CT image f(x, y) in the image dataset to a grayscale image, where x and y represent the abscissa and ordinate of the CT image respectively.

[0073] (2.2) Construct a two-dimensional Gabor filter; where the two-dimensional Gabor filter includes two Gabor filter kernels with a size of K, a wavelength of λ, a phase offset of a shape aspect ratio of γ, and parallel stripe directions θ of 0° and 90° respectively, and the standard deviation of the two-dimensional Gabor filter on the time domain axis is σ.

[0074] (2.3) Transform the coordinates of the CT image f(x, y) adjusted to a grayscale image through a coordinate transformation equation to obtain the CT image f′(x′, y′) with adjusted coordinates, where x′ and y′ represent the abscissa and ordinate of the CT image with adjusted coordinates respectively.

[0075] Furthermore, the expression of the coordinate transformation equation is:

[0076] x′ = xcosθ + ysinθ (1)

[0077] y′ = -xsinθ + ycosθ (2)

[0078] where θ represents the direction of the parallel stripes of the Gabor filter kernel.

[0079] It should be understood that the CT image f′(x′, y′) with adjusted coordinates is in the form of a grayscale image.

[0080] (2.4) Input the CT image f′(x′, y′) with adjusted coordinates into the two-dimensional Gabor filter, and sequentially perform step-by-step feature extraction on the CT image with adjusted coordinates using the 0° Gabor filter kernel and the 90° Gabor filter kernel to obtain a first image feature and a second image feature.

[0081] Furthermore, the calculation formula of the two-dimensional Gabor filter is:

[0082]

[0083] where the coordinates x and y are transformed to obtain the coordinates x′ and y′ through coordinate transformation, λ is the filter wavelength, is the phase offset, γ is the shape aspect ratio, σ is the standard deviation of the two-dimensional Gabor filter on the time domain x-axis, and θ is the direction of the parallel stripes.

[0084] It should be noted that according to the size K of the Gabor filter kernel, the range size extracted once can be determined as K*K.

[0085] (2.5) Feature fusion is performed on the first image feature and the second image feature through feature splicing to obtain a fused feature map corresponding to the current CT image. This fused feature map is an image of internal battery defects with a significant difference in distinguishability from the electrode background.

[0086] Exemplarily, after performing Gabor filter feature extraction and feature fusion on the internal battery defects in step (2), the processing results of four types of defects, namely electrode breakage, electrode wrinkle, internal foreign matter, and burr in the battery, are as Figure 3 shown. As can be seen from Figure 3 it, in the original defect image, the distinction between the defect and the background is not obvious enough, and there are certain difficulties in defect detection. By fusing the first image feature and the second image feature extracted by the 0° and 90° Gabor filter kernels through feature splicing, the darker features and the brighter features extracted are complementary, effectively suppressing the complex electrode background and enhancing the distinguishability between the defect and the background. The CT image of the battery defect obtained after two-dimensional Gabor filter feature extraction and feature fusion processing is the input for the subsequent improved YOLOv8 network model.

[0087] (3) The YOLOv8 network model is improved by combining the Convolutional Block Attention Module (CBAM) module and the Group Shuffle Convolution (GSConv) module to obtain an improved YOLOv8 network model, and its structure is as Figure 4 shown, which can improve the accuracy of the YOLOv8 network model and reduce the number of parameters of the YOLOv8 network model at the same time.

[0088] (3.1) After the last C2f layer in the backbone network of the YOLOv8 network model and before the SPPF module, a CBAM module is added to obtain an improved backbone network.

[0089] It should be noted that through the CBAM module, the channel attention and spatial attention of the feature map can be extracted for correlation, so as to strengthen important features and suppress unimportant features; then connecting an SPPF module can ensure the unity in dimension.

[0090] In this embodiment, the CBAM module includes a channel attention module and a spatial attention module. As Figure 5 shown, the CBAM module takes the input feature map F∈R C×H×WAfter passing through the channel attention module, the channel attention map M is obtained. c ∈R C×1×1 Then, the channel attention map M c ∈R C×1×1 and the input feature map F ∈ R C×H×W are multiplied element-wise to obtain the feature map after channel processing; then, the feature map after channel processing is input into the spatial attention module to obtain the spatial attention map M s ∈R 1 ×H×W Then, the spatial attention map M s ∈R 1×H×W and the feature map after channel processing are multiplied element-wise to obtain the reconstructed feature map. Among them, H is the height, W is the width, and C is the number of channels.

[0091] The calculation process of this CBAM module can be summarized as:

[0092]

[0093]

[0094] Among them, F′ represents the feature map after channel processing, and F″ represents the reconstructed feature map. denotes element-wise multiplication.

[0095] Furthermore, the channel attention module includes an average pooling layer, a max pooling layer, a shared network generated by a multilayer perceptron (MLP), and a sigmoid activation function. As Figure 6 shown, the input of the channel attention module is the feature map F ∈ R C×H×W . First, it passes through the average pooling layer and the max pooling layer respectively to obtain two feature maps. Then, the two feature maps are fed into the shared network, and the two features output by the shared network are summed element-wise to obtain a combined feature vector. Then, it undergoes an activation operation through the sigmoid activation function, and the activation size is R C / r×1×1 , where r represents the reduction rate, to obtain the channel attention map M c ∈R C×1×1 .

[0096] The calculation method of this channel attention module can be expressed as:

[0097] M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) (6)

[0098] Among them, M c(F) is the channel attention map of the feature map F, σ is the sigmoid function, MLP represents the shared network, AvgPool represents the average pooling layer, and MaxPool represents the max pooling layer.

[0099] Furthermore, the spatial attention module includes an average pooling layer, a max pooling layer, a convolutional layer, and a sigmoid activation function. As Figure 7 shown, the input of the spatial attention module is the feature map after channel processing. First, it passes through the average pooling layer and the max pooling layer respectively to obtain two feature maps. Then, the two feature maps are concatenated in the channel dimension. The concatenated feature map is convolved through the convolutional layer, and finally, it undergoes an activation operation through the sigmoid activation function to obtain the spatial attention map M s ∈R 1×H×W .

[0100] The calculation method of this spatial attention module can be expressed as:

[0101] M s (F′) = σ(f 7×7 ([AvgPool(F′)+MaxPool(F′)])) (7)

[0102] where, M s (F′) is the spatial attention map of the feature map F′ after channel processing, and f 7×7 is a 7×7 convolutional operation.

[0103] (3.2) Replace all CBS modules in the feature fusion (Neck) module of the YOLOv8 network model with GSConv modules to obtain an improved feature fusion module.

[0104] It should be noted that the role of Neck is to fuse the high-level features and low-level features processed by the Backbone features and transmit them to the detection head (Head) for detection.

[0105] In this embodiment, the structure of the GSConv module is as Figure 8 shown. The input of the GSConv module first undergoes downsampling through the convolutional (Conv) layer, then uses the depthwise convolution (DWConv) layer to perform a convolutional operation on a single channel and then apply it to multiple channels. Secondly, the output results of the convolutional layer and the depthwise convolutional layer are concatenated to obtain the feature map after convolutional concatenation, and then a shuffle operation is performed to mix the features output by the dense convolutional operation into the output of the GSConv module with a small computational amount.

[0106] It should be understood that the GSConv module effectively avoids the irrelevant consumption of extracting channel correlation in the CBS module through a depth convolution layer that only focuses on the spatial correlation between channels, thereby reducing the number of parameters in the improved YOLOv8 network model.

[0107] (3.3) Construct an improved YOLOv8 network model based on the improved backbone network, the improved feature fusion module, and the detection head of the YOLOv8 network model.

[0108] As Figure 4 shown, the improved backbone network includes two CBS modules, a C2f module, a CBS module, a C2f module, a CBS module, a C2f module, a CBS module, a C2f module, a CBAM module, and an SPPF module connected in sequence; the improved feature fusion module includes an Unsample module, a C2f module, an Unsample module, a C2f module, a GSConv module, a C2f module, a GSConv module, and a C2f module. Among them, the purpose of the Unsample module is to increase the dimension, magnify the high-level small feature map for easy fusion with the low-level large feature map, and then detect through the detection head of the YOLOv8 network model.

[0109] (4) Input the training set after two-dimensional Gabor filter feature extraction and feature fusion into the improved YOLOv8 network model for training, and adjust the parameters of the improved YOLOv8 network model according to the CIoU loss function to obtain a trained YOLOv8 network model.

[0110] Specifically, input the fused feature map corresponding to the CT image in the training set after two-dimensional Gabor filter feature extraction and feature fusion in step (2) into the improved YOLOv8 network model for iterative training, output the predicted defect image with predicted bounding box annotations and a txt file containing the center coordinates of the predicted bounding box of the defect image and the length and width of the rectangular box, and calculate the CIoU loss function based on the output result and the corresponding label (i.e., the coordinates of the ground truth box and the length and width of its corresponding rectangular box). Take minimizing the CIoU loss function value as the optimization objective, adjust the parameters of the improved YOLOv8 network model until the CIoU loss function value is less than the preset loss threshold or reaches the preset number of training epochs, and stop training to obtain a trained YOLOv8 network model.

[0111] Furthermore, the CIoU loss function is calculated by the following formula:

[0112]

[0113] where CIoU Loss represents the value of the CIoU loss function, IoU is the intersection over union of the localization box, b is the center point of the predicted box, bgt is the center point of the ground truth box, ρ 2 (b, b gt ) represents the Euclidean distance between the center point b and the center point b gt c represents the diagonal length of the smallest bounding rectangle that can contain both the predicted box and the ground truth box, α represents the weight of the localization box, and υ is the coefficient for measuring the consistency of the aspect ratio of the localization box.

[0114] Furthermore, the calculation formula for the intersection over union (IoU) of the localization box is:

[0115]

[0116] where A represents the ground truth box, B represents the predicted box, and the intersection over union IoU of the localization box is used to characterize the accuracy of the YOLOv8 network model in localization.

[0117] Furthermore, the calculation formula for the coefficient for measuring the consistency of the aspect ratio of the localization box is:

[0118]

[0119] where w / h is the aspect ratio of the predicted box, w gt / h gt is the aspect ratio of the ground truth box.

[0120] Furthermore, the calculation formula for the weight of the localization box is:

[0121]

[0122] where IoU is the intersection over union of the ground truth box A and the predicted box B.

[0123] (5) Input the test set after two-dimensional Gabor filter feature extraction and feature fusion into the trained YOLOv8 network model, and output the type of the defect image corresponding to each CT image in the test set, as well as a txt file containing the center coordinates of the predicted box of the defect image and the length and width of the rectangle box, to obtain the defect detection results of each CT image in the test set.

[0124] Specifically, inputting the test set into the final model can obtain the detection results. As Figure 9 shown, the performance of the model can thus be judged. The detected model outputs a detection image with the defect selected by the predicted box, and the type of the defect will be marked above the rectangle box. At the same time, a txt file containing the center coordinates of the predicted box of the defect and the length and width of the rectangle box will be output for each image. The detected defect is framed by a rectangle box, and the type of the defect is marked above the rectangle box, where the mark ① is the electrode fold defect, the mark ② is the electrode fracture defect, the mark ③ is the internal foreign object defect, and the mark ④ is the burr defect, as Figure 9 shown.

[0125] Furthermore, after making predictions using the test set, the performance metrics of the YOLOv8 network model can be calculated to evaluate the performance of the YOLOv8 network model. The performance metrics of the YOLOv8 network model include accuracy, recall, precision, mean average precision, and the number of parameters, and their calculation formulas are as follows:

[0126]

[0127]

[0128]

[0129]

[0130] Among them, P is the accuracy, R is the recall, TP represents the true value predicted as positive by the model, FP represents the false value predicted as positive by the model, and FN represents the false value predicted as negative by the model; AP is the precision, mAP is the mean average precision, n is the total number of samples in the test set, and i is the i-th sample in the test set.

[0131] It should be understood that the accuracy represents the proportion of true values predicted as positive by the model among all the results predicted as positive, and the recall represents the proportion of correctly predicted positive samples among all the correctly predicted samples. For example, assuming the detection of whether there is a defect such as electrode wrinkles in an image, then TP means that the image actually has electrode wrinkles and is detected, FP means that there are no electrode wrinkles in the image but the model detects it as having a defect, and FN means that there is no defect in the image and it is detected that there is no defect in the image.

[0132] Exemplarily, the YOLOv8 original model is trained using the CT image dataset of battery internal defects, and the training results are compared with those obtained by improving the YOLOv8 network model of the present invention. The comparison results are shown in Table 1.

[0133] Table 1: Comparison of training situations of different models

[0134]

[0135] It can be seen from the experimental results that the improved YOLOv8 model described in the present invention has the best detection performance for the CT images of battery defects. When the IoU is set to 0.5, the mAP reaches 98.2%, and when the IoU is set to 0.75, the mAP reaches 90.2%. Based on the original YOLOv8 model, when the IoU of the improved YOLOv8 model described in the present invention is set to 0.75, the average precision is improved by 4.1%, the number of parameters is reduced by 25,000, and the detection precision for the four types of defects is improved, with the highest precision improvement being 8.5%. Therefore, compared with the original YOLOv8 model, the improved YOLOv8 model described in the present invention has a high detection precision while having a relatively small number of parameters, verifying the superiority of the improved YOLOv8 model described in the present invention.

[0136] After the YOLOv8 deep learning algorithm is improved by the method described in the present invention, the detection precision of each defect when the IoU is 0.5 is as Figure 10 shown. The detection precision of each defect is above 96.9%. The categories of the defects are labeled through tags, and the specific categories of the defects can be effectively detected, and the coordinates of the defects are recorded in a txt file. Therefore, when the method described in the present invention is applied to the detection of the four types of defects, namely internal electrode fracture, electrode fold, internal foreign matter, and burr in the battery, the defects can be accurately framed in the picture and the categories of the defects can be given, so that the main defects inside the battery can be effectively identified.

[0137] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A battery CT image defect detection method based on improved YOLOv8, characterized in that: The following steps are involved: (1) Using CT scanning equipment to obtain CT images containing internal defects of the battery to construct an image dataset, and divide it into a training set, a validation set, and a test set; (2) extracting features from the CT image in the image data set by using a two-dimensional Gabor filter, and fusing the extracted features to obtain a fused feature map corresponding to the current CT image; the step (2) includes the following sub-steps: (2.1) Adjust the CT image f(x, y) in the image data set to a grayscale image, where x and y represent the horizontal and vertical coordinates of the CT image, respectively; (2.2) constructing a two-dimensional Gabor filter; wherein the two-dimensional Gabor filter includes a size of K, a wavelength of λ, and a phase shift of Two Gabor filter kernels with shape aspect ratio γ and parallel stripes with directions θ of 0° and 90° respectively, and the standard deviation of the two-dimensional Gabor filter in the x-axis of the time domain is σ; (2.3) The coordinates of the CT image f(x, y) adjusted to a grayscale image are transformed by using the coordinate transformation equation to obtain the CT image f′(x′) after coordinate adjustment. ′ ,y ′ ), where x ′ and ′ Respectively represent the abscissa and ordinate of the CT image after coordinate adjustment; the expression of the coordinate transformation equation is: x ′ =xcosθ+ysinθ y ′ =-xsinθ+ycosθ Among them, θ represents the direction of the parallel stripes of the Gabor filter kernel; (2.4) The CT image f′(x ′ ,y ′ ) is input into a two-dimensional Gabor filter, and a 0° Gabor filter kernel and a 90° Gabor filter kernel are used to perform step-by-step feature extraction on the coordinate-adjusted CT image in turn to obtain a first image feature and a second image feature; wherein the calculation formula of the two-dimensional Gabor filter is: Among them, the coordinates x and y are transformed to obtain the coordinate x ′ ,y ′ , λ is the filter wavelength, is the phase offset, γ is the shape aspect ratio, σ is the standard deviation of the two-dimensional Gabor filter in the time domain x-axis, and θ is the direction of the parallel stripes; (2.5) performing feature fusion on the first image feature and the second image feature by feature splicing to obtain a fusion feature map corresponding to the current CT image; (3) Combining the dual attention mechanism module with the lightweight convolution shuffle module to improve the YOLOv8 network model, obtaining an improved YOLOv8 network model; the step (3) includes the following sub-steps: (3.1) After the last C2f layer in the backbone network of the YOLOv8 network model and before the SPPF module, a dual attention mechanism module is added to obtain an improved backbone network; (3.2) All CBS modules in the feature fusion module of the YOLOv8 network model are replaced with lightweight convolutional shuffle modules to obtain an improved feature fusion module; (3.3) constructing an improved YOLOv8 network model based on the improved backbone network, the improved feature fusion module and the detection head of the YOLOv8 network model; wherein the improved backbone network includes two CBS modules, a C2f module, a CBS module, a C2f module, a CBS module, a C2f module, a CBS module, a C2f module, a dual attention mechanism module and an SPPF module connected in sequence; the improved feature fusion module includes an Unsample module, a C2f module, an Unsample module, a C2f module, a lightweight convolution shuffle module, a C2f module, a lightweight convolution shuffle module and a C2f module; (4) Input the training set after two-dimensional Gabor filter feature extraction and feature fusion into the improved YOLOv8 network model for training, and adjust the parameters of the improved YOLOv8 network model according to the CIoU loss function to obtain a trained YOLOv8 network model; wherein the CIoU loss function is calculated by the following formula: Among them, CIoU Loss represents the CIoU loss function value, IoU is the intersection over union ratio of the positioning box, b is the center point of the prediction box, and b gt is the center point of the real box, ρ 2 (b,b gt ) represents the center point b and the center point b gt The Euclidean distance between them, c represents the diagonal length of the smallest enclosing rectangle that can contain both the predicted box and the true box, α represents the weight of the positioning box, and υ is the coefficient for measuring the consistency of the aspect ratio of the positioning box; The calculation formula of the intersection-over-union ratio of the positioning frame is: Among them, A represents the real box, B represents the predicted box, and the intersection over union (IoU) of the positioning box is used to characterize the positioning accuracy of the YOLOv8 network model; The calculation formula of the coefficient of measuring the consistency of the aspect ratio of the positioning frame is: Among them, w / h is the aspect ratio of the prediction box, w gt / h gt is the aspect ratio of the real frame; The calculation formula of the weight of the positioning box is: Among them, IoU is the intersection over union ratio of the real box A and the predicted box B; (5) The test set after two-dimensional Gabor filter feature extraction and feature fusion is input into the trained YOLOv8 network model, and the type of defect image corresponding to each CT image in the test set and a txt file containing the center coordinates of the defect image prediction box and the length and width of the rectangular box are output to obtain the defect detection results of each CT image in the test set.

2. The battery CT image defect detection method based on improved YOLOv8 according to claim 1, characterized in that: The internal defects of the battery include electrode breakage, electrode wrinkles, internal foreign matter and burrs.

3. The battery CT image defect detection method based on improved YOLOv8 according to claim 1, characterized in that: The dual attention mechanism module includes a channel attention module and a spatial attention module. The dual attention mechanism module takes the input feature map F∈R C×H×W After the channel attention module, the channel attention map M is obtained c ∈R C×1×1 , and then map the channel attention to M c ∈R C×1×1 And the input feature map F∈R C×H×W Perform element-by-element multiplication to obtain the feature map after channel processing; then input the feature map after channel processing into the spatial attention module to obtain the spatial attention map M s ∈R 1×H×W , and then map the spatial attention to M s ∈R 1×H×W Multiply the feature map after channel processing element by element to obtain the reconstructed feature map; where H is the height, W is the width, and C is the number of channels.

4. The battery CT image defect detection method based on improved YOLOv8 according to claim 3 is characterized in that: The channel attention module includes an average pooling layer, a maximum pooling layer, a shared network generated by a multi-layer perceptron and a sigmoid activation function. The input of the channel attention module is the feature map F∈R C×H×W First, after passing through the average pooling layer and the maximum pooling layer, two feature maps are obtained. Then the two feature maps are passed into the shared network, and the two features output by the shared network are summed element by element to obtain the merged feature vector, which is then activated by the sigmoid activation function to obtain the channel attention map M. c ∈R C×1×1 ; The spatial attention module includes an average pooling layer, a maximum pooling layer, a convolution layer and a sigmoid activation function. The input of the spatial attention module is a feature map after channel processing. First, after passing through the average pooling layer and the maximum pooling layer, two feature maps are obtained. Then, the two feature maps are spliced ​​in the channel dimension. The spliced ​​feature map is convolved through a convolution layer, and finally activated by a sigmoid activation function to obtain a spatial attention map M. s ∈R 1×H×W .

5. The battery CT image defect detection method based on improved YOLOv8 according to claim 1, characterized in that: The input of the lightweight convolution shuffle module is first downsampled through the convolution layer, and then a deep convolution layer is used to implement a convolution operation on a single channel, which is then applied to multiple channels. Next, the output result of the convolution layer and the output result of the deep convolution layer are spliced ​​to obtain a convolution-joined feature map, and then a shuffle operation is performed to mix the features output by the dense convolution operation into the output of the lightweight convolution shuffle module with a small amount of computation.

Citation Information

Patent Citations

  • Gabor feature fused YOLOv3 aviation composite material surface defect target detection method

    CN116823741A

  • System and method for detecting surface defects of lithium battery pole piece

    CN116934762A

  • Pedestrian target detection method based on improved YOLOv8 lightweight network model

    CN117423135A

  • Light-weight road crack detection method and system capable of self-adapting to crack size

    CN117765373A