Automatic detection method for capsule defects based on infrared and visible light fusion

By fusing infrared and visible light images and utilizing a combined network of the YOLOX algorithm and spatial channel attention mechanism, we have achieved rapid and accurate detection of millimeter- and micrometer-level defects on the surface of large airship capsules. This solves the problem of high detection difficulty in existing technologies and improves detection efficiency and accuracy.

CN116091479BActive Publication Date: 2025-11-21HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310177506.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-11-21
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

Existing technologies make it difficult to quickly and effectively detect millimeter- and micrometer-level defects on the surface of large airship capsules, resulting in significant challenges in airtightness testing and impacting the airship's hovering performance and safety.

Method used

By employing an infrared and visible light image fusion method, and utilizing a joint network of the YOLOX algorithm and spatial channel attention mechanism, pores and debonding defects on the surface of the capsule are automatically labeled through image feature fusion and iterative training.

Benefits of technology

It enables rapid and accurate detection of millimeter- and micrometer-level defects on the surface of capsules, improving detection speed and accuracy, and avoiding the time-consuming and labor-intensive nature of manual inspection and potential secondary damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116091479B_ABST
    Figure CN116091479B_ABST
Patent Text Reader

Abstract

The capsule defect automatic detection method based on infrared and visible light fusion belongs to the technical field of automatic detection, and particularly relates to a capsule defect automatic detection method based on infrared and visible light fusion. The capsule defect automatic detection method based on infrared and visible light fusion is provided. The capsule defect automatic detection method based on infrared and visible light fusion comprises the following steps: step 1: infrared images and visible light images of the surface of the capsule are respectively acquired, and the pinhole defects and the debonding defects at the bonding positions of the skin and the skeleton in the images are manually marked; step 2: the images obtained in step 1 are divided into a training set, a verification set and a test set, and the visible light images and the infrared images with the same frame rate are fused by using a Laplace method; and step 3: the training set is sent into a joint network based on YOLOX and a spatial channel attention mechanism for training to obtain an algorithm model, and the weight parameters of the network are reversely adjusted and iteratively trained by using the operation results of the verification set.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of automatic detection, and particularly relates to a capsule defect automatic detection method based on infrared and visible light fusion. BACKGROUND

[0002] At the end of the 20th century, with the rapid development of modern material technology and radar electronic technology, new types of aerostats have regained vitality in the fields of low-altitude early warning detection, air defense and anti-missile guidance. Aerostats have many advantages such as long loitering time, wide field of view, high efficiency and cost ratio, and can undertake defensive tasks such as escort, coastal patrol and anti-submarine by taking advantage of their aerial superiority and endurance, so they have great development value in scientific research and military fields.

[0003] As a key material of aerostats, the defect detection of the capsule directly contributes to the safety and reliability evaluation. Flight tests show that the gas tightness of the capsule is the most important performance affecting the long loitering of the aerostat, and a small hole of millimeter or even micrometer level on the surface of the capsule will greatly affect the loitering performance of the aerostat. For a medium-sized airship, the leakage of a micro-hole with an equivalent diameter of 1mm is equivalent to the permeability of the entire airship skin reaching 5000mL / (atm.m2.d), so the micro-hole has a fatal impact on long-haul flight. The larger the capsule, the more difficult the processing, the more serious the damage, and the greater the permeability of the capsule. However, the size of hundreds of square meters on the surface of the aerostat makes it very difficult to find a millimeter-level small hole on the surface. In order to ensure the gas tightness of the capsule, comprehensive leak detection should be carried out at each stage before forming, assembling integration, transportation and flight test. The capsule gas tightness detection, especially the rapid detection and positioning technology of the gas tightness of large capsules, is one of the difficulties in the field of aerostats at present, and is also a key point affecting and restricting the leap-forward development of aerostats. However, due to the huge size of the capsule, manual soap bubble method is time-consuming and labor-intensive, and may cause secondary damage to the capsule. Moreover, due to the limitation of site conditions, full coverage of capsule leak detection cannot be achieved.

[0004] Therefore, there is an urgent need for effective full-size capsule gas tightness defect inspection means to provide protection for the gas tightness before flight and long-haul flight of the capsule. SUMMARY

[0005] The present application is aimed at the above problems, and provides a capsule defect automatic detection method based on infrared and visible light fusion.

[0006] To achieve the above-mentioned purpose, the present application adopts the following technical solutions. The present application comprises the following steps:

[0007] Step 1: respectively collect infrared images and visible light images of the surface of the capsule, and manually mark the air hole defects, debonding defects at the bonding position of the skin and skeleton in the images;

[0008] Step 2: the images obtained in step 1 are divided into a training set, a validation set and a test set, and the visible light images and infrared images of the same frame rate are fused by using the Laplace method;

[0009] Step 3: the training set is sent into the joint network based on YOLOX and the spatial channel attention mechanism for training and an algorithm model is obtained, and the operation results of the validation set are used to reversely adjust the weight parameters of the network and perform iterative training; after the iterative training, the last model obtained is selected, and after verification in the test set, the model with the highest accuracy is selected for the system;

[0010] Step 4: the fused images in the test set are input into the algorithm model obtained by training, and the algorithm model automatically marks the positions of the pores and debonding, thereby completing the detection; the CSPDarknet structure realizes extraction of image features, and for the extracted high-dimensional image features, the PANet network based on the CBAM mechanism realizes feature enhancement.

[0011] As a preferred scheme, in step 1 of the present application, 5000-7000 infrared images are collected, and 5000-7000 visible light images are collected.

[0012] As another preferred scheme, in step 1 of the present application, the labelimg software is used to manually mark the pores defects, debonding defects between the skin and the skeleton in the images.

[0013] As another preferred scheme, in step 2 of the present application, the images obtained in step 1 are divided into a training set, a validation set and a test set according to a ratio of 7:2:1.

[0014] As another preferred scheme, in step 3 of the present application, the number of iterative training is 300 times, and the algorithm model with the smallest training loss and the highest average precision is selected from the obtained algorithm models.

[0015] As another preferred scheme, the thermal imager of the present application adopts a spatial resolution of 700x1020 pixels, and the collection frame rate of the thermal imager and the visible light camera is set to 50Hz.

[0016] As another preferred scheme, in step 3 of the present application, the optimizer for iterative training is the stochastic gradient descent method, and the learning rate setting strategy is a cosine annealing algorithm in the range of 1x10 -4 to 0.01.

[0017] As another preferred solution, in step 4, the fused image in the test set is input into the trained algorithm model, and the algorithm model recognizes and marks the positions of the pores and debonding after operation, to complete the automatic detection of defects; the image is first subjected to the Focus structure in the CSPDarknet structure, the structure performs pixel elimination every other pixel for the input data, reduces the image size to 2 times the original size, and expands the channel number to 4 times the original number, and concentrates the width and height information into the channel information; the image is subjected to two convolution structures to realize the adjustment of the image size, is subjected to the CSPLayer structure to complete the first feature extraction, and the extracted feature is referred to as the first layer feature; the second feature extraction is completed through once again convolution structure and the CSPLayer to obtain the second layer feature; then, after the convolution structure, the SPP structure is subjected to feature extraction through different size pooling kernels to improve the receptive field of the network, and the third layer feature is obtained through the third feature extraction through the CSPLayer; the extracted three layer features are sent into the PANet structure to extract high-dimensional image features;

[0018] In the PANet structure, after the third layer feature is subjected to convolution processing, upsampling and CBAM mechanism weight adjustment are performed, and the second layer feature is completed through Concat splicing and feature extraction, the obtained feature is subjected to convolution processing, upsampling and CBAM mechanism weight adjustment are performed, and the first layer feature is completed through Concat splicing and feature extraction, and the result is output to the YoloHead structure; meanwhile, after downsampling and CBAM structure weight adjustment, the result is completed through Concat splicing and feature extraction with the result after convolution of the second layer feature, the second output is subjected to downsampling and CBAM structure weight adjustment again, and the result is completed through Concat splicing and feature extraction with the result after convolution of the first layer feature, to complete the third output.

[0019] As another preferred solution, the convolution structure is composed of a convolution layer, a BN layer and an activation function, the convolution layer can adjust the size and channel number of the image, the BN layer and the activation function can improve the feature expression ability of the neural network, effectively solve the gradient disappearance problem and accelerate the network convergence.

[0020] As another preferred solution, the CBAM mechanism described in the application, after the feature layer is subjected to global maximum pooling and global average pooling, two 1x1xchannel number feature maps are obtained, which are respectively sent into a two-layer neural network and a Relu activation function, then an addition operation and a sigmoid activation function are performed, and the obtained value is subjected to a multiplication operation with the input feature layer; the output feature layer is used as the input feature layer of the next part, and after layer-based global maximum pooling and global average pooling are performed on the input feature layer, two HxWx1 feature maps are obtained, H and W are the length and width of the feature map respectively, the two feature maps are concatenated, then subjected to 7x7 convolution operation and dimension reduction into HxWx1 feature map, and then subjected to sigmoid activation function and multiplication with the feature layer output by the previous part to obtain the final feature layer.

[0021] As another preferred solution, in step 1 of the application, the collection mode of the infrared image and the visible light image is as follows: the robot moves on the surface of the large capsule, the excitation light source heats and illuminates along the surface of the capsule, the computer control system 8 controls the thermal imager 2 and the visible light camera 4 to collect images, the images are sent into the joint network of YOLOX algorithm and spatial channel attention mechanism, and the labels of defects are obtained after operation; the excitation light source, the thermal imager and the visible light camera are arranged on the robot.

[0022] As another preferred solution, in step 3 of the application, 4 spatial channel attention mechanisms are added after the 2 up-sampling processes and the 2 down-sampling processes of the YOLOX enhanced feature layer part.

[0023] Secondly, in step 3 of the application, the loss function when the algorithm predicts is defined as follows:

[0024] Loss=a Loss α +b Loss β (1)

[0025]

[0026] Wherein, Loss represents the total prediction loss function, Loss α and Loss β represent the loss function of the visible light image prediction and the loss function of the infrared image prediction respectively; a and b are the weight coefficients of the two losses; formula (2) is the specific calculation formula of the loss function; L cls represents the classification loss of the prediction frame; L reg represents the positioning loss of the prediction frame; L obj represents the loss of the prediction category; λ represents the balance coefficient of the positioning loss; N pos represents the number of prediction frames classified as positive samples.

[0027] In addition, a in the present application is 0.3, b is 0.7, x is 1 and 2, and lambda is 5.0.

[0028] Advantages of the present application.

[0029] The present application collects infrared images and visible light images on the surface of red cysts, and uses an algorithm model to detect and mark defects automatically; it is convenient to operate and fast to detect.

[0030] The present application uses a dual-mode form combining visible light imaging and infrared imaging and image fusion (using Laplace method for feature fusion of images) to complete non-destructive testing of cysts.

[0031] The present application sends infrared image and visible light image data into a model trained based on an improved YOLOX algorithm, so as to realize rapid positioning and labeling of defects.

[0032] The present application improves the network of YOLOX, and after adding the attention mechanism, the detection accuracy of the network can be effectively improved without affecting the detection rate. BRIEF DESCRIPTION OF DRAWINGS

[0033] The present application will be further described below in combination with the drawings and specific embodiments. The protection scope of the present application is not limited to the following descriptions.

[0034] Figure 1 It is a schematic diagram of the hardware embodiment of the present application.

[0035] Figure 1 In the figure, 1 is a robot, 2 is a thermal imager, 3 is an excitation light source, 4 is a visible light camera, 5 is a light shield, 6 is a robot control signal line, 7 is an excitation control unit, and 8 is a computer control system.

[0036] Figure 2 It is a structure diagram of the YOLOX network.

[0037] Figure 3 It is a structure diagram of the spatial channel attention mechanism. DETAILED DESCRIPTION

[0038] As shown in the figure, the present application includes the following steps:

[0039] Step 1: respectively collect infrared images and visible light images on the surface of cysts, and manually mark the air hole defects, and the debonding defects of the adhesion between the skin and the skeleton in the images;

[0040] Step 2: divide the images obtained in step 1 into training set, validation set and test set, and use Laplace method to fuse the features of the visible light images and infrared images with the same frame rate;

[0041] The training set is used for feature learning of the network, the verification set is used for evaluation and reverse correction of parameters in the training process, and the test set is used for verification of the final model; the spatial channel attention mechanism module is placed in the YOLOX network to strengthen the position after each up-sampling and down-sampling of the feature layer, a total of four; the improved network of YOLOX is used to obtain the final trained weight model.

[0042] Step 3: The training set is input into the joint network based on YOLOX and spatial channel attention mechanism for training to obtain an algorithm model, and the operation result of the verification set is used to adjust the weight parameters of the network in reverse and perform iterative training; after iterative training, the last model obtained is selected, and after verification in the test set, the model with the highest accuracy is selected for the system;

[0043] The YOLOX algorithm is a single-stage algorithm, which consists of a backbone part, a strengthened feature layer and a detection head; the backbone part adopts the CSPDarknet structure, which can extract features from the image; the strengthened feature layer adopts PANet, which further extracts features from the input feature layer after up-sampling and down-sampling, and the channel spatial attention mechanism is added in the up-sampling and down-sampling process to adjust the weight of the network to improve the detection accuracy; the decoupled head is used in the detection head part, which can separately predict classification, regression and prediction probability, and the parameters are not shared in the prediction process to improve the detection accuracy;

[0044] Step 4: The fusion image in the test set is input into the trained algorithm model, and the algorithm model recognizes and marks the positions of the pores and debonding after operation, completing the automatic detection of defects. As shown in Figure 2 As shown in

[0045] As shown in Figure 3As shown, in the PANet structure, after the third layer features are processed by convolution, upsampling and CBAM mechanism, the weight is adjusted, and the second layer features are concatenated and extracted, and after the obtained features are processed by convolution, upsampling and CBAM mechanism (the structure of the CBAM mechanism is as shown in Figure 3 As shown, after the feature layer is subjected to global maximum pooling and global average pooling, two 1x1xchannel number feature maps are obtained, which are respectively sent into a two-layer neural network and a Relu activation function, and then subjected to addition operation and sigmoid activation function, and the obtained value is subjected to multiplication operation with the input feature layer; the output feature layer is used as the input feature layer of the next part, and after global maximum pooling and global average pooling based on the layer are performed, two HxWx1 feature maps are obtained, H and W are the length and width of the feature map respectively, the two feature maps are concatenated and subjected to 7x7 convolution operation to reduce the dimension to HxWx1 feature map, and then subjected to sigmoid activation function and multiplication with the feature layer output by the previous part to obtain the final feature layer) to adjust the weight, and the first layer features are concatenated and extracted, and the result is output to the YoloHead structure, and after downsampling and CBAM structure weight adjustment, the result is concatenated and extracted with the second layer features after convolution, and after the second output, the result is again subjected to downsampling and CBAM structure weight adjustment, and concatenated and extracted with the first layer features after convolution, and the third output is completed. The CBAM mechanism can be understood as combining the channel attention and spatial attention of the extracted high-dimensional image features, thereby effectively improving the accuracy of defect detection and positioning; the output three layer features are respectively sent into the YoloHead structure to predict the defect type and position and calculate the prediction probability and loss, and the defect recognition and classification are completed.

[0046] In step 1, 5000-7000 infrared images and 5000-7000 visible light images are collected; the accuracy of the model is ensured to prevent overfitting in the network operation process and ensure the generality of the model.

[0047] In step 1, the labelimg software is used to manually mark the pore defects, skin and skeleton adhesion defects in the image.

[0048] In step 2, the images obtained in step 1 are divided into training set, validation set and test set according to the ratio of 7:2:1.

[0049] In step 3, the number of iterations of training is 300, and the algorithm model with the minimum training loss and the highest average accuracy (the training loss and average accuracy of each algorithm model are parameters / quantities that can be obtained during the iterative training process of the model) is selected from the obtained algorithm models.

[0050] The thermal imager adopts a spatial resolution of 700*1020 pixels (which can obtain accurate detailed topography of defects), and the acquisition frame rate of the thermal imager and the visible light camera is set to 50 Hz (which can reduce the operation amount of the network and the consumption of computing resources as much as possible on the basis of ensuring the acquisition frequency).

[0051] In step 3, the optimizer for iterative training is the stochastic gradient descent method, and the learning rate setting strategy is the cosine annealing algorithm in the range of 1*10 -4 to 0.01.

[0052] In step 1, the collection mode of the infrared image and the visible light image is as follows: the capsule skin is uniformly laid on a plane, a robot is used to move on the surface of a large capsule, an excitation light source is used to heat and illuminate the surface of the capsule, a computer control system 8 is used to control the thermal imager 2 and the visible light camera 4 to collect images, the images are sent to the joint network of the YOLOX algorithm and the spatial channel attention mechanism, and the labels of defects are obtained after operation; the excitation light source, the thermal imager, and the visible light camera are arranged on the robot. The robot is used to facilitate movement on the surface of a large capsule and complete image data collection. The camera and the thermal imager are used to collect image and temperature data of the surface and are synchronized with the excitation system. The light source excitation module heats the surface of the capsule, then an infrared camera is used to collect temperature data for a fixed time, and the data are sent to the completed model of the defect detection module. After model operation, the size and position distribution of defects can be automatically obtained.

[0053] During operation, the thermal imager, the excitation light source, and the visible light camera are fixed on the robot, and the positions of the thermal imager and the camera are adjusted so that the images captured by the two cameras are approximately the same.

[0054] In the computer control system, the acquisition frequencies of the thermal imager and the camera are both set to 50 Hz. First, the robot is controlled to start moving, the excitation light source system is turned on, and the thermal imager and the camera are used to synchronously collect surface data.

[0055] The excitation control unit 7 (after receiving the acquisition signal issued by the computer, the excitation parameters of the excitation light source and the acquisition parameters of the thermal imager and the camera are set) sets the waveform start frequency, end frequency, detection time length and frame frequency of the thermal imager and the camera and other parameters. The computer control system 8 is physically connected with the camera and the thermal imager through a network cable, and acquires the acquisition image of the visible light and the thermal image of the thermal imager. After the visible light image and the thermal image are preprocessed, the defect detection model obtained by the improved network of YOLOX is used to complete the labeling of the millimeter-level defects and the debonding defects in the image, and the real-time processing of the image signal is realized.

[0056] The excitation light source is used to generate thermal excitation of a specified waveform, and can be an LED, a halogen lamp or the like.

[0057] A light shield can be arranged outside the excitation light source. The light shield is a conical plastic hard plate, and a layer of surface metal film (aluminum film can be used) is attached inside the plastic hard plate. The metal film has high reflectivity to the light emitted by the excitation light source, so that most of the light rays directed to the shield wall are reflected, the light energy is confined in the light shield, and finally the light energy received by the surface of the capsule is maximized and uniform, and the excitation light source can be used to generate heat.

[0058] In step 3, 4 spatial channel attention mechanisms are added after the 2 up-sampling processes and the 2 down-sampling processes in the YOLOX enhanced feature layer part.

[0059] In step 3, the loss function during algorithm prediction is defined as follows:

[0060] Loss=a Loss α +b Loss β (1)

[0061]

[0062] Wherein, Loss represents the total prediction loss function, Loss α and Loss β represent the loss function of visible light image prediction and infrared image prediction respectively; a and b are the weight coefficients of the two losses; formula (2) is the specific calculation formula of the loss function, x is α or β, representing the weight of visible light data and infrared data; L cls represents the classification loss of the prediction frame; L reg represents the positioning loss of the prediction frame; L obj represents the loss of the prediction category; λ represents the balance coefficient of the positioning loss; N pos represents the number of prediction frames classified as positive samples.

[0063] The a is 0.3, b is 0.7 (the proportion of enhanced infrared image data, supplemented by visible light data), and lambda is 5.0 (increasing the proportion of positioning loss of the prediction box).

[0064] In step 4, the trained algorithm model is an algorithm network capable of real-time detection of the fusion image of the thermal imager and the visible light camera, and the real-time detection of defects is guaranteed by the ''collection-processing'' synchronous control software developed based on the python code in the computer, and the real-time detection of defects is completed.

[0065] The present application is an automatic defect detection method combining visible light and infrared, which excites the light source to heat and illuminate, and uses a visible light camera and a thermal imager to collect images, and sends the collected images into a pre-trained algorithm model to mark the positions of pores and debonding. The present application is convenient for nondestructive testing of micron-level defects of large capsules.

[0066] The specific embodiments of the present application will be further described in detail below in combination with the drawings and examples.

[0067] In this embodiment, a 20cm*20cm*3cm size capsule material sample is automatically and quickly detected, and the sample includes debonding areas with lengths of 1.1cm, 1.5cm, 2.5cm and 3.5cm, a cutting defect with a length of 3cm, and a natural crack with a length of 15.5cm.

[0068] Step 1: deploy an infrared thermal imager, a laser pulse excitation light source and a visible light camera on a robot, wherein the thermal imager has a spatial resolution of 700*1020 pixels, the acquisition frame rate of the thermal imager and the visible light camera is set to 50Hz, the exposure time of the laser pulse excitation light source is 10s, a total of 1000 frames of thermal imaging sequences are acquired, and light-heat synchronous data acquisition is realized by computer control.

[0069] Step 2: control the robot to sample the sample at multiple angles through the signal control line of the robot, and fuse the acquired infrared and visible light image sequences by using Laplace method, label the fusion images after quality control, and the training set contains a total of 3272 labeled images.

[0070] Step 3: import the training set data into Figure 2The joint network based on YOLOX and spatial channel attention mechanism is trained and an algorithm model is obtained, and the operation results of the validation set are used to adjust the weight parameters of the network in reverse and perform iterative training, wherein the joint network based on YOLOX and spatial channel attention mechanism is designed based on Python 3.7.13 and PyTorch1.12.0-GP platform, and the training computer has Intel(R) Xeon(R) CPU, 64GB RAM, and GeForce RTX2060 12GB GPU. The minimum batch sample size used for training is 16, the training optimizer is the stochastic gradient descent method, the learning rate setting strategy is the cosine annealing algorithm in the range of 1x10 -4 0.01, after 300 iterations of training, the last 10 models obtained are selected, and the model with the highest accuracy is selected for the system after verification in the test set.

[0071] Step 4: input the fusion image in the test set into the algorithm model obtained by training, and realize the extraction of image features by the CSPDarknet structure shown in Figure 2 Step 5: realize feature enhancement for the extracted high-dimensional image features by the PANet network based on the CBAM mechanism shown in Figure 3 The CBAM mechanism can be understood as combining the extracted high-dimensional image features from the channel attention and spatial attention directions, thereby effectively improving the accuracy of defect detection and positioning.

[0072] According to the above embodiment, the YOLOX network based on the CBAM attention mechanism and the YOLO series network without considering the attention mechanism proposed in the present application are compared in terms of 1) overall class average accuracy; 2) processing frame rate; 3) model parameter amount and other indicators.

[0073] Table 1 is a performance comparison of the YOLOX network based on the CBAM attention mechanism and the YOLO series detection network without considering the attention mechanism. From the comparison results, it is not difficult to find that the YOLOX network based on the CBAM attention mechanism has excellent performance in terms of detection accuracy and parameter amount. On the one hand, the YOLOX network based on the CBAM attention mechanism has only 5.08M parameters, which is 136.3M, 56.87M and 0.98M less than Faster R-CNN, YOLOv3 and YOLOv4 Tiny, respectively. On the other hand, in terms of overall class average accuracy, the YOLOX network based on the CBAM attention mechanism is 10.48%, 15.89% and 13.46% higher than the other three networks, respectively.

[0074] In summary, the network can significantly improve the detection accuracy, and save the memory for computer operation, and therefore is more suitable for automatic detection of large capsule defects than other detection networks.

[0075] Table 1 Comparison results with other networks

[0076]

[0077] It can be understood that the above specific description of the present application is only used to illustrate the present application and is not limited to the technical solutions described in the embodiments of the present application. Those skilled in the art should understand that the present application can still be modified or replaced equivalently to achieve the same technical effect. As long as the use needs are met, it is within the protection scope of the present application.

Claims

1. A method for automatic detection of capsule defects based on infrared and visible light fusion, characterized in that The method comprises the following steps: Step 1: respectively collect infrared images and visible light images of the surface of the capsule, and manually mark the pores and the debonding defects between the skin and the skeleton in the images; Step 2: divide the images obtained in step 1 into a training set, a validation set and a test set, and fuse the visible light images and the infrared images of the same frame rate using the Laplace method; Step 3: input the training set into a joint network based on YOLOX and a spatial channel attention mechanism for training and obtaining an algorithm model, and use the operation results of the validation set to adjust the weight parameters of the network in reverse and perform iterative training; after iterative training, select the last model obtained, verify it in the test set, and select the model with the highest accuracy for the system; Step 4: input the fused images in the test set into the algorithm model obtained by training, and the algorithm model automatically marks the positions of the pores and the debonding, completing the detection; the CSPDarknet structure realizes extraction of image features, and the PANet network based on the CBAM mechanism realizes feature enhancement on the extracted high-dimensional image features; In the PANet structure, after the third layer features are processed by convolution, upsampling and CBAM mechanism weight adjustment, the features are concatenated with the second layer features for feature extraction; after the extracted features are processed by convolution, upsampling and CBAM mechanism weight adjustment, the features are concatenated with the first layer features for feature extraction, and the result is output to the YoloHead structure; meanwhile, after downsampling and CBAM structure weight adjustment, the result is concatenated with the result of the convolution of the second layer features for feature extraction, and the second output is again subjected to downsampling and CBAM structure weight adjustment, and then concatenated with the result of the convolution of the first layer features for feature extraction, and the third output is completed.

2. The method for automatic detection of capsule defects based on infrared and visible light fusion according to claim 1, characterized in that In step 2, the images obtained in step 1 are divided into a training set, a validation set and a test set according to a ratio of 7:2:

1. 3.The method of claim 1, wherein In step 3, the number of iterations of the iterative training is 300, and the algorithm model with the smallest training loss and the highest average precision is selected from the obtained algorithm models. 4.The method of claim 1, wherein The thermal imager adopts a spatial resolution of 700x1020 pixels, and the collection frame rate of the thermal imager and the visible light camera is set to 50Hz.

5. The method for automatic detection of capsule defects based on infrared and visible light fusion according to claim 1, characterized in that In step 3, the optimizer for iterative training is the stochastic gradient descent method, and the learning rate setting strategy is the cosine annealing algorithm in the range of 1 x 10 -4 to 0.

01. 6.The method of claim 1, wherein In step 4, the fused image in the test set is input into the trained algorithm model, and the algorithm model recognizes and marks the positions of pores and debonding after operation, completing automatic detection of defects; the image first passes through the Focus structure in the CSPDarknet structure, which performs pixel elimination every other pixel for the input data, reduces the image size to 2 times the original size, and expands the channel number to 4 times the original number, concentrating the width and height information into the channel information; the image passes through two convolution structures to realize image size adjustment, and passes through the CSPLayer structure to complete the first feature extraction, and the extracted features are called first-layer features; then, the second feature extraction is completed through a convolution structure and a CSPLayer, and second-layer features are obtained; then, after passing through a convolution structure, a SPP structure is used to extract features through different size pooling kernels, improve the receptive field of the network, and complete the third feature extraction through a CSPLayer to obtain third-layer features; the extracted three-layer features are sent to the PANet structure to extract high-dimensional image features.

7. The method for automatic detection of capsule defects based on infrared and visible light fusion according to claim 6, characterized in that The convolution structure is composed of a convolution layer, a BN layer and an activation function, the convolution layer can adjust the size and channel number of the image, the BN layer and the activation function can improve the feature expression ability of the neural network, effectively solve the gradient disappearance problem and accelerate the network convergence.

8. The method for automatic detection of capsule defects based on infrared and visible light fusion according to claim 6, characterized in that The CBAM mechanism, after global maximum pooling and global average pooling, two 1x1xchannel number feature maps are obtained, which are respectively sent into a two-layer neural network and a Relu activation function, then added and operated through a sigmoid activation function, and the obtained value is multiplied with the input feature layer; the output feature layer is used as the input feature layer of the next part, and after global maximum pooling and global average pooling based on the layer, two HxWx1 feature maps are obtained, H and W are the length and width of the feature map respectively, the two feature maps are concatenated and then reduced to HxWx1 feature map through 7x7 convolution operation, and then multiplied with the output feature layer of the previous part through sigmoid activation function to obtain the final feature layer. 9.The method of claim 1, wherein In step 1, the infrared image and the visible light image are collected in the following way: the robot moves on the surface of a large capsule, uses an excitation light source to heat and illuminate the surface of the capsule, and uses a computer control system 8 to control the thermal imager 2 and the visible light camera 4 to collect images, which are input into the joint network of YOLOX algorithm and spatial channel attention mechanism, and the labels of defects are obtained after operation; the excitation light source, the thermal imager and the visible light camera are arranged on the robot.

10. The method for automatic detection of capsule defects based on infrared and visible light fusion according to claim 1, characterized in that In step 3, the loss function during algorithm prediction is defined as follows: Loss = a Lossα+ b Loss β (1) wherein, Loss represents the total prediction loss function, Loss α and Loss β respectively represent the loss function of visible light image prediction and the loss function of infrared image prediction; a and b are the weight coefficients of the two losses; formula (2) is the specific calculation formula of the loss function; L cls represents the classification loss of the prediction frame; L reg represents the positioning loss of the prediction frame; L obj represents the loss of the prediction category; λ represents the balance coefficient of the positioning loss; N pos represents the number of prediction frames classified as positive samples.

Citation Information

Patent Citations

  • Insulator defect detection method and system, electronic equipment and readable storage medium

    CN113298789A

  • Target detection method based on space attention and channel attention

    CN114882237A