An ammonia leakage mixed self-attention lightweight infrared detection method and system
By constructing an infrared detection method for shuffling and self-attention lightweight ammonia leak, the problems of small range, poor sensitivity and low real-time performance in the prior art are solved, and real-time accurate monitoring at a longer distance is achieved to ensure safety and reduce losses.
Patent Information
- Application Number
- CN202210516324.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-12
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-05-12
AI Technical Summary
The existing ammonia leak detection device has a small detection range, poor sensitivity, low real-time performance, and cannot locate the leakage source. It is impossible to realize real-time accurate monitoring of ammonia leaks, which poses safety hazards and economic losses.
By constructing an infrared detection method for shuffling self-attention lightweight infra-red detection of ammonia leakage, using infrared image detection data set for normalization and data annotation, a shuffling self-attention network structure is constructed, and a shuffling self-attention network model is trained and tested to achieve real-time accurate monitoring of ammonia leakage.
Realize real-time and accurate monitoring of ammonia leakage at a longer distance, ensure the safety of staff, respond to ammonia leakage in a timely manner, and reduce economic losses.
Smart Images

Figure CN114964628B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and gas leakage detection, and in particular to a method and system for detecting ammonia leakage using a shuffled self-attention lightweight infrared method. Background Art
[0002] Ammonia, as an important basic industrial raw material, is widely used in industrial production such as cold chain logistics and aerospace. However, ammonia leakage is a major safety hazard in the production process. If leakage is not handled in time, it will not only cause poisoning but also pose an explosion hazard. For a long time, many major ammonia leakage accidents have caused huge economic losses, and there is an urgent need to develop real-time ammonia leak detection technology.
[0003] Conventional ammonia leak detection methods use gas sensors for point measurement. Their contact-based principle prevents access to many test areas, resulting in limited detection range, poor sensitivity, limited real-time performance, and an inability to locate the leak source, posing a significant threat to worker safety. Infrared imaging detection technology can display images of hazardous gas leaks on a display, enabling large-scale, real-time, and visual detection of hazardous gas leaks. The key to this problem lies in designing a non-contact detection model for hazardous gas leaks based on infrared imaging.
[0004] Non-contact ammonia leak detection can image and locate leaking gas at a distance, thus ensuring worker safety. While some progress has been made, the model is complex and cumbersome, making implementation difficult and making real-time ammonia leak detection difficult. Infrared images captured by thermal imagers contain high noise and contrast, resulting in poor detection results. Furthermore, ammonia leaks often occur at irregular times and locations. To ensure safety, a higher level of real-time performance is required. Existing detection methods are no longer sufficient, and accurate, real-time detection of ammonia leaks has become a pressing technical challenge. Summary of the Invention
[0005] The purpose of this application is to provide a mixed self-attention lightweight infrared detection method and system for ammonia leakage, which is used to solve the technical problems of conventional ammonia leakage detection devices in the prior art, such as small detection range, poor sensitivity, low real-time performance, and inability to locate the source of leakage, so as to achieve real-time and accurate monitoring of ammonia leakage at a long distance, ensure the safety of staff and respond in time when ammonia leakage occurs, thereby reducing economic losses.
[0006] In view of the above problems, the present application provides a lightweight infrared detection method and system for ammonia leak shuffling self-attention.
[0007] In the first aspect of the present application, a lightweight infrared detection method for ammonia leakage with shuffled self-attention is provided. The method includes: obtaining an infrared image detection data set; normalizing the infrared image detection data set and then dividing it into a training set and a test set, and performing ammonia leakage infrared data annotation on the infrared image detection data set; constructing a shuffled self-attention network structure; training the shuffled self-attention network structure with the training set having the ammonia leakage infrared data annotation to obtain a shuffled self-attention network model; detecting the test set through the shuffled self-attention network model, calling the final shuffled self-attention network model and the test program, and inputting an ammonia leakage infrared image to determine the detection result.
[0008] In the second aspect of the present application, a lightweight infrared detection system for ammonia leakage with shuffled self-attention is provided. The system includes: a first obtaining unit configured to obtain an infrared image detection data set; a first processing unit configured to normalize the infrared image detection data set and then divide it into a training set and a test set, and perform ammonia leakage infrared data annotation on the infrared image detection data set; a first constructing unit configured to construct a shuffled self-attention network structure; a second processing unit configured to train the shuffled self-attention network structure with the training set having the ammonia leakage infrared data annotation to obtain a shuffled self-attention network model; a third processing unit configured to detect the test set through the shuffled self-attention network model, call the final shuffled self-attention network model and the test program, and input an ammonia leakage infrared image to determine the detection result.
[0009] In the third aspect of the present application, a lightweight infrared detection system for ammonia leakage with shuffled self-attention is provided, including: a processor coupled to a memory, the memory being configured to store a program, and when the program is executed by the processor, the system is caused to execute the steps of the method as described in the first aspect.
[0010] One or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0011] A lightweight infrared detection method for ammonia leakage with shuffled self-attention provided by this application obtains an infrared image detection data set; normalizes the infrared image detection data set and divides it into a training set and a test set, and annotates the infrared data of ammonia leakage for the infrared image detection data set; constructs a shuffled self-attention network structure; uses the training set with the infrared data annotation of ammonia leakage to train the shuffled self-attention network structure to obtain a shuffled self-attention network model; detects the test set through the shuffled self-attention network model, calls the final shuffled self-attention network model and the test program, and inputs the infrared image of ammonia leakage to determine the detection result. It solves the technical problems in the prior art that the conventional ammonia leakage detection device has a small detection range, poor sensitivity, low real-time performance, and inability to locate the leakage source, and achieves the technical effect of being able to realize real-time and accurate monitoring of ammonia leakage at a relatively long distance, ensuring the safety of staff and making a timely response when ammonia leaks, and reducing economic losses.
[0012] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically gives the specific implementation manners of this application. Brief Description of the Drawings
[0013] Figure 1 It is a schematic flowchart of a lightweight infrared detection method for ammonia leakage with shuffled self-attention provided by this application;
[0014] Figure 2 It is a schematic diagram of the network structure of the shuffled self-attention network model (SSANet) of a lightweight infrared detection method for ammonia leakage with shuffled self-attention provided by this application;
[0015] Figure 3 It is a schematic diagram of the structure of the SK5 Block module of the lightweight feature extraction network with a stride of 1 (Stride = 1) in a lightweight infrared detection method for ammonia leakage with shuffled self-attention provided by this application;
[0016] Figure 4 It is a schematic diagram of the structure of the SK5 Block module of the lightweight feature extraction network with a stride of 1 (Stride = 2) in a lightweight infrared detection method for ammonia leakage with shuffled self-attention provided by this application;
[0017] Figure 5 It is a schematic diagram of the structure of the Transformer encoding layer in a lightweight infrared detection method for ammonia leakage with shuffled self-attention provided by this application;
[0018] Figure 6Schematic diagram of the Transformer module in a method for ammonia leakage mixed-shuffle self-attention lightweight infrared detection provided by this application;
[0019] Figure 7 Schematic diagram of the structure of a system for ammonia leakage mixed-shuffle self-attention lightweight infrared detection provided by this application;
[0020] Figure 8 Schematic diagram of the structure of an exemplary electronic device of this application.
[0021] Explanation of reference numerals: First acquisition unit 11, first processing unit 12, first construction unit 13, second processing unit 14, third processing unit 15, electronic device 300, memory 301, processor 302, communication interface 303, bus architecture 304. Detailed implementation manners
[0022] This application provides a method and system for ammonia leakage mixed-shuffle self-attention lightweight infrared detection, which are used to solve the technical problems in the prior art that the detection range of conventional ammonia leakage detection devices is small, the sensitivity is poor, the real-time performance is not high, and the leakage source cannot be located, and achieve the technical effect of realizing real-time and accurate monitoring of ammonia leakage at a relatively long distance, ensuring the safety of staff and making timely responses when ammonia leaks, and reducing economic losses.
[0023] In view of the above technical problems, the general idea of the technical solution provided by this application is as follows:
[0024] The method provided by the embodiments of this application obtains an infrared image detection data set; after normalizing the infrared image detection data set, it is divided into a training set and a test set, and the infrared data of ammonia leakage in the infrared image detection data set is labeled; a mixed-shuffle self-attention network structure is constructed; the training set with the infrared data label of ammonia leakage is used to train the mixed-shuffle self-attention network structure to obtain a mixed-shuffle self-attention network model; the test set is detected through the mixed-shuffle self-attention network model, and the final mixed-shuffle self-attention network model and the test program are called, and the infrared image of ammonia leakage is input to determine the detection result.
[0025] After introducing the basic principle of this application, below, the technical solutions in this application will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the exemplary embodiments described here. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application. Additionally, it should be noted that for the sake of description, only the parts related to this application are shown in the accompanying drawings rather than all of them.
[0026] Example 1
[0027] As Figure 1 shown, this application provides a method for lightweight infrared detection of ammonia leakage by shuffling self-attention, and the method includes:
[0028] S100: Obtain an infrared image detection data set;
[0029] Specifically, build an image acquisition system, and use an infrared thermal imager to collect ammonia leakage videos and then perform frame extraction to obtain infrared images; for the ammonia leakage infrared images collected by the infrared thermal imager with high noise and low contrast, in order to facilitate and accelerate the model's processing of data, avoid the increase in network training time caused by the existence of singular sample data, and may cause the network to not converge. First, use a non-local mean denoising method to remove randomly distributed noise points in the background area, and at the same time, better maintain the local features of the edge pixels in the ammonia cloud area. Then, use the adaptive histogram equalization algorithm with limited contrast to preprocess the ammonia leakage infrared images for contrast enhancement to establish an ammonia leakage infrared image detection data set.
[0030] S200: After normalizing the infrared image detection data set, divide it into a training set and a test set, and perform ammonia leakage infrared data annotation on the infrared image detection data set;
[0031] Specifically, before processing the data, first normalize the data to eliminate the adverse effects caused by singular sample data, accelerate the speed of gradient descent to find the optimal solution, and improve the accuracy of the model in processing data.
[0032] Normalize the image sizes in the infrared image detection data set and divide them into a training set and a test set. Use open-source tools to perform ammonia leakage infrared data annotation on the obtained infrared image detection data set. Preferably, the annotation information in this embodiment may include the position of the leaked ammonia and the size of the leakage area in the annotation bounding box coordinate area.
[0033] S300: Construct a shuffling self-attention network structure;
[0034] S400: Use the training set with the ammonia leakage infrared data annotation to train the shuffling self-attention network structure to obtain a shuffling self-attention network model;
[0035] Specifically, in the embodiments of the present application, a Shuffle Self-Attention Network Model (SSANet) is constructed to achieve real-time non-contact detection of ammonia leakage by infrared. To improve the detection accuracy of the ammonia leakage infrared detection model while meeting the detection speed, the 3×3 depthwise separable convolution kernel is designed as a 5×5 SK5 Block module to reconstruct the feature extraction network. According to the leakage area at different times of ammonia leakage, the receptive field is expanded while reducing the model inference calculation amount and compressing the model weight size to achieve real-time detection of the model. The SSANet model uses the Transformer module as the bottleneck layer from bottom to top of the model feature pyramid to realize the bottom-up fusion of multi-head attention of the leakage area, obtain multi-scale feature information, and thus realize the fusion of global features and local features to improve the accuracy of the model.
[0036] The shuffle self-attention network structure is trained and learned using the training set composed of the ammonia leakage infrared data annotations to obtain a shuffle self-attention network model.
[0037] S500: Detect the test set through the shuffle self-attention network model, call the final shuffle self-attention network model and the test program, and input the ammonia leakage infrared image to determine the detection result.
[0038] Specifically, the shuffle self-attention network model obtained in the above steps is tested through the test set composed of the ammonia leakage infrared data annotations. When the output result of the model reaches the expected output result, the final shuffle self-attention network model and the test program are called, and the detection result is output by inputting the ammonia leakage infrared image into the final shuffle self-attention network model, realizing lightweight real-time detection of ammonia leakage by infrared, and providing an effective real-time detection method for developing a non-contact detection device for ammonia leakage to ensure the safe production and stable operation of ammonia-related enterprises.
[0039] The method provided by the present application obtains an infrared image detection data set; normalizes the infrared image detection data set and divides it into a training set and a test set, and performs ammonia leakage infrared data annotation on the infrared image detection data set; constructs a shuffle self-attention network structure; trains the shuffle self-attention network structure using the training set with the ammonia leakage infrared data annotations to obtain a shuffle self-attention network model; detects the test set through the shuffle self-attention network model, calls the final shuffle self-attention network model and the test program, and inputs the ammonia leakage infrared image to determine the detection result. It solves the technical problems in the prior art that the conventional ammonia leakage detection device has a small detection range, poor sensitivity, low real-time performance, and inability to locate the leakage source, and achieves the technical effect of being able to realize real-time and accurate monitoring of ammonia leakage at a relatively long distance, ensuring the safety of staff and making a timely response when ammonia leakage occurs, and reducing economic losses.
[0040] Step S100 in the method provided by the embodiment of the present application includes:
[0041] S110: Collect an ammonia leakage video through an infrared thermal imager, perform frame extraction on the ammonia leakage video to obtain an infrared image set;
[0042] S120: Perform non-local means denoising on the infrared image set to obtain a denoised infrared image set;
[0043] S130: Perform adaptive histogram equalization processing based on the denoised infrared image set to obtain the infrared image detection data set.
[0044] Specifically, the infrared imaging detection technology can present the harmful gas leakage image on the display, making it possible to detect harmful gas leakage in a large range, in real time, and visually.
[0045] After collecting the ammonia leakage video with an infrared thermal imager and performing frame extraction, infrared images are obtained. Aiming at the high noise and low contrast of the ammonia leakage infrared images collected by the infrared thermal imager, first, the non-local means denoising method is used to remove the randomly distributed noise points in the background area, and at the same time, the local features of the edge pixels in the ammonia cloud area are well preserved. Then, the contrast-limited adaptive histogram equalization algorithm is used to perform preprocessing on the ammonia leakage infrared images to enhance the contrast and establish the ammonia leakage infrared detection data set.
[0046] Exemplarily, after collecting the ammonia leakage video with a long-wave infrared thermal imager with a wavelength of 8 - 14 μm and performing frame extraction, infrared images are obtained; the non-local means denoising (Non-Local Means, NL-means) and contrast-limited adaptive histogram equalization (Contrast Limited Adaptive Histgram Equalization, CLAHE) algorithms are used to perform preprocessing on the ammonia leakage infrared images for noise removal and contrast enhancement;
[0047] The non-local means denoising algorithm makes full use of the redundant information and information with similar structures in the image. By searching for all similar blocks within the entire image range, it performs weighted averaging on the similar structures to eliminate noise while maximizing the preservation of the detailed features of the image. The non-local means denoising algorithm is used to process the infrared image of ammonia leakage for image denoising to reduce the noise in the infrared image. After being processed by the non-local means denoising algorithm, the infrared image is obtained. The large area of white noise dots in the image disappears, making the image clearer, and no details are lost, with an obvious denoising effect. On the basis of denoising the infrared image of ammonia leakage by the non-local means, to further improve the image quality, the contrast-limited adaptive histogram equalization (CLAHE) algorithm is used to process the infrared image of ammonia leakage. CLAHE makes the processed area finer than the original image. While enhancing the contrast, it can suppress noise. CLAHE first divides the image into individual sub-blocks, calculates the histogram in the sub-blocks, and then uses the method of limiting the contrast to reasonably distribute the parts of each sub-block that exceed the clipping limit to other areas in the histogram, achieving the effect of redistributing each sub-block of the histogram. On the basis of the infrared image denoised by the non-local means, using the contrast-limited adaptive histogram equalization algorithm to process the infrared image of ammonia leakage makes the gas cloud of ammonia leakage more obvious and the details more abundant, which is beneficial to the detection of the model.
[0048] Furthermore, to analyze the effectiveness of the image preprocessing algorithm, three quantitative evaluation indicators of the image, namely peak signal-to-noise ratio (PSNR), average gradient (AG), and information entropy (IE), are selected to quantitatively evaluate the enhanced infrared image of the present invention by taking the mean value. The peak signal-to-noise ratio is an evaluation index for the quality of the reference image, which describes the difference degree and noise resistance between the denoised image and the original image. The larger its value, the better the denoising effect and the better the overall visual effect of the human eye. The average gradient represents the contrast of the fine parts of the image. The larger the evaluation gradient, the clearer the image. The information entropy represents the amount of information in the image. The larger the information entropy, the richer the detailed information of the image. The above three indicators are used to evaluate the original infrared image of ammonia leakage and the preprocessed image of the image, and the results are shown in Table 1.
[0049] Table 1. Quantitative evaluation indicators of image preprocessing
[0050]
[0051] As can be seen from Table 1, the peak signal-to-noise ratio of the preprocessed image is as high as 23.40 dB compared with the original image, indicating that the image has improved noise resistance and a better denoising effect. The preprocessed image is higher than the original image in terms of the average gradient and information entropy indicators, indicating that the preprocessed image effectively improves the contrast of the image, making the image clearer and the detailed texture more abundant. The comparison of the algorithm performance is shown in Table 2. The model obtained by training with the preprocessed image is denoted as Prep-YOLOv5s.
[0052] Table 2 Comparison of Network Performance before and after Image Preprocessing
[0053]
[0054] As can be seen from Table 2, the Prep-YOLOv5s model has a 1.00% increase in mAP compared to the model trained with the original infrared images under the premise that the speed is basically the same. This shows that the image preprocessing method can improve the accuracy of ammonia leakage detection while ensuring the unchanged detection speed of the model. Experiments verify that the denoising effect is obvious without losing details through image preprocessing, and the gas cloud is more obvious, which is helpful for ammonia leakage detection.
[0055] Step S200 in the method provided by the embodiments of this application further includes:
[0056] S210: Perform image size preprocessing on the infrared image detection data in the infrared image detection dataset;
[0057] S220: Based on the infrared image detection data after size preprocessing, determine the target detection area, and set a bounding box according to the target detection area;
[0058] S230: Generate a label file in the format of a preset dataset according to the bounding box of the target detection area;
[0059] S240: Annotate the label file in the format of the preset dataset, and the annotation information includes the ammonia leakage position and the leakage area size of the bounding box;
[0060] S250: Convert the label file in the format of the preset dataset to determine the ammonia leakage infrared data annotation, where the ammonia leakage infrared data annotation is a label file in the first preset format and includes the center point coordinates of the bounding box.
[0061] Specifically, after normalizing the size of the obtained ammonia leakage infrared images, they are divided into a training set and a test set. The open-source tool LabeImg is used to manually annotate the obtained ammonia infrared image data, that is, the bounding boxes of the corresponding targets are marked for all ammonia leakage areas to be detected on each image, generating an xml label file that conforms to the VOC2007 dataset format. The annotation information includes the position of the leaked ammonia and the size of the leakage area in the coordinate area of the annotated bounding box. The xml file is converted into a txt label file, and the txt label file includes the center point coordinates of the bounding box of the ammonia leakage area.
[0062] Exemplarily, the 906×720 image output by the infrared camera is converted into a 640×640 size image suitable for input to the network model. The obtained ammonia infrared image data is manually annotated using the open-source tool LabeImg, that is, the bounding boxes corresponding to the targets are marked for all the ammonia leakage areas to be detected on each image, generating an xml tag file in the dataset format conforming to VOC2007. The annotation information includes the position of the leaked ammonia in the coordinate area of the annotated bounding box and the size of the leakage area. The xml file is converted into a txt tag file, and the txt tag file includes the central point coordinates of the bounding box of the ammonia leakage area.
[0063] As Figure 2 shown, step S300 in the method provided by the embodiment of the present application further includes:
[0064] S310: Construct a detection network model SSANet;
[0065] S320: Use the lightweight SK5 Block module to construct the feature extraction network of the detection network model SSANet, where the feature extraction network includes 5×5 depthwise separable convolution and a channel shuffle structure;
[0066] S330: Add a Transformer module with a self-attention mechanism to the feature extraction network as the bottleneck layer from bottom to top of the model feature pyramid, where the Transformer module includes an image slicing part, a data embedding part, and an encoding layer part.
[0067] Specifically, ammonia is a colorless, flammable, explosive and toxic gas, and the leakage has the characteristics of occurring at uncertain times and locations, which poses extremely high requirements for the sensitivity and real-time performance of its detection. A shuffle self-attention network model (SSANet) is designed to achieve real-time non-contact detection of ammonia leakage by infrared.
[0068] To enable the ammonia detection model to have less inference computation and accelerate the inference speed of the model to achieve real-time ammonia leakage detection, the feature extraction network of the model is reconstructed to reduce the computation during model inference and compress the size of the model weights. By comparing the metrics of different feature extraction networks, it is found that SK5-YOLOv5s has obvious advantages in terms of model weight size, speed, accuracy, etc. In the embodiments of this application, the lightweight SK5 Block module is preferably used to construct the feature extraction network of the model to achieve a balance between speed and accuracy. Specifically, in this embodiment, the evaluation metrics of different feature extraction networks are compared. The feature extraction network of the YOLOv5s model is replaced with the lightweight architectures GhostNet, MobileNetv3, and ShuffleNetv2 respectively and compared with the method of the present invention in a comparative experiment. When the feature extraction network of the model is constructed by the SK5 Block module during the experiment, the same training and testing schemes are followed. The experimental results are shown in Table 3.
[0069] Table 3 Comparison of evaluation metrics of different feature extraction networks
[0070]
[0071] As can be seen from Table 3, the model weight size of the SK5-YOLOv5s model improved by lightweight is compressed to 3.4M, and the detection speed of a single image reaches 2.7ms. Compared with the ShuffleNetv2-YOLOv5s model, the number of parameters increases slightly by 4%, but the accuracy improvement reaches 94.4%. This verifies the effectiveness of using the SK5 Block module to construct the feature extraction network of the model, and at the same time verifies the effectiveness of expanding the 3×3 depthwise separable convolution kernel to 5×5 to expand the receptive field and improve the model detection accuracy without increasing too much computation, meeting the balance between model lightweight and model detection accuracy. When using MobileNetv3 and GhostNet as the feature extraction network of the model for lightweight, although the detection speed of the model is increased and the model weight size is reduced, the detection accuracy of the model is not as good as that of the SK5-YOLOv5s model. Compared with the SK5-YOLOv5s model, the number of parameters and the model weight are larger, and the speed is slower. By comprehensively comparing the number of parameters, model weight size, speed, and accuracy of the models, SK5-YOLOv5s has obvious advantages in all aspects. Therefore, the SK5-YOLOv5s structure is selected as the feature extraction network of the lightweight infrared detection model for ammonia leakage.
[0072] The SK5 Block module is mainly designed with 1×1 ordinary convolution (Conv1×1), 3×3 depthwise separable convolution kernel expanded to 5×5 depthwise convolution (Depthwise Convolution, DWConv) and channel shuffle (Channel Shuffle) structure.
[0073] In the case of ammonia leakage, there are phenomena such as the irregular movement of the ammonia cloud and the low contrast and unclear features in the diffusion area in the infrared image, resulting in the defect that the model focuses on local features while ignoring global features and has low detection accuracy. Introducing the Transformer module with self-attention mechanism can convert the infrared image into a sequence and then input it into the model for processing. Based on the feature extraction network reconstructed by the SK5 Block module, Transformer is used as the bottom-up bottleneck layer of the model feature pyramid to obtain multi-scale feature information, thus realizing the fusion of global features and local features.
[0074] The Transformer module can capture global information and rich context information and perform global inference on images and specific predicted targets. Applying the Transformer module to low-resolution feature maps can reduce expensive computing and storage costs while increasing the importance of leakage features and improving the detection accuracy of the model. Specifically, in this embodiment, performance comparisons were made by adding different bottleneck layer structures to the bottom-up bottleneck layer of the feature pyramid in the Neck part of the model. The different bottleneck layer structures added were CSPBottleNeck, GhostBottleNeck, CbamBottleNeck, and TransformerBottleNeck. Among them, the model trained by adding the TransformerBottleNeck structure is denoted as SSANet, and the experimental results are compared as shown in Table 4.
[0075] Table 4 Comparison of network performance of different BottleNeck structures
[0076]
[0077] As can be seen from Table 4, the feature pyramid of the Neck part of the model has a significant impact on the model when different bottleneck layer structures are added to the bottom-up bottleneck layer. The SSANet model uses the Transformer module, and the detection speed is reduced by 1.85% compared with the SK5-YOLOv5s model. The detection speed of a single image reaches 3.2 ms, and the detection accuracy is improved by 1.90% to reach 96.30%. When other different bottleneck layer structures are added, although the detection speed is improved, the detection accuracy is not as good as that of the Transformer module structure. The experiment verifies that the Transformer module adopted in this paper integrates feature embedding to aggregate the feature map information of different scales, and at the same time makes full use of the feature interaction across space and scales through the multi-head attention mechanism to enhance the attention to the ammonia leakage target. Therefore, the Transformer module is selected as the bottom-up bottleneck layer of the feature pyramid of the Neck part of the model for feature fusion to extract the target attention to ensure the improvement of accuracy for the lightweight infrared real-time detection of ammonia leakage. Among them, the Transformer module includes an image slicing part, a data embedding part, and an encoding layer part.
[0078] Before step S310 in the method provided by the embodiment of the present application, the method further includes:
[0079] Step 1: Obtain a first sample according to the infrared image detection data set, and use the first sample as the first initial clustering center;
[0080] Step 2: Calculate the distance between each data sample in the infrared image detection data set and the first initial clustering center, and select the minimum distance from all sample distances as the first clustering;
[0081] Step 3: Based on the calculated distances of all samples, determine the sample with the maximum distance, and use the sample with the maximum distance as the second clustering center;
[0082] Step 4: Repeat Step 2 - Step 3 until K clustering centers are obtained, and perform clustering calculation on the K clustering centers to obtain the annotation box size;
[0083] Step 5: According to the annotation box size, obtain the ammonia leakage size information, where the ammonia leakage size information is the area and aspect ratio information of the leakage area of ammonia leakage at different times, and use the ammonia leakage size information as the candidate box parameter of the detection network model SSANet.
[0084] Specifically, the K-means algorithm is used to perform clustering analysis on the ammonia leakage infrared data set to analyze the area and aspect ratio of the ammonia leakage area at different times to be applicable to the candidate box of ammonia leakage infrared detection and the candidate box parameter model parameters of the preset target detection model;
[0085] There is only one category for ammonia leakage detection, and ammonia leakage diffuses from nothing to something. As an iterative clustering K-means clustering algorithm, distance is used as the similarity index to discover K classes in a given dataset, and the center of each class is obtained according to the mean value of all numerical values in the class. The center of each class is described by the clustering center. The value of the clustering center is updated successively through an iterative method until the best clustering result is obtained. Randomly select a sample from the ammonia leakage infrared detection dataset as the first initialized clustering center;
[0086] Calculate the distance between each sample point in the sample and the initialized clustering center, and select the shortest distance among them as the first clustering;
[0087] Based on the calculated distances of all samples, determine the sample with the maximum distance, and use the sample with the maximum distance as the second clustering center;
[0088] Repeat the above two steps until k clustering centers are selected;
[0089] Use the K-means algorithm to calculate the final clustering result for the k clustering centers to obtain the size of the bounding box;
[0090] Use the size of the bounding box obtained by clustering analysis to obtain the area and aspect ratio of the leakage area at different times of ammonia leakage, and modify the candidate box parameters of the lightweight non-contact ammonia leakage infrared real-time detection SSANet model.
[0091] Preferably, in the embodiment of the present application, 9 clustering centers are used to cluster the size of the bounding box in the ammonia leakage infrared detection dataset, classify the size of the bounding box close to the center point, and update the value of the clustering center successively through an iterative method until the optimal clustering result is obtained, and use the size of the bounding box obtained by clustering analysis to modify the candidate box parameters of the object detection model.
[0092] Furthermore, there is only one category for ammonia leakage detection, and ammonia leakage diffuses from nothing to something. The size of the candidate box changes regularly from small to large. The candidate box parameters of the YOLOv5s basic model cannot meet the actual needs of ammonia leakage infrared image detection, and it is necessary to redesign the candidate box parameters that change regularly from small to large to meet the size requirements during the ammonia leakage process. The K-means clustering algorithm is used to perform clustering calculations on the size of the ground truth boxes annotated in the dataset. The K-means clustering algorithm is an iterative clustering algorithm, which enables the model to select more accurate candidate boxes to better reflect the characteristics of the target, avoids the model blindly searching during training, improves the detection effect of the model, and helps the model converge quickly to achieve real-time performance.
[0093] YOLOv5s has a total of three detection layers. Each detection layer has three candidate boxes with different aspect ratios to identify and locate the target, for a total of nine candidate boxes. Therefore, an image with a size of 640×640 pixels is used as the input. K-means clustering analysis is performed on the ammonia leakage infrared detection dataset images with the number of candidate boxes being nine. The comparison of the candidate box sizes of the three detection layers before and after clustering is shown in Table 5.
[0094] Table 5 Initial candidate box sizes of the detection layer
[0095]
[0096] As can be seen from Table 5, before clustering, the change distribution range of the aspect ratios of the candidate boxes of the YOLOv5s base model is 0.70 - 2.03, while after clustering analysis of the ammonia leakage infrared detection dataset, the aspect ratio distribution range of the obtained candidate boxes is mainly 0.19 - 1.17. The change in the aspect ratio distribution range of the candidate boxes before and after clustering is relatively large, which is consistent with the situation where the true boxes marked gradually increase from small to large, indicating that the effect of using the K-means algorithm for clustering analysis of the ammonia leakage infrared detection dataset is obvious. The candidate box parameters obtained by clustering are helpful for the model to identify and locate the target, and can improve the convergence speed and accuracy of the model.
[0097] The YOLOv5s model is trained using the original image without using K-means clustering candidate boxes, the original image with K-means clustering candidate boxes, and the preprocessed image with K-means clustering candidate boxes. Among them, the model trained using the original image with K-means clustering candidate boxes is denoted as Kms-YOLOv5s, and the model trained using the preprocessed image with K-means clustering candidate boxes is denoted as Kms-Prep-YOLOv5s. The performance comparison experiments of the three are shown in Table 6.
[0098] Table 6 Comparison of network performance before and after clustering
[0099]
[0100] As can be seen from Table 6, re-pre-setting the candidate box parameters has no effect on the size of the model weights and the detection speed. The size of all model weights is 14.40M and the detection speed of a single image remains 3.6ms. However, the detection accuracy has been significantly improved. The average precision mAP of the Kms-YOLOv5s model trained using the original image with K-means clustering candidate boxes has increased by nearly 1.00%; the average precision mAP of the Kms-Prep-YOLOv5s trained using the preprocessed image with K-means clustering candidate boxes has increased by 1.50%, indicating that the aspect ratios of the candidate boxes pre-set using the K-means clustering algorithm conform to the true size of ammonia leakage and effectively improve the accuracy of the model for ammonia leakage detection.
[0101] Step S400 in the method provided by the embodiment of the present application further includes:
[0102] S410: Using K clustering centers to perform annotation box size clustering on the ammonia leakage infrared data set in the training set, obtaining a clustering result, and using the clustering result as the model candidate box parameter;
[0103] S420: Extracting features from the ammonia leakage infrared data through the feature extraction network to obtain feature detection extraction information;
[0104] S430: Performing bottom-up fusion on the feature detection extraction information through the Transformer module to obtain multi-scale feature information;
[0105] S440: Training the detection network model SSANet using each ammonia leakage infrared data in the training set to obtain the shuffle self-attention network model, where the shuffle self-attention network model is used to extract the multi-scale feature information to detect ammonia leakage information.
[0106] Specifically, first, the K-means algorithm is used to perform clustering analysis on the ammonia leakage infrared data set to analyze the area and aspect ratio of the ammonia leakage area at different times to obtain candidate boxes suitable for ammonia leakage infrared detection and candidate box parameter models of the preset target detection model; secondly, considering that ammonia is a colorless, flammable, explosive and toxic gas and the leakage occurs at unpredictable times and locations, which poses extremely high requirements for the real-time detection, a shuffle self-attention network (SSANet) model is designed. The 3×3 depthwise separable convolution kernel is redesigned as a 5×5 SK5 Block module to reconstruct the feature extraction network, expanding the receptive field and improving the detection accuracy while reducing the model weight; finally, the Transformer module is used as the bottleneck layer of the model feature pyramid from bottom to top to realize the bottom-up fusion of the multi-head attention of the leakage area, obtaining multi-scale feature information so as to realize the fusion of global features and local features to improve the accuracy of the model. The SSANet is used to train the ammonia leakage infrared detection data set to obtain the final ammonia leakage infrared detection model. Exemplarily, taking K = 9 as an example, the specific steps for detecting ammonia leakage information are as follows:
[0107] Step 1: Using 9 clustering centers to perform clustering on the annotation box sizes in the ammonia leakage infrared detection data set, classifying the annotation box sizes close to the center point, and gradually updating the clustering center values through an iterative method until the optimal clustering result is obtained, and modifying the candidate box parameters of the target detection model using the annotation box sizes obtained by the clustering analysis;
[0108] Step 2: Since ammonia is a colorless, flammable, explosive and toxic gas, and the leakage occurs at unpredictable times and locations, extremely high requirements are put forward for its real-time detection. To enable the ammonia detection model to have less inference calculation and accelerate the inference speed of the model to achieve real-time ammonia leakage detection. The present invention reconstructs the feature extraction network of the model to reduce the calculation amount during model inference and compress the size of the model weights. The Shuffle Self-Attention Network model (SSANet) reconstructs the lightweight shuffle feature extraction network by designing the 3×3 depthwise separable convolution kernel into a 5×5 SK5Block module. The SK5 Block module is mainly designed with a 1×1 ordinary convolution and a 3×3 depthwise separable convolution kernel, expanded into a 5×5 depthwise convolution (DWConv) and a channel shuffle structure.
[0109] According to different moving strides (Stride) of the convolution kernel, two lightweight convolution modules are designed. One is a lightweight SK5 Block module with a stride of 1 (Stride = 1), and the other is a lightweight SK5 Block module with a stride of 2 (Stride = 2).
[0110] Step 3: In the case of ammonia leakage, there is an irregular movement of the ammonia cloud and a low contrast and unclear feature in the diffusion area in the infrared image, resulting in the defect that the model focuses on local features and ignores global features, with low detection accuracy. Introducing a Transformer module with a self-attention mechanism can convert the infrared image into a sequence and then input it into the model for processing. Based on the feature extraction network reconstructed by the SK5 Block module, Transformer is used as the bottom-up bottleneck layer of the model feature pyramid to obtain multi-scale feature information, thereby realizing the fusion of global features and local features. The Transformer module can capture global information and rich context information and perform global inference on images and specific predicted targets. Applying the Transformer module to low-resolution feature maps can reduce expensive computing and storage costs while increasing the importance of leakage features and improving the detection accuracy of the model.
[0111] Step 4: In the Transformer encoding layer part of the SSANet model, after performing multiple groups of self-attention processing on the original input sequence, two fully connected layers (Linear) are used to replace the original layer normalization process. Two linear fully connected transformations are performed to reduce the computational complexity while effectively reducing the impact of the sample batch size. In this way, by aggregating the feature information of different branches, the feature space extracted by the backbone network is enriched. The Transformer module fuses the feature embeddings to aggregate the feature map information of different scales. At the same time, through the multi-head attention mechanism, the feature interactions across space and scales are fully utilized to enhance the attention to the ammonia leakage target. Compared with other different bottleneck layer structures, for lightweight infrared real-time detection of ammonia leakage, the Transformer module is selected as the feature pyramid bottom-up bottleneck layer of the model Neck part for feature fusion to extract target attention and ensure the improvement of accuracy.
[0112] As Figure 3 and Figure 4 shown, step S420 in the method provided by the embodiment of the present application further includes:
[0113] S421: When it is the first-step lightweight convolution module, divide the input feature map channels into a first branch and a second branch. The first branch retains its own information and passes it down to obtain the first-channel feature. The second branch passes through a 1×1 convolution channel and is fused with a 5×5 depthwise separable convolution to obtain the second-channel feature;
[0114] S422: Shuffle based on the first-channel feature and the second-channel feature through the channel shuffle structure to obtain the feature detection and extraction information;
[0115] S423: When it is the second-step lightweight convolution module, divide the input feature map into two equally mapped branches to obtain two branch features, and shuffle the two branch features to obtain the feature detection and extraction information for output.
[0116] Specifically, to make the ammonia detection model have less inference calculation amount and accelerate the inference speed of the model to achieve real-time detection of ammonia leakage. The feature extraction network of the model is reconstructed to reduce the calculation amount during model inference and compress the size of the model weights. The lightweight SK5 Block module is used to construct the feature extraction network of the model to achieve a balance between speed and accuracy. The SK5 Block module is mainly designed with 1×1 ordinary convolution and 3×3 depthwise separable convolution kernels, expanded into a 5×5 depthwise separable convolution (Depthwise Convolution, DWConv) and a channel shuffle structure.
[0117] According to different moving strides (Stride) of the convolutional kernel, two lightweight convolutional modules are designed. One is the lightweight SK5 Block module with a stride of 1 (Stride = 1), as shown in Figure 2 ; the other is the lightweight SK5 Block module with a stride of 2 (Stride = 2), as shown in Figure 3 .
[0118] In the lightweight SK5 Block module with a stride of 1 (Stride = 1), the channel split grouping operation is started. The input feature channels are split into two parts. One branch directly retains its own information and passes it down. The other branch passes through a 1x1 ordinary convolution to improve the detection speed of the model. At the same time, depthwise separable convolution is fused to reduce the computational complexity of model training. Finally, channel shuffle is used to achieve feature interaction between channels, aiming to improve the model accuracy. Splitting the feature channels into two parts in the lightweight convolutional module is beneficial for network parallelism, reducing the number of model parameters and improving the running speed. In the lightweight SK5 Block module with a stride of 2 (Stride = 2), the input feature map is divided into two branches with the same mapping. When outputting, the concat operation is used to integrate the feature map information. The channel shuffle and concat operations can be combined into an element-wise operation, expanding the channel dimension and transmitting information between channels. This operation can effectively improve the generalization of the model and also improve the model rate.
[0119] As shown in Figure 5 , step S430 in the method provided by the embodiment of the present application further includes:
[0120] S431: Slice the input infrared image through the image slicing part, and convert the sliced image sub-blocks into sequences;
[0121] S432: Input the sequence into the detection network model SSANet, and perform position embedding on the image sub-blocks through the data embedding part to obtain the embedded image information, where the position embedding is the position information between the pixel values of the embedded image sub-blocks;
[0122] S433: Add classification tags to the embedded image information, input the image information with classification tags into the encoding layer part, fuse the multi-head attention mechanism for feature extraction to obtain multiple groups of self-attention processed features, and aggregate the multiple groups of attention processed features through a fully connected layer to obtain the multi-scale feature information.
[0123] Specifically, in the case of ammonia leakage, there are phenomena such as the irregular movement of the ammonia cloud and the low contrast and unclear features in the diffusion area in the infrared image, resulting in the defect that the model focuses on local features and ignores global features, with low detection accuracy. Introducing the Transformer module with self-attention mechanism can convert the infrared image into a sequence and then input it into the model for processing. Based on the feature extraction network reconstructed by the SK5 Block module, Transformer is used as the bottom-up bottleneck layer of the model feature pyramid to obtain multi-scale feature information, thus realizing the fusion of global features and local features.
[0124] The Transformer module can capture global information and rich context information and perform global inference on images and specific predicted targets. Applying the Transformer module to low-resolution feature maps can reduce expensive computing and storage costs while enhancing the importance of leakage features and improving the detection accuracy of the model. As Figure 6 shown, the Transformer module is divided into three parts:
[0125] (1) Image slicing part: The Transformer module first divides the input infrared image into N sub-blocks (Patches) of P×P×C, and converts them into N vectors of P 2 C dimensions through a flattening operation. After the image is converted into a sequence, it can be input into the model for processing;
[0126] (2) Data embedding part: After processing the infrared image, in order to avoid the influence of the sub-block size on the model structure, linear mapping is used to convert different flattened sub-blocks into D-dimensional vectors, and position embedding (Position Embedding) is added to avoid losing the position information between image pixel values. The position embedding uses a trainable parameter, and its dimension is the same as that of the image transformation, so as to capture the relationship between pixels in the high-dimensional vector space and reduce the model's dependence on the number of input images;
[0127] (3) Transformer encoding layer part: Under the infrared imaging technology, the black-gray cloud feature presented in the ammonia leakage area has a phenomenon of low contrast and unclear features with the background. The Transformer encoding layer can consider various attention distributions and focus on different aspects of the leakage information, improving the detection accuracy of the model for ammonia leakage in complex scenarios while maintaining less spatio-temporal complexity. Before the image data is input into the Transformer encoding layer (Encoder), a [class] token for classification needs to be added, and the multi-head attention mechanism (Multi-Head Attention) is fused to extract features.
[0128] In the Transformer encoding layer, the SSANet model performs multiple self-attention processes on the original input sequence and then uses two fully connected layers (Linear) to replace the original layer normalization process. These two linear fully connected transformations reduce computational complexity while effectively mitigating the impact of sample batch size. This enriches the feature space extracted by the backbone network by aggregating feature information from different branches.
[0129] In summary, the embodiments of the present application have at least the following technical effects:
[0130] 1. The method provided in an embodiment of the present application obtains an infrared image detection dataset; normalizes the infrared image detection dataset and divides it into a training set and a test set, annotates the infrared image detection dataset with ammonia leak infrared data; constructs a shuffled self-attention network structure; trains the shuffled self-attention network structure using the training set with the annotated ammonia leak infrared data to obtain a shuffled self-attention network model; detects the test set using the shuffled self-attention network model, calls the final shuffled self-attention network model and test program, and inputs an ammonia leak infrared image to determine the detection result. This method solves the technical problems of conventional ammonia leak detection devices in the prior art, such as small detection range, poor sensitivity, low real-time performance, and inability to locate the leak source. It achieves real-time and accurate monitoring of ammonia leaks at a long distance, ensuring the safety of workers and enabling timely response when ammonia leaks occur, thereby reducing economic losses.
[0131] 2. Conventional ammonia leak detection uses a point measurement gas sensor, but its contact principle makes many areas to be tested inaccessible, resulting in problems such as small detection range, poor sensitivity, low real-time performance, and inability to locate the source of the leak, posing a great threat to the safety of workers. The present invention combines infrared imaging detection technology with ammonia leak detection, and proposes a shuffled self-attention network model (SSANet) to achieve infrared real-time non-contact detection of ammonia leaks. Because the ammonia leak images obtained by the infrared thermal imager have high noise and low contrast, this application first uses a non-local mean denoising method to remove randomly distributed noise points in the background area, while better maintaining the local features of the edge pixels of the ammonia cloud area. Then, a contrast-limited adaptive histogram equalization algorithm is used to perform contrast enhancement preprocessing on the ammonia leak infrared image to establish an ammonia leak infrared detection dataset.
[0132] 3. This application uses the K-means clustering algorithm to cluster and calculate the real frame size of the annotation based on the actual spread of ammonia leakage from nothing to something, and analyzes the area and aspect ratio of the ammonia leakage area at different times to obtain candidate frames suitable for infrared detection of ammonia leakage and preset target detection model candidate frame parameters to improve the detection accuracy of the model.
[0133] 4. In view of the fact that ammonia is a colorless, flammable, explosive and toxic gas, and the leakage has the characteristics of unpredictable occurrence time and location, this application poses extremely high requirements for the real-time detection of ammonia. A shuffle self-attention network model (SSANet) is designed. The 3×3 depthwise separable convolution kernel is redesigned as a 5×5 SK5 Block module to reconstruct the feature extraction network, expanding the receptive field and improving the detection accuracy while reducing the model weight. Finally, the Transformer module is used as the bottleneck layer of the model's feature pyramid from bottom to top to achieve the multi-head attention fusion of the leakage area from bottom to top, obtaining multi-scale feature information, thereby realizing the fusion of global features and local features to improve the model's accuracy. The SSANet model is used to train the ammonia leakage infrared detection dataset, and the final ammonia leakage infrared detection model is obtained to detect ammonia leakage. The weight of the SSANet model is reduced by 76.4% compared with the YOLOv5s basic model, dropping to 3.4M; the average detection speed of a single image is increased by 1.1%, reaching 3.2 ms; the average detection accuracy is increased by 3.5%, reaching 96.3%. The present invention realizes the lightweight real-time detection of ammonia leakage infrared, providing an effective real-time detection method for the development of non-contact detection devices for ammonia leakage to ensure the safe production and stable operation of ammonia-related enterprises.
[0134] Embodiment 2
[0135] Based on the same inventive concept as the ammonia leakage shuffle self-attention lightweight infrared detection method in the foregoing embodiment, as Figure 7 shown, this application provides an ammonia leakage shuffle self-attention lightweight infrared detection system, wherein the system includes:
[0136] A first acquisition unit 11, which is used to acquire an infrared image detection dataset;
[0137] A first processing unit 12, which is used to normalize the infrared image detection dataset and divide it into a training set and a test set, and perform ammonia leakage infrared data annotation on the infrared image detection dataset;
[0138] A first construction unit 13, which is used to construct a shuffle self-attention network structure;
[0139] A second processing unit 14, which is used to train the shuffle self-attention network structure with the training set having the ammonia leakage infrared data annotation to obtain a shuffle self-attention network model;
[0140] The third processing unit 15, which is used to detect the test set through the shuffle self-attention network model, call the final shuffle self-attention network model and the test program, and input the infrared image of ammonia leakage to determine the detection result.
[0141] Further, the system further includes:
[0142] A fourth processing unit, which is used to collect the ammonia leakage video through an infrared thermal imager, perform frame extraction processing on the ammonia leakage video, and obtain an infrared image set;
[0143] A fifth processing unit, which is used to perform non-local mean denoising on the infrared image set to obtain a denoised infrared image set;
[0144] A sixth processing unit, which is used to perform adaptive histogram equalization processing based on the denoised infrared image set to obtain the infrared image detection data set.
[0145] Further, the system further includes:
[0146] A seventh processing unit, which is used to perform image size preprocessing on the infrared image detection data in the infrared image detection data set;
[0147] An eighth processing unit, which is used to determine the target detection area based on the infrared image detection data after size preprocessing, and set a bounding box according to the target detection area;
[0148] A ninth processing unit, which is used to generate a label file in a preset data set format according to the bounding box of the target detection area;
[0149] A tenth processing unit, which is used to annotate the label file in the preset data set format, and the annotation information includes the ammonia leakage position of the bounding box and the size of the leakage area;
[0150] An eleventh processing unit, which is used to convert the format of the label file in the preset data set format to determine the ammonia leakage infrared data annotation, where the ammonia leakage infrared data annotation is a first preset format label file and includes the center point coordinates of the bounding box.
[0151] Further, the system further includes:
[0152] A second construction unit, which is used to construct a detection network model SSANet;
[0153] The twelfth processing unit is used to construct the feature extraction network of the detection network model SSANet by using the lightweight SK5 Block module. Among them, the feature extraction network includes 5×5 depthwise separable convolution and channel shuffle structure;
[0154] The thirteenth processing unit is used to add a Transformer module with self-attention mechanism to the feature extraction network as the bottom-up bottleneck layer of the model feature pyramid. Among them, the Transformer module includes an image slicing part, a data embedding part, and an encoding layer part.
[0155] Furthermore, the system further includes:
[0156] The fourteenth processing unit is used to execute step 1: obtain a first sample according to the infrared image detection dataset, and use the first sample as the first initial clustering center;
[0157] The fifteenth processing unit is used to execute step 2: calculate the distance between each data sample in the infrared image detection dataset and the first initial clustering center, select the minimum distance from all sample distances as the first clustering;
[0158] The sixteenth processing unit is used to execute step 3: based on the calculated distances of all samples, determine the maximum distance sample, and use the maximum distance sample as the second clustering center;
[0159] The seventeenth processing unit is used to execute step 4: repeat step 2-step 3 until K clustering centers are obtained, and perform clustering calculation on the K clustering centers to obtain the annotation box size;
[0160] The eighteenth processing unit is used to execute step 5: obtain ammonia leakage size information according to the annotation box size. The ammonia leakage size information is the area and aspect ratio information of the ammonia leakage area at different times, and use the ammonia leakage size information as the candidate box parameters of the detection network model SSANet.
[0161] Furthermore, the system further includes:
[0162] The nineteenth processing unit is used to perform annotation box size clustering on the ammonia leakage infrared dataset in the training set by using K clustering centers to obtain a clustering result, and use the clustering result as the model candidate box parameters;
[0163] The twentieth processing unit, which is used to extract features from the infrared data of ammonia leakage through the feature extraction network to obtain feature detection and extraction information;
[0164] The twenty-first processing unit, which is used to perform bottom-up fusion on the feature detection and extraction information through the Transformer module to obtain multi-scale feature information;
[0165] The twenty-second processing unit, which is used to train the detection network model SSANet using each piece of infrared data of ammonia leakage in the training set to obtain the shuffled self-attention network model, and the shuffled self-attention network model is used to extract the multi-scale feature information to detect ammonia leakage information.
[0166] Furthermore, the system further includes:
[0167] The twenty-third processing unit, which is used to divide the input feature map channels into a first branch and a second branch when it is the first-step lightweight convolution module. The first branch retains its own information and passes it down to obtain the first-channel feature. The second branch passes through a 1×1 convolution channel and is fused with a 5×5 depthwise separable convolution to obtain the second-channel feature;
[0168] The twenty-fourth processing unit, which is used to shuffle based on the first-channel feature and the second-channel feature through the channel shuffle structure to obtain the feature detection and extraction information;
[0169] The twenty-fifth processing unit, which is used to divide the input feature map into two branches with the same mapping to obtain two branch features when it is the second-step lightweight convolution module, shuffle the two branch features, and output the feature detection and extraction information.
[0170] Furthermore, the system further includes:
[0171] The twenty-sixth processing unit, which is used to slice the input infrared image through the image slicing part and convert the sliced image sub-blocks into sequences;
[0172] The twenty-seventh processing unit, which is used to input the sequence into the detection network model SSANet and perform position embedding on the image sub-blocks through the data embedding part to obtain the embedded image information, where the position embedding is the position information between the pixel values of the embedded image sub-blocks;
[0173] The twenty-eighth processing unit is configured to add classification tags to the post-embedded image information, input the image information with added classification tags into the encoding layer part, perform feature extraction by integrating the multi-head attention mechanism to obtain multiple groups of self-attention processing features, and aggregate the multiple groups of attention processing features through a fully connected layer to obtain the multi-scale feature information.
[0174] Embodiment III
[0175] Based on the same inventive concept as the ammonia leakage shuffling self-attention lightweight infrared detection method in the foregoing embodiment, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method in Embodiment I is implemented.
[0176] Exemplary electronic device
[0177] Reference is made below Figure 8 to describe the electronic device of the present application.
[0178] Based on the same inventive concept as the ammonia leakage shuffling self-attention lightweight infrared detection method in the foregoing embodiment, the present application also provides an ammonia leakage shuffling self-attention lightweight infrared detection system, including: a processor, the processor is coupled to a memory, and the memory is used to store a program. When the program is executed by the processor, the system is enabled to execute the steps of the method described in Embodiment I.
[0179] The electronic device 300 includes: a processor 302, a communication interface 303, and a memory 301. Optionally, the electronic device 300 may further include a bus architecture 304. Among them, the communication interface 303, the processor 302, and the memory 301 may be interconnected through the bus architecture 304; the bus architecture 304 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus architecture 304 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0180] The processor 302 may be a CPU, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present application solution.
[0181] A communication interface 303, using any device such as a transceiver, is used to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), wired access networks, etc.
[0182] The memory 301 can be a ROM or other types of static storage devices that can store static information and instructions, a RAM or other types of dynamic storage devices that can store information and instructions, or it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor through the bus architecture 304. The memory can also be integrated with the processor.
[0183] Among them, the memory 301 is used to store computer execution instructions for implementing the solution of this application, and is controlled by the processor 302 for execution. The processor 302 is used to execute the computer execution instructions stored in the memory 301, so as to implement a method for lightweight infrared detection of ammonia leakage with shuffled self-attention provided in the above embodiments of this application.
[0184] Those of ordinary skill in the art can understand that: The various digital numbers such as the first and the second involved in this application are only for the convenience of description and are not used to limit the scope of this application, nor do they represent the order of precedence. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one" means one or more. At least two means two or more. "At least one", "any one", or their similar expressions refer to any combination of these items, including any combination of single item (s) or plural item (s). For example, at least one (item, kind) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0185] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wire (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state disk (SSD)).
[0186] The various illustrative logical units and circuits described in this application can be implemented or operate the described functions by a design of a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of the above. The general-purpose processor can be a microprocessor. Optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.
[0187] The steps of the methods or algorithms described in this application can be directly embedded in hardware, software units executed by a processor, or a combination of the two. The software units can be stored in a RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be provided in an ASIC, and the ASIC can be provided in a terminal. Optionally, the processor and the storage medium can also be provided in different components of the terminal. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, thereby providing instructions executed on the computer or other programmable device for implementing the steps in Figure 1 one process or multiple processes and / or blocks Figure 1 the steps of the functions specified in one block or multiple blocks.
[0188] Although this application has been described in conjunction with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of this application. Accordingly, this specification and the drawings are only exemplary descriptions of this application and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art can make various changes and modifications to this application without departing from the scope of this application. Thus, if these modifications and variations of this application fall within the scope of this application and its equivalent technologies, this application is intended to include these changes and modifications.
Claims
1. An ammonia leakage mixed self-attention lightweight infrared detection method, characterized in that The method includes: Obtaining an infrared image detection data set; Normalizing the infrared image detection data set and then dividing it into a training set and a test set, and performing ammonia leakage infrared data annotation on the infrared image detection data set; Constructing a shuffled self-attention network structure; Training the shuffled self-attention network structure with the training set having the ammonia leakage infrared data annotation to obtain a shuffled self-attention network model; Detecting the test set through the shuffled self-attention network model, calling the final shuffled self-attention network model and the test program, and inputting an ammonia leakage infrared image to determine the detection result; The constructing of the shuffled self-attention network structure includes: Constructing a detection network model; Using a lightweight SK5 Block module to construct a feature extraction network of the detection network model, where the feature extraction network includes a 5×5 depthwise separable convolution and a channel shuffle structure; Adding a Transformer module with a self-attention mechanism to the feature extraction network as the bottleneck layer from bottom to top of the model feature pyramid, where the Transformer module includes an image slicing part, a data embedding part, and an encoding layer part; Before constructing the detection network model, it includes: Step 1: According to the infrared image detection data set, obtaining a first sample and using the first sample as the first initialization clustering center; Step 2: Calculating the distance between each data sample in the infrared image detection data set and the first initialization clustering center, and selecting the minimum distance from all sample distances as the first cluster; Step 3: Based on the calculated distances of all samples, determining the maximum distance sample and using the maximum distance sample as the second clustering center; Step 4: Repeating Step 2 - Step 3 until K clustering centers are obtained, and performing clustering calculation on the K clustering centers to obtain the annotation box size; Step 5: According to the annotation box size, obtaining ammonia leakage size information, where the ammonia leakage size information is the area and aspect ratio information of the ammonia leakage area at different times, and using the ammonia leakage size information as the candidate box parameters of the detection network model; The training of the shuffled self-attention network structure with the training set having the ammonia leakage infrared data annotation to obtain a shuffled self-attention network model includes: Using K clustering centers to perform annotation box size clustering on the ammonia leakage infrared data set in the training set to obtain a clustering result, and using the clustering result as the model candidate box parameters; Performing feature extraction on the ammonia leakage infrared data through the feature extraction network to obtain feature detection extraction information; Performing bottom-up fusion on the feature detection extraction information through the Transformer module to obtain multi-scale feature information; Training the detection network model with each ammonia leakage infrared data in the training set to obtain the shuffled self-attention network model, where the shuffled self-attention network model is used to extract the multi-scale feature information to detect ammonia leakage information.
2. The method according to claim 1, wherein The obtaining of the infrared image detection data set includes: Collect the ammonia leakage video through an infrared thermal imager, perform frame extraction on the ammonia leakage video to obtain an infrared image set; Perform non-local mean denoising on the infrared image set to obtain a denoised infrared image set; Perform adaptive histogram equalization processing based on the denoised infrared image set to obtain the infrared image detection data set.
3. The method according to claim 1, characterized in that, Perform ammonia leakage infrared data annotation on the infrared image detection data set, including: Perform image size preprocessing on the infrared image detection data in the infrared image detection data set; Based on the infrared image detection data after size preprocessing, determine the target detection area and set a bounding box according to the target detection area; Generate a label file in a preset data set format according to the bounding box of the target detection area; Annotate the label file in the preset data set format, and the annotation information includes the ammonia leakage position of the bounding box and the size of the leakage area; Convert the format of the label file in the preset data set format to determine the ammonia leakage infrared data annotation, where the ammonia leakage infrared data annotation is a first preset format label file and includes the center point coordinates of the bounding box.
4. The method according to claim 1, wherein The feature extraction network includes a first-step lightweight convolution module and a second-step lightweight convolution module. The feature extraction network performs feature extraction on the ammonia leakage infrared data to obtain feature detection extraction information, including: When it is the first-step lightweight convolution module, divide the input feature map channels into a first branch and a second branch. The first branch retains its own information and passes it down to obtain the first channel feature. The second branch passes through a 1×1 convolution channel and is fused with a 5×5 depthwise separable convolution to obtain the second channel feature; Perform shuffling on the first channel feature and the second channel feature through the channel shuffling structure to obtain the feature detection extraction information; When it is the second-step lightweight convolution module, divide the input feature map into two branches with equal mapping to obtain two branch features, shuffle the two branch features, and output the feature detection extraction information.
5. The method according to claim 1, characterized in that, The self-bottom-up fusion of the feature detection extraction information through the Transformer module to obtain multi-scale feature information includes: Slice the input infrared image through the image slicing part and convert the sliced image sub-blocks into sequences; Input the sequence into the detection network model, and perform position embedding on the image sub-blocks through the data embedding part to obtain the embedded image information, where the position embedding is the position information between the pixel values of the embedded image sub-blocks; Add a classification mark to the embedded image information, input the image information with the classification mark added into the encoding layer part, fuse the multi-head attention mechanism for feature extraction to obtain multiple groups of self-attention processed features, and aggregate the multiple groups of attention processed features through a fully connected layer to obtain the multi-scale feature information.
6. An ammonia leakage mixed self-attention lightweight infrared detection system, characterized in that, The system includes: A first obtaining unit, and the first obtaining unit is used to obtain an infrared image detection data set; A first processing unit, which is used to normalize the infrared image detection dataset and then divide it into a training set and a test set, and perform ammonia leakage infrared data annotation on the infrared image detection dataset; A first construction unit, which is used to construct a shuffled self-attention network structure; A second processing unit, which is used to train the shuffled self-attention network structure by using the training set with the ammonia leakage infrared data annotation to obtain a shuffled self-attention network model; A third processing unit, which is used to detect the test set through the shuffled self-attention network model, call the final shuffled self-attention network model and the test program, and input the ammonia leakage infrared image to determine the detection result; The system further includes: A second construction unit, which is used to construct a detection network model; A twelfth processing unit, which is used to construct a feature extraction network of the detection network model by using a lightweight SK5 Block module, where the feature extraction network includes a 5×5 depthwise separable convolution and a channel shuffle structure; A thirteenth processing unit, which is used to add a Transformer module with a self-attention mechanism to the feature extraction network as a bottleneck layer from bottom to top of the model feature pyramid, where the Transformer module includes an image slicing part, a data embedding part, and an encoding layer part; The system further includes: A fourteenth processing unit, which is used to execute step 1: obtain a first sample according to the infrared image detection dataset and use the first sample as a first initial clustering center; A fifteenth processing unit, which is used to execute step 2: calculate the distance between each data sample in the infrared image detection dataset and the first initial clustering center, select the minimum distance from all sample distances as the first cluster; A sixteenth processing unit, which is used to execute step 3: determine the maximum distance sample based on the calculated distances of all samples and use the maximum distance sample as the second clustering center; A seventeenth processing unit, which is used to execute step 4: repeat steps 2 - 3 until K clustering centers are obtained, and perform clustering calculation on the K clustering centers to obtain the annotation box size; An eighteenth processing unit, which is used to execute step 5: obtain the ammonia leakage size information according to the annotation box size, where the ammonia leakage size information is the area and aspect ratio information of the leakage area of ammonia leakage at different times, and use the ammonia leakage size information as the candidate box parameter of the detection network model; The system further includes: A nineteenth processing unit, which is used to perform annotation box size clustering on the ammonia leakage infrared dataset in the training set by using K clustering centers to obtain a clustering result, and use the clustering result as the model candidate box parameter; The twentieth processing unit, which is used to extract features from the ammonia leakage infrared data through the feature extraction network to obtain feature detection and extraction information; The twenty-first processing unit, which is used to perform bottom-up fusion on the feature detection and extraction information through the Transformer module to obtain multi-scale feature information; The twenty-second processing unit, which is used to train the detection network model by using each ammonia leakage infrared data in the training set to obtain the shuffled self-attention network model, and the shuffled self-attention network model is used to extract the multi-scale feature information to detect ammonia leakage information.
7. An ammonia leakage mixed self-attention lightweight infrared detection system, characterized in that, Comprising: A processor, the processor is coupled to a memory, and the memory is used to store a program, and when the program is executed by the processor, the system is caused to execute the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Submarine pipeline and leakage point detection method
CN112085728A
Image target object real-time detection method and system, terminal and storage medium
CN113222064A