Target enhancement detection method and device for oil tank detection
Through the combination of the image fusion model and the YOLOv11 model, the problem of high misjudgment rate and leakage detection in oil tank detection is solved, and high-precision oil tank detection is achieved.
Patent Information
- Application Number
- CN202510418150.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems with high misjudgment rate and leakage detection rate in oil tank detection under complex backgrounds, especially detection performance bottlenecks caused by the noise of synthetic aperture radar image due to the limitations of the meteorological conditions.
The image fusion model is used to fusion and enhance optical images and synthetic aperture radar images. The DenseNet structure is used to extract features and input the YOLOv11 model for target detection. The shallow and deep features are fused by the feature extractor for information measurement.
The organic fusion of optical image scene details and synthetic aperture radar image scattering characteristics is achieved, the error judgment rate and leakage detection rate are reduced, and the accuracy of oil tank detection is improved.
Smart Images

Figure CN120339835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to a target enhancement detection method and device for oil tank detection. Background Art
[0002] Multi-source remote sensing image fusion is an important technical means to assist in obtaining high-quality ground object observation data. Remote sensing earth observation is the core key technology for environmental protection. With the rapid development of remote sensing technology, rich multi-source remote sensing imaging means such as multi-spectral, hyperspectral, synthetic aperture radar, lidar, and thermal infrared have played important roles in many fields. Multi-source heterogeneous remote sensing data fusion aims to overcome the limitations of a single sensor in terms of observation content and acquisition time, and comprehensively utilize multi-dimensional observation information to interpret the observation scene more accurately.
[0003] Target detection is a research hotspot in the field of remote sensing image processing. Its core task is to accurately locate the spatial position of specific targets in the image and achieve efficient classification and recognition, providing key technical support for the intelligent analysis of remote sensing images. However, with the development of earth observation technology towards sub-meter high resolution, the complex and variable imaging conditions and ground object background interference have significantly increased the difficulty of target detection. At the same time, the inherent physical imaging characteristics of different remote sensing sensors lead to limitations in the representation of single-source data: although optical images have rich spectral texture characteristics that conform to human visual perception, their imaging quality is easily restricted by meteorological conditions such as clouds, fog, rain, and snow; the complementary synthetic aperture radar images have all-weather penetration ability and can effectively obtain special scene information such as vegetation coverage and camouflaged targets, but are difficult to visually interpret due to speckle noise. The inherent deficiency in the representation ability of this single-modal data makes the target detection method based on a single sensor face significant performance bottlenecks in complex scenes. The multi-source remote sensing image fusion technology can organically integrate the detail representation advantages of optical images and the physical penetration characteristics of synthetic aperture radar images through the information coordination mechanism at the feature level and decision level, not only significantly improving the interpretability of the images, but also constructing a multi-dimensional target feature expression system, laying a data foundation for highly robust detection algorithms. Although existing research has verified the enhancement effect of visible light-infrared fusion on target detection, the research on detection technology for optical-synthetic aperture radar cross-modal fusion is still in its infancy, and there are still theoretical method gaps in key links such as complex background suppression and multi-scale feature preservation, which urgently need in-depth exploration and innovative breakthroughs. Summary of the Invention
[0004] The present invention provides a target enhancement detection method and device for oil tank detection, aiming to improve the accuracy of target detection.
[0005] To achieve the above object, the present invention provides a target enhancement detection method for oil tank detection, including:
[0006] Step 1, obtain the optical image and synthetic aperture radar image of the target area;
[0007] Step 2, input the optical image and synthetic aperture radar image into the image fusion model for image fusion enhancement to obtain the fused and enhanced image;
[0008] Step 3, input the fused and enhanced image into the trained YOLOv11 model for target detection to obtain the oil tank detection result of the target area;
[0009] The image fusion model includes a feature extractor and a DenseNet structure. The DenseNet structure is used to fuse the optical image and synthetic aperture radar image to obtain the fused and enhanced image. The feature extractor is used to extract the shallow features and deep features of the optical image and synthetic aperture radar image, and the shallow features and deep features are used for information measurement between the fused and enhanced image and the optical image and synthetic aperture radar image.
[0010] Furthermore, before inputting the optical image and synthetic aperture radar image into the image fusion model for image fusion enhancement, it also includes:
[0011] Preprocess the synthetic aperture radar image to obtain the preprocessed synthetic aperture radar image.
[0012] Furthermore, preprocessing the synthetic aperture radar image to obtain the preprocessed synthetic aperture radar image includes:
[0013] Use the Gamma filtering algorithm to suppress the noise of the synthetic aperture radar image to obtain the denoised synthetic aperture radar image;
[0014] Perform format conversion on each band image of the denoised synthetic aperture radar image, and during the format conversion process, perform normalization and enhancement processing to obtain multiple single-channel grayscale images as the preprocessed synthetic aperture radar image.
[0015] Furthermore, before inputting the optical image and synthetic aperture radar image into the image fusion model for image fusion enhancement, it also includes:
[0016] Select an optical RGB image with a resolution of 0.5m and a synthetic aperture radar image with a resolution of 0.5m and VV polarization mode as training data to train the unsupervised image fusion network to obtain the image fusion model. The image fusion model includes a feature extractor and a DenseNet structure.
[0017] Furthermore, the DenseNet structure is composed of ten convolutional layers connected in sequence;
[0018] Among them, the LeakyReLU activation function is adopted in the first nine convolutional layers, and the tanh activation function is adopted in the last convolutional layer;
[0019] In the first seven convolutional layers, a short connection is added between every two convolutional layers.
[0020] Furthermore, the loss function of the image fusion model is:
[0021] L(θ,D) = L sim (θ,D) + λL ewc (θ,D)
[0022] where θ represents the parameters of the DenseNet structure, D represents the training data, and L sim (θ,D) represents the similarity loss between the fused enhanced image and the training data, and L ewc (θ,D) represents the learning loss, and λ represents the hyperparameter.
[0023] The present invention also provides a target enhancement detection device for oil tank detection, including:
[0024] An acquisition module for acquiring the optical image and the synthetic aperture radar image of the target area;
[0025] An enhancement module for inputting the optical image and the synthetic aperture radar image into the image fusion model for image fusion enhancement to obtain a fused enhanced image;
[0026] A detection module for inputting the fused enhanced image into the trained YOLOv11 model for target detection to obtain the oil tank detection result of the target area;
[0027] The image fusion model includes a feature extractor and a DenseNet structure. The DenseNet structure is used to fuse the optical image and the synthetic aperture radar image to obtain a fused enhanced image. The feature extractor is used to extract the shallow features and deep features of the optical image and the synthetic aperture radar image. The shallow features and deep features are used for information measurement between the fused enhanced image and the optical image and the synthetic aperture radar image.
[0028] The above scheme of the present invention has the following beneficial effects:
[0029] The optical image and synthetic aperture radar image of the target area obtained by the present invention are input into an image fusion model for image fusion enhancement to obtain a fusion-enhanced image; the fusion-enhanced image is input into a trained YOLOv11 model for target detection to obtain the oil tank detection result of the target area; the image fusion model includes a feature extractor and a DenseNet structure, and the DenseNet structure is used to fuse the optical image and the synthetic aperture radar image to obtain a fusion-enhanced image, and the feature extractor is used to extract the shallow features and deep features of the optical image and the synthetic aperture radar image, and the shallow features and deep features are used for information measurement between the fusion-enhanced image and the optical image and the synthetic aperture radar image; compared with the prior art, the present invention uses an image fusion model to realize the organic fusion of the rich scene details of the high-resolution optical image and the unique scattering characteristics of the synthetic aperture radar image, completes the complementary advantages of multi-modal information, and can more comprehensively express the target features by fusing the two kinds of data, thereby effectively reducing the false judgment rate and the missed detection rate, and further improving the accuracy of target detection.
[0030] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a schematic flow chart of an embodiment of the present invention;
[0032] Figure 2 is a schematic structural diagram of the image fusion model in an embodiment of the present invention;
[0033] Figure 3 is a schematic structural diagram of the DenseNet structure in an embodiment of the present invention;
[0034] Figure 4 is a schematic structural diagram of the C2PSA module in the trained YOLOv11 model;
[0035] Figure 5 is a schematic structural diagram of the detection head in the trained YOLOv11 model;
[0036] Figure 6 is a schematic structural diagram of the target enhancement detection device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the drawings and specific embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] In the description of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0039] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it can be a locking connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0040] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0041] The present invention provides a target enhancement detection method and device for oil tank detection in view of existing problems.
[0042] As Figure 1 、 Figure 2 shown, an embodiment of the present invention provides a target enhancement detection method for oil tank detection, including:
[0043] Step 1, obtaining an optical image and a synthetic aperture radar image of a target area;
[0044] Step 2, inputting the optical image and the synthetic aperture radar image into an image fusion model for image fusion enhancement to obtain a fusion enhanced image;
[0045] Step 3, inputting the fusion enhanced image into a trained YOLOv11 model for target detection to obtain an oil tank detection result of the target area;
[0046] The image fusion model includes a feature extractor and a DenseNet structure. The DenseNet structure is used to fuse the optical image and the synthetic aperture radar image to obtain a fusion enhanced image. The feature extractor is used to extract the shallow features and deep features of the optical image and the synthetic aperture radar image. The shallow features and deep features are used for information measurement between the fusion enhanced image and the optical image and the synthetic aperture radar image.
[0047] It should be noted that the target area mentioned in the embodiments of the present invention can be an oil processing plant or a port oil storage area. The optical image refers to the image data of the oil processing plant or the port oil storage area captured by optical means, and the synthetic aperture radar image is the image data of the oil processing plant or the port oil storage area captured by a high-resolution imaging radar. The purpose of the embodiments of the present invention is to detect or identify oil tanks for storing crude oil, refined oil, liquefied natural gas or chemical products in the image data of the oil processing plant or the port oil storage area.
[0048] In order to remove noise, facilitate computer display, and improve the accuracy of subsequent detection, in the embodiments of the present invention, before inputting the optical image and the synthetic aperture radar image into the image fusion model for image fusion enhancement, it is necessary to preprocess the synthetic aperture radar image to obtain the preprocessed synthetic aperture radar image.
[0049] Specifically, preprocessing the synthetic aperture radar image to obtain the preprocessed synthetic aperture radar image includes:
[0050] Using the Gamma filtering algorithm to suppress noise in the synthetic aperture radar image to effectively reduce speckle noise and other interferences in the image, improve the image quality, and obtain the denoised synthetic aperture radar image;
[0051] Perform format conversion on each band image in the denoised synthetic aperture radar image, convert each band image from the original 32-bit floating-point format to an 8-bit depth single-channel grayscale image, which is convenient for subsequent image processing and analysis. And in order to ensure that the details of the image are retained to the greatest extent, during the format conversion process, normalization and enhancement processing are adopted, so that the image has good performance in both visual effects and subsequent algorithm applications, and obtain multiple single-channel grayscale images as the preprocessed synthetic aperture radar image.
[0052] Through the above preprocessing operations, clearer and more reliable input data can be provided for subsequent target detection and recognition tasks.
[0053] Specifically, before inputting the optical image and the synthetic aperture radar image into the image fusion model for image fusion enhancement, it also includes:
[0054] Select an optical RGB image with a resolution of 0.5m and a synthetic aperture radar image with a resolution of 0.5m and VV polarization mode as training data to train an unsupervised image fusion network to obtain an image fusion model. The image fusion model includes a feature extractor and a DenseNet structure.
[0055] It should be noted that since it is difficult to obtain registered optical and synthetic aperture radar data in the same area at the same time, the training data of the embodiments of the present invention is from the publicly available dataset SpaceNet6, which covers an area of 120 km in a certain port in southern Netherlands. 2 The dataset includes 3,401 spatio-temporally synchronized optical images and synthetic aperture radar images. The optical image dataset includes single-band panchromatic images with a resolution of 0.5 m, multi-spectral images with a resolution of 2 m, panchromatic sharpened images with a resolution of 0.5 m, and RGB images. The synthetic aperture radar image dataset includes X-band datasets with four polarization modes of HH, HV, VH, and VV with a resolution of 0.5 m. According to the experimental requirements, the embodiments of the present invention screen out the images containing oil tank targets from the dataset, and finally select optical RGB images with a resolution of 0.5 m and synthetic aperture radar images with a resolution of 0.5 m and VV polarization mode as the training data to train the unsupervised image fusion network.
[0056] The unsupervised image fusion network in the embodiments of the present invention aims to achieve efficient fusion of various types of images through deep learning technology, and its structure is as Figure 2 shown. This network is particularly suitable for the case where there is no paired training data. Through adaptive feature extraction and fusion strategies, it generates high-quality fused images. It uses a unified model and unified parameters to solve different functional problems, is applicable to various image fusion tasks, and overcomes the common obstacle in most image fusion problems, that is, the lack of common labels and no-reference metrics, by constraining the similarity between the fused image and the source images.
[0057] In the stage of training the unsupervised image fusion network in the embodiments of the present invention, a feature extractor is used to extract five layers of shallow features and deep features from the images. Each layer of features comes from the convolutional layer before the max-pooling layer and is used to represent information at different levels. Among them, the shallow features contain local information such as texture and edges, and the deep features contain deep information such as structure and content. These features are used to measure the information between the fused enhanced image and the optical image and the synthetic aperture radar image to judge the importance of the optical image and the synthetic aperture radar image and determine the weights during fusion.
[0058] The embodiments of the present invention use gradient calculation to evaluate the information richness of the feature maps, and the calculation expression is:
[0059]
[0060] where g I represents the information richness, represents the k-th feature map of the j-th layer, represents the Laolacian gradient calculation, H j 、W j 、Dj Denote the height, width, and number of channels of the feature map;
[0061] Based on this information richness, calculate the adaptive information protection degree for controlling the similarity between the fused enhanced image and the optical image and the synthetic aperture radar image. The calculation expression is:
[0062]
[0063] Where, v1 and v2 represent the adaptive information protection degree, c represents a predefined constant for balancing the numerical range of information measurement. The higher v1 is, the more important the information of the source image I1 (i.e., the optical image) is, and it should be closer to I1 during fusion. The higher v2 is, the more important the information of the source image I2 (i.e., the synthetic aperture radar image) is, and it should be closer to I2 during fusion.
[0064] These information protection degree values are used to define the loss function, and then the loss function is optimized through the DenseNet structure. In the test stage, only the DenseNet fusion step is required, and there is no need to process the information protection degree values. This design avoids the dependence on labels.
[0065] In the test stage of the unsupervised image fusion network according to the embodiments of the present invention, the DenseNet structure is used for image fusion. Its input is the data (I1, I2) after splicing the optical image and the synthetic aperture radar image, and the output is the fused image.
[0066] Specifically, as Figure 3 shown, the DenseNet structure consists of ten consecutive convolutional layers. The convolutional kernel size of each convolutional layer is 3×3, the stride is 1, and Padding adopts the mirroring method;
[0067] In order to retain more original image information, no pooling layer is used. Therefore, in the DenseNet structure of the embodiments of the present invention, the first nine convolutional layers all adopt the LeakyReLU activation function with a slope set to 0.2, and the last convolutional layer adopts the tanh activation function;
[0068] Among the first seven convolutional layers, a short connection is added between every two convolutional layers to improve the training efficiency and reduce the occurrence of the gradient vanishing problem. The number of feature map channels is set to 44, and the number of channels of the last 4 layers decreases layer by layer until it drops to 1.
[0069] Specifically, the loss function of the image fusion model is:
[0070] L(θ,D)=L sim (θ,D)+λL ewc (θ,D)
[0071] Among them, θ represents the parameters of the DenseNet structure, D represents the training data, and L sim (θ, D) represents the similarity loss between the fused enhanced image and the training data, and L ewc (θ, D) represents the learning loss, and λ represents the hyperparameter.
[0072] In the embodiments of the present invention, the similarity loss between the fused enhanced image and the training data consists of two aspects: structural similarity and intensity distribution. Among them, the structural similarity index measure (SSIM) is calculated from the similarity of information such as light, contrast, and structure. The DenseNet structure uses SSIM to constrain the structural similarity between I1, I2, and I f The information degrees are controlled by ω1 and ω2. The expression of the loss function is:
[0073]
[0074] Among them, ω1 and ω2 represent hyperparameters.
[0075] In the embodiments of the present invention, the training data used for model training is constructed by registering and slicing optical and synthetic aperture radar data of the same area, similar time phases, and the same resolution. The image size is 500×520 pixels. A total of 652 pairs of registered images are obtained for network training to minimize the information loss between the fused image and the source image, so as to ensure that the fused result adaptively retains the key information of the source image. The network parameters are optimized through the backpropagation algorithm, and an image fusion model is trained.
[0076] The optical image and synthetic aperture radar image of the target area in the embodiments of the present invention are input into the image fusion model for image fusion enhancement to obtain a fused enhanced image, which not only retains the scene details in the optical image but also combines the scattering characteristics of the synthetic aperture radar image, thereby realizing the complementary enhancement of visual information and structural information.
[0077] Before training the YOLOv11 model in the embodiments of the present invention, it is also necessary to automatically annotate the training data using X-Anylabeling to improve the annotation efficiency. The annotation process is as follows:
[0078] First, access the GitHub repository to download the X-Anylabeling installation package, and configure the corresponding dependency environment according to the instruction document in the repository. After completing the environment setup, enter python anylabeling / app.py in the terminal to start the application;
[0079] Next, import the folder of images to be labeled, and load the YOLOv11 model that has been trained on the DOTA-v1.0 dataset. Select the "Multiple Labeling Mode" to achieve batch automatic labeling, so that all the images in the folder can be automatically labeled at once. The labeling process will combine the existing model for rapid calibration. However, due to the limitations of automated labeling, errors or missed labels may occur. Therefore, after the preliminary labeling is completed, manually check and modify and edit the inaccurate or missed targets to ensure the accuracy and consistency of the labeling.
[0080] After the inspection is completed, click the "Export" button and select a suitable export format (such as YOLO format) to save the labeled result file for subsequent model training.
[0081] It should be noted that the YOLOv11 model adopted in the embodiments of the present invention is the latest one-stage object detection algorithm in the YOLO series, with significant improvements in architecture and training methods, aiming to provide higher detection accuracy, faster processing speed, and more efficient computing performance. At the same time, it supports the rotated object detection task. The main improvements include two parts. First, the C2PSA module is proposed, which is a module that combines the multi-head attention mechanism and is used to enhance the feature extraction ability. The specific structure diagram is as Figure 4 shown; secondly, two depthwise convolution operations (DWConv) are added to the classification detection head in the original decoupled head, as Figure 5 shown. Different from the standard convolution, the depthwise convolution processes each channel of the input separately, that is, each channel has its own convolution kernel for convolution. This operation can reduce the computational complexity and the number of parameters, achieving a faster processing speed.
[0082] While maintaining the accuracy, the YOLOv11 model reduces the number of parameters by 22% compared to the YOLOv8 model, effectively improving the computing efficiency. Whether in computer vision tasks such as object detection, instance segmentation, image classification, pose estimation, or oriented bounding box detection (OBB), the YOLOv11 model can demonstrate excellent performance.
[0083] The loss function of the YOLOv11 model is mainly composed of classification loss and regression loss. The BCE (Binary Cross Entropy) is used as the classification loss, and each category judges "whether it is this category" and outputs the confidence. The expression of the loss function is:
[0084]
[0085] Simply put, BCE calculates the entropy of two pieces of information (is this category, is not this category) and calculates them together, so that the losses calculated by this loss function in the above two cases are not zero.
[0086] Use the method of CIOU loss plus Distribution Focal Loss (DFL) in terms of regression loss. After introducing Anchor-Free Center-based methods (based on the center point) compared with YOLOv5 in CIOU loss, the model changes from outputting "anchor box size offset" to "predicting the distances lrtb = left, top, right, bottom from the left, top, right, and bottom borders of the target box to the center point of the target". And DFL, in the form of cross-entropy, optimizes the probabilities of the two positions on the left and right closest to the label y, so that the network can focus on the distribution of the target position and its adjacent areas faster. The expression is:
[0087] DFL(S i ,S i+1 ) = -((y i+1 - y) log(S i ) + (y - y i ) log(S i+1 ))
[0088] Where S i , S i+1 are the "predicted values" and "adjacent predicted values" output by the network. y, y i , y i+1 are the "actual values", "label integral values", and "adjacent label integral values" of the label.
[0089] In the embodiments of the present invention, 70% of the sample data is selected as the training sample, and the separate optical image, the separate synthetic aperture radar image, and the fused enhanced image are used as inputs respectively. Based on the Windows 11 operating system, the deep learning development framework PyTorch 2.0.1 (CUDA 11.7) and the development language Python 3.9.21 are used. In terms of hardware, an AMD Ryzen 9 5900X 12-core processor and an NVIDIA GeForce RTX 3060 GPU are adopted. During the training process, the parameter settings are as follows: the batch size is set to 16, the initial learning rate is 0.0001, and the number of training epochs is 100. Taking minimizing the classification loss and the regression loss as the optimization goal, the network parameters are optimized through the backpropagation algorithm, and the trained YOLOv11 model is obtained through separate training.
[0090] Next, the embodiments of the present invention further illustrate the effectiveness and accuracy of the provided method in combination with the image quality evaluation index and the target detection model evaluation index. The process is as follows:
[0091] First, in order to evaluate the information content of the fused image compared to single optical images and synthetic aperture radar images, the following quality evaluation metrics are selected to quantitatively evaluate the experiment. The quality evaluation metrics include Natural Image Quality Evaluation (NIQE), variance (Var), entropy (EN), and edge intensity (EI).
[0092] Among them, NIQE (Natural Image Quality Evaluation) does not require a training model and is directly evaluated through statistical characteristics. A smaller value indicates better quality. The image entropy represents the average information content of the image. The larger the entropy value, the more abundant the information contained in the image and the better the quality. The calculation formula is as follows:
[0093]
[0094] where L represents the number of gray levels, and p i is the distribution probability of each gray level;
[0095] The edge intensity EI can reflect the clarity of the image. The larger the value, the better the image quality. The calculation formula is as follows:
[0096]
[0097] where W represents the width of the image and H represents the height of the image.
[0098] s x = F * h x , s y = F * h y
[0099] where h x , h y are the Sobel operators in the x and y directions; s x , s y are the results after convolution with the Sobel operator.
[0100] Var (variance) measures the distribution of image gray values. The larger the value, the richer the contrast and details. The formula is:
[0101]
[0102] At the same time, accuracy, recall, and mAP50 are selected to evaluate the performance of the model in oil tank detection. The formulas are as follows:
[0103]
[0104] where P represents the detection accuracy, R represents the recall, TP represents the number of true positive samples, FP represents the number of false positive samples, FN represents the number of false negative samples, and n represents the total number of detection target categories.
[0105] According to the analysis of the quantitative evaluation indicators in Table 1 below, the enhanced fusion image shows advantages in many aspects:
[0106] (1) Contrast and detail performance (Var): The variance (Var) of the fusion image is higher than that of the synthetic aperture radar image, indicating that its contrast and detail richness are better, and it can present target features more clearly, improving visual recognition.
[0107] (2) Information content (EN) and clarity (EI): The information entropy (EN) and edge intensity (EI) of the fusion image are both better than those of the optical image, indicating that while maintaining scene details, the fusion image also enhances the expression ability of target features, which is helpful for subsequent target detection tasks.
[0108] (3) Naturalness and visual effect (NIQE): The fusion image performs best in the NIQE index with the lowest value, meaning that its visual quality is better, the overall effect is more natural, and the distortion degree is lower, making it more intuitive and effective for human eye observation and computer processing.
[0109] In summary, the fusion image is superior to the optical image in terms of NIQE, information content (EN), and clarity (EI), and is close to or better than the synthetic aperture radar image in terms of information entropy and edge intensity. In addition, it effectively reduces the noise of the synthetic aperture radar image and increases the information content of the optical image, demonstrating the advantages of multi-source remote sensing data fusion.
[0110] Table 1 Quantitative comparison of data before and after fusion using four quality evaluation indicators
[0111] Data source\Indicator Var(↑) EN(↑) EI(↑) NIQE(↓) Optical image 106.87 6.13 52.10 5.74 Synthetic aperture radar image 62.66 6.47 85.46 10.50 Enhanced fusion image 87.78 6.43 54.30 5.31
[0112] The evaluation of the training results of the YOLOv11 model based on different data sources is shown in Table 2 below:
[0113] Table 2 Comparison of model training results based on different data sources
[0114] Data source Precision / % Recall / % mAP@50 / % Optical image 96.5 98.6 98.7 Synthetic aperture radar image 76.3 53.9 67.3 Enhanced fusion image 95.2 98.1 99.0
[0115] It can be analyzed from Table 2 above that:
[0116] (1) The model trained with the fusion image has the best comprehensive performance, and the mAP@50 (99.0%) is higher than that of the optical image (98.7%) and the synthetic aperture radar image (67.3%). This shows that the model trained with the enhanced fusion image performs best in the target detection task, and the precision and recall are also close to those of the model trained with the optical image, indicating that multi-source data fusion can effectively combine the advantages of optical images and synthetic aperture radar images to improve detection performance.
[0117] (2) The model trained with synthetic aperture radar images has poor performance, with both Precision and Recall being low, indicating that single synthetic aperture radar data has certain limitations in target detection.
[0118] (3) The fused image combines the advantages of optical images and synthetic aperture radar images, approaching or even outperforming optical images in all metrics, while overcoming the problem of weak recognition ability of synthetic aperture radar images and achieving overall performance optimization.
[0119] In the embodiment of the present invention, the obtained optical image and synthetic aperture radar image of the target area are input into an image fusion model for image fusion enhancement to obtain a fused and enhanced image; the fused and enhanced image is input into a trained YOLOv11 model for target detection to obtain the oil tank detection result of the target area; the image fusion model includes a feature extractor and a DenseNet structure. The DenseNet structure is used to fuse the optical image and the synthetic aperture radar image to obtain a fused and enhanced image, and the feature extractor is used to extract the shallow features and deep features of the optical image and the synthetic aperture radar image. The shallow features and deep features are used for information measurement between the fused and enhanced image and the optical image and the synthetic aperture radar image. Compared with the prior art, the present invention uses an image fusion model to achieve an organic fusion of the rich scene details of high-resolution optical images and the unique scattering characteristics of synthetic aperture radar images, completes the complementary advantages of multi-modal information, and can more comprehensively express target features by fusing the two types of data, thereby effectively reducing the false positive rate and missed detection rate, and further improving the accuracy of target detection.
[0120] Corresponding to the target enhancement detection method for oil tank detection described in the above embodiment, as Figure 6 shown, the embodiment of the present invention also provides a target enhancement detection device 100 for oil tank detection. The target enhancement detection device 100 includes:
[0121] An acquisition module 101, configured to acquire an optical image and a synthetic aperture radar image of a target area;
[0122] An enhancement module 102, configured to input the optical image and the synthetic aperture radar image into an image fusion model for image fusion enhancement to obtain a fused and enhanced image;
[0123] A detection module 103, configured to input the fused and enhanced image into a trained YOLOv11 model for target detection to obtain the oil tank detection result of the target area;
[0124] The image fusion model includes a feature extractor and a DenseNet structure. The DenseNet structure is used to fuse the optical image and the synthetic aperture radar image to obtain a fused enhanced image. The feature extractor is used to extract the shallow features and deep features of the optical image and the synthetic aperture radar image. The shallow features and deep features are used for information measurement between the fused enhanced image and the optical image and the synthetic aperture radar image.
[0125] It should be noted that for the content such as information interaction and execution process between the above-mentioned devices / units, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.
[0126] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment, and details will not be elaborated here.
[0127] The above is the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A target enhancement detection method for oil tank detection, characterized in that, Including: Step 1: Obtain the optical image and synthetic aperture radar (SAR) image of the target area; Step 2: Input the optical image and the SAR image into the image fusion model for image fusion enhancement to obtain the fused and enhanced image; Step 3: Input the fused and enhanced image into the trained YOLOv11 model for target detection to obtain the oil tank detection result of the target area; The image fusion model includes a feature extractor and a DenseNet structure. The DenseNet structure is used to fuse the optical image and the SAR image to obtain the fused and enhanced image. The feature extractor is used to extract the shallow features and deep features of the optical image and the SAR image, and the shallow features and the deep features are used for information measurement between the fused and enhanced image and the optical image and the SAR image.
2. The target enhancement detection method for oil tank detection according to claim 1, wherein Before inputting the optical image and the SAR image into the image fusion model for image fusion enhancement, it further includes: Preprocess the SAR image to obtain the preprocessed SAR image.
3. The target enhancement detection method for oil tank detection according to claim 2, wherein Preprocessing the SAR image to obtain the preprocessed SAR image includes: Use the Gamma filtering algorithm to suppress the noise of the SAR image to obtain the denoised SAR image; Perform format conversion on each band image in the denoised SAR image, and during the format conversion process, adopt normalization and enhancement processing to obtain multiple single-channel grayscale images as the preprocessed SAR image.
4. The target enhancement detection method for oil tank detection according to claim 2, characterized in that, Before inputting the optical image and the SAR image into the image fusion model for image fusion enhancement, it further includes: Select an optical RGB image with a resolution of 0.5m and a SAR image with a resolution of 0.5m and VV polarization mode as training data to train an unsupervised image fusion network to obtain the image fusion model, and the image fusion model includes a feature extractor and a DenseNet structure.
5. The target enhancement detection method for oil tank detection according to claim 4, characterized in that, The DenseNet structure consists of ten convolutional layers connected in sequence; Among them, the first nine convolutional layers all use the LeakyReLU activation function, and the last convolutional layer uses the tanh activation function; Among the first seven convolutional layers, a short connection is added between every two convolutional layers.
6. The target enhancement detection method for oil tank detection according to claim 5, characterized in that, The loss function of the image fusion model is: L(θ, D) = L sim (θ, D) + λL ewc (θ, D) Among them, θ represents the parameters of the DenseNet structure, D represents the training data, and L sim (θ, D) represents the similarity loss between the fused enhanced image and the training data, and L ewc (θ, D) represents the learning loss, and λ represents the hyperparameter.
7. An object enhancement detection device for oil tank detection, characterized in that, Including: An acquisition module, used to obtain the optical image and the SAR image of the target area; An enhancement module, used to input the optical image and the SAR image into the image fusion model for image fusion enhancement to obtain the fused and enhanced image; A detection module, used to input the fused and enhanced image into the trained YOLOv11 model for target detection to obtain the oil tank detection result of the target area; The image fusion model includes a feature extractor and a DenseNet structure. The DenseNet structure is used to fuse the optical image and the synthetic aperture radar image to obtain a fused enhanced image. The feature extractor is used to extract the shallow features and deep features of the optical image and the synthetic aperture radar image. The shallow features and the deep features are used for information measurement between the fused enhanced image and the optical image and the synthetic aperture radar image.
Citation Information
Cited By
Multi-source sensor oil tank data monitoring and abnormity early warning method and system
CN120820202A