A method for enhancing the detection of thermal infrared security targets in oilfields under low-light dusty conditions

By constructing a dust enhancement network containing training samples of both clear and degraded images and multi-path convolutional branches, and performing end-to-end joint training with a target detection network, the stability and accuracy issues of target detection in low-light dusty environments in oil fields were resolved, achieving synergistic optimization of image enhancement and detection.

CN122493393APending Publication Date: 2026-07-31DAQING ANRUIDA TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DAQING ANRUIDA TECH DEV CO LTD
Filing Date
2026-05-11
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In the low-light dust environment of oilfields, existing technologies struggle to simultaneously improve image signal-to-noise ratio, recover target edges, and preserve small target features. Furthermore, the models lack generalization ability under different well sites and dust conditions, resulting in insufficient stability and accuracy in thermal infrared security target detection.

Method used

Training samples containing clear thermal infrared images and low-light dust degradation images were constructed. A dust removal enhancement network with multiple convolutional branches and feature fusion layers with different receptive fields was used. The network was jointly trained end-to-end with the target detection network. A joint loss function was constructed by constructing reconstruction loss and detection loss to optimize the network parameters.

Benefits of technology

It improves the stability and accuracy of target detection in low-light dusty environments, enhances target contrast and edge details in images, adapts to the feature extraction requirements of target detection networks, and improves the model's adaptability in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493393A_ABST
    Figure CN122493393A_ABST
Patent Text Reader

Abstract

This invention discloses a method for enhancing and detecting thermal infrared security targets in oilfields under low-light dusty conditions. Belonging to the field of oilfield security monitoring technology, this method addresses the problems of feature degradation in thermal infrared images under low-light dusty conditions and insufficient target detection stability caused by the separation of enhancement and detection training. It constructs training samples containing clear thermal infrared images and low-light dusty degraded images, sets up multiple dust-removing enhancement networks with different receptive fields, and performs end-to-end joint training between the dust-removing enhancement network and the target detection network. This allows detection loss to participate in the parameter update of the dust-removing enhancement network, achieving improved stability of thermal infrared security target detection in oilfields. This method is applicable to thermal infrared security monitoring in oilfield well sites, gathering and transportation stations, and field production areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of oilfield security monitoring technology, specifically involving thermal infrared image enhancement, deep learning target detection, and end-to-end joint training technology for enhanced detection in low-light dust environments. Background Technology

[0002] Oilfield well sites, gathering and transportation stations, and other production areas are typically scattered, placing high demands on the continuity, real-time performance, and stability of monitoring systems for nighttime inspections and perimeter security. Thermal infrared imaging, which does not rely on visible light illumination, can acquire the thermal radiation characteristics of targets such as personnel, vehicles, and equipment at night and under low-light conditions, thus gradually becoming an important technical means in oilfield security monitoring. With the development of deep learning target detection technology, thermal infrared image enhancement, infrared target detection, and image preprocessing methods for complex environments have been introduced into security monitoring scenarios to improve target recognition capabilities in harsh environments.

[0003] In existing technologies, oilfield infrared security solutions typically focus on the deployment of infrared imaging equipment, panoramic scanning, moving target detection, and coordinated monitoring and command. Low-light target detection solutions often combine low-light image enhancement networks with target detection networks to improve detection performance in low-light conditions. Infrared image enhancement solutions frequently employ histogram equalization, filtering, detail enhancement, noise reduction, or image layering to improve image quality. Complex weather target detection solutions typically enhance images of fog, dust storms, or low visibility conditions before inputting the enhanced images into the target detection network for detection. These solutions are effective for general infrared monitoring, low-light detection, and image enhancement tasks.

[0004] However, the degradation of thermal infrared images in low-light dusty environments in oilfields is complex. Low light levels reduce the image signal-to-noise ratio and cause insufficient target contrast, while dust can lead to scattering blurring, loss of edge details, and weakening of small target features. Existing single enhancement or general denoising methods struggle to simultaneously suppress dust noise, restore target edges, and preserve small target details. While some solutions incorporate target detection networks, the enhancement and detection modules are often trained separately. The enhancement results primarily serve visual perception and fail to use detection loss to inversely constrain the enhancement process. This results in enhanced images that may not meet the feature extraction requirements of the detection network, and may even lead to missed or false detections due to artifacts or over-smoothing.

[0005] Furthermore, the acquisition of thermal infrared security data in oilfields is affected by on-site safety regulations, equipment type, environmental conditions, and dust levels, resulting in a limited number of real samples and insufficient scene coverage. Existing solutions lack unified processing of synthetic simulation data and real equipment data, and the generalization ability of models is prone to fluctuations under different well sites and dust conditions. Therefore, it is necessary to propose a thermal infrared security target enhancement detection method for low-illuminance dusty environments in oilfields. This method involves constructing training samples containing clear thermal infrared images and low-illuminance dust-degraded images, combining multiple dust-removing enhancement networks with different receptive fields, and performing end-to-end joint training of the dust-removing enhancement network and the target detection network. This allows detection loss to participate in the parameter update of the enhancement network, thereby improving the stability and accuracy of thermal infrared security target detection in harsh environments. Summary of the Invention

[0006] To address the issues of feature degradation in thermal infrared images under low-light dusty environments and insufficient target detection stability caused by separation training between enhancement and detection in existing technologies, this invention proposes the following solutions: A method for enhancing the detection of thermal infrared security targets in oilfields under low-illumination dusty conditions, the method comprising: S1. Acquire thermal infrared data of the oilfield thermal infrared security scene, and construct training samples containing clear thermal infrared images and low-light dust degradation images based on the thermal infrared data. S2. Construct a dust removal enhancement network, which includes multiple convolutional branches with different receptive fields, a feature fusion layer, and a reconstruction output layer. The network extracts multi-scale features from the low-light dust degradation image through the multiple convolutional branches with different receptive fields. S3. The multi-scale features output by the convolution branches with different receptive fields are fused and an enhanced thermal infrared image is output through the reconstruction output layer. S4. Concatenate the dust removal enhancement network and the target detection network to construct an end-to-end joint training framework. Input the training samples into the end-to-end joint training framework. Obtain an enhanced thermal infrared image based on the low-light dust degradation image. Determine the reconstruction loss based on the enhanced thermal infrared image and the clear thermal infrared image. Determine the detection loss based on the detection prediction results of the target detection network. S5. Construct a joint loss function based on the reconstruction loss and the detection loss, and perform joint optimization on the dust removal enhancement network and the target detection network through backpropagation, so that the detection loss participates in the parameter update of the dust removal enhancement network, and obtain the trained dust removal enhancement network weights and detection network weights. S6. Load the trained dust removal enhancement network weights and detection network weights, perform enhanced detection end-to-end inference on the thermal infrared image or video frame to be detected, and output the target detection result.

[0007] Furthermore, the thermal infrared data in S1 includes synthetic simulation data and real device data. The synthetic simulation data includes at least one of OpenCV physical simulation synthetic data and noise-injected mathematical simulation synthetic data. The real device data includes at least one of Anruida thermal infrared camera real data and FLIR public thermal infrared dataset.

[0008] Further, S1, the construction of training samples containing clear thermal infrared images and low-light dust degradation images based on the thermal infrared data, includes: when the thermal infrared data is synthetic simulation data, performing low-light brightness attenuation on the clear thermal infrared image, and then superimposing Gaussian dust noise and Gaussian blur to generate the low-light dust degradation image; or, simulating low-light attenuation through multiplicative noise and simulating the particle scattering effect of dust through Poisson noise to generate the low-light dust degradation image; when the thermal infrared data is real device data, loading the thermal infrared image corresponding to the real device data, performing grayscale conversion and size normalization preprocessing on the thermal infrared image, and generating the corresponding low-light dust degradation image through a degradation algorithm.

[0009] Furthermore, the multiple convolutional branches with different receptive fields mentioned in S2 include 3×3 convolutional branches, 5×5 convolutional branches, and 7×7 convolutional branches. Each convolutional branch includes two convolutional layers with the same kernel size and a ReLU activation function, and padding is used to ensure that the output size is consistent with the input size.

[0010] Furthermore, S3, which involves fusing the multi-scale features output from the convolutional branches of the multiple receptive fields and outputting an enhanced thermal infrared image through the reconstructed output layer, includes: stitching the multi-scale features output from the convolutional branches of the multiple receptive fields along the channel dimension and performing feature fusion through the feature fusion layer; sequentially passing the fused multi-scale features through a 3×3 convolutional layer with 96 input channels and 64 output channels, and another 3×3 convolutional layer with 64 input channels and 1 output channel, and outputting an enhanced thermal infrared image with values ​​normalized to [0,1] through the Sigmoid function.

[0011] Further, the target detection network in S4 is a YOLO detection network; the cascading of the dust removal enhancement network and the target detection network in S4 includes: performing channel copying on the single-channel enhanced thermal infrared image output by the dust removal enhancement network to obtain a 3-channel input tensor, and inputting the 3-channel input tensor into the YOLO detection network.

[0012] Furthermore, the reconstruction loss in S4 is the MSE reconstruction loss, which is determined based on the pixel difference between the enhanced thermal infrared image and the clear thermal infrared image. The detection loss is the target detection loss output by the target detection network. The joint loss function is composed of the MSE reconstruction loss and the target detection loss.

[0013] Furthermore, the joint optimization of the dust removal enhancement network and the target detection network by backpropagation as described in S5 includes: using the AdamW optimizer to include the trainable parameters of the dust removal enhancement network and the target detection network in the optimization range, and sequentially completing enhancement forward inference, detection forward inference, joint loss calculation, backpropagation and parameter update in each training batch.

[0014] Further, S6 describes performing enhanced detection end-to-end inference on the thermal infrared image or video frame to be detected, including: performing grayscale conversion, size normalization to 256×256, floating-point tensor conversion, and [0,1] interval normalization on the input thermal infrared image or video frame; adding batch dimension and channel dimension before inputting it into the trained dust removal enhancement network to obtain the enhanced image tensor; converting the enhanced image tensor into a single-channel image; then converting it into a 3-channel BGR format before inputting it into the trained target detection network; and outputting the target bounding box coordinates, category ID, confidence score, and visualized detection results.

[0015] Based on the same inventive concept, the present invention also proposes a computer storage medium storing a computer program, which executes the above-mentioned method for enhancing the detection of oilfield thermal infrared security targets in low-light dusty environments when running on a processor.

[0016] Compared with the prior art, the present invention has the following beneficial effects: By constructing training samples containing clear thermal infrared images and low-light dust degradation images based on thermal infrared security scenarios in oil fields, the training data can cover the characteristics of low-light brightness decay and dust degradation, thereby improving the model's adaptability to low-light dust environments in oil fields.

[0017] By setting up a dust removal enhancement network that includes multiple convolutional branches with different receptive fields, feature fusion layers, and reconstruction output layers, features of different scales in low-light dust degradation images can be extracted and fused, thereby improving the problems of insufficient target contrast and weakened edge details in thermal infrared images.

[0018] By cascading the dust removal enhancement network with the object detection network and constructing a joint loss function based on the reconstruction loss and detection loss, the detection loss is made involved in the parameter update of the dust removal enhancement network, thereby reducing the feature requirement mismatch problem caused by separate training of enhancement and detection.

[0019] Compared with existing technologies, this invention improves the stability of target detection results in low-light dusty environments by incorporating detection loss into the parameter update of the dust removal enhancement network, making the enhanced thermal infrared image more suitable for the feature extraction requirements of the target detection network.

[0020] By loading the trained dust removal enhancement network weights and detection network weights, end-to-end inference for enhanced detection is performed on the thermal infrared image or video frame to be detected, so that image enhancement and target detection are completed continuously in the same inference process, thereby outputting the target bounding box coordinates, category ID, confidence score and visual detection results.

[0021] This invention has the ability to synergistically optimize thermal infrared image enhancement and target detection in low-light dusty environments, which can improve the detection stability of targets such as personnel, vehicles and equipment in oilfield security scenarios. It is applicable to thermal infrared security monitoring in oilfield well sites, gathering and transportation stations and field production areas. Attached Figure Description

[0022] Figure 1 This is a flowchart of the oilfield thermal infrared security target enhancement detection method under low-illuminance sand and dust environment described in the implementation method; Figure 2 This is the overall system architecture diagram described in the implementation method; Figure 3 This is a diagram of the multi-scale convolutional fusion dust removal enhancement network structure described in the implementation method; Figure 4 This is a schematic diagram of the end-to-end joint training framework described in the implementation method. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Implementation Method 1 like Figure 1 As shown, a method for enhancing the detection of thermal infrared security targets in oilfields under low-light dusty conditions is described, the method comprising: S1. Acquire thermal infrared data of the oilfield thermal infrared security scene, and construct training samples containing clear thermal infrared images and low-light dust degradation images based on the thermal infrared data. S2. Construct a dust removal enhancement network, which includes multiple convolutional branches with different receptive fields, a feature fusion layer, and a reconstruction output layer. The network extracts multi-scale features from the low-light dust degradation image through the multiple convolutional branches with different receptive fields. S3. The multi-scale features output by the convolution branches with different receptive fields are fused and an enhanced thermal infrared image is output through the reconstruction output layer. S4. Concatenate the dust removal enhancement network and the target detection network to construct an end-to-end joint training framework. Input the training samples into the end-to-end joint training framework. Obtain an enhanced thermal infrared image based on the low-light dust degradation image. Determine the reconstruction loss based on the enhanced thermal infrared image and the clear thermal infrared image. Determine the detection loss based on the detection prediction results of the target detection network. S5. Construct a joint loss function based on the reconstruction loss and the detection loss, and perform joint optimization on the dust removal enhancement network and the target detection network through backpropagation, so that the detection loss participates in the parameter update of the dust removal enhancement network, and obtain the trained dust removal enhancement network weights and detection network weights. S6. Load the trained dust removal enhancement network weights and detection network weights, perform enhanced detection end-to-end inference on the thermal infrared image or video frame to be detected, and output the target detection result.

[0025] By constructing a complete processing chain that includes training samples, a dust removal enhancement network, a target detection network, a joint loss function, and an end-to-end inference process, the enhancement processing of low-light dust-degraded images can be continuously linked with the target detection process, thereby improving the stability of thermal infrared security target detection results.

[0026] Furthermore, the thermal infrared data in S1 includes synthetic simulation data and real device data. The synthetic simulation data includes at least one of OpenCV physical simulation synthetic data and noise-injected mathematical simulation synthetic data. The real device data includes at least one of Anruida thermal infrared camera real data and FLIR public thermal infrared dataset.

[0027] By using synthetic simulation data and real equipment data as the source of thermal infrared data, it is possible to take into account both the construction of low-illuminance dust degradation features and the introduction of real thermal infrared image features, thereby improving the coverage of training samples for oilfield thermal infrared security scenarios.

[0028] Further, S1, the construction of training samples containing clear thermal infrared images and low-light dust degradation images based on the thermal infrared data, includes: when the thermal infrared data is synthetic simulation data, performing low-light brightness attenuation on the clear thermal infrared image, and then superimposing Gaussian dust noise and Gaussian blur to generate the low-light dust degradation image; or, simulating low-light attenuation through multiplicative noise and simulating the particle scattering effect of dust through Poisson noise to generate the low-light dust degradation image; when the thermal infrared data is real device data, loading the thermal infrared image corresponding to the real device data, performing grayscale conversion and size normalization preprocessing on the thermal infrared image, and generating the corresponding low-light dust degradation image through a degradation algorithm.

[0029] Preferably, the low-light dust degradation image corresponds to the clear thermal infrared image, so that the reconstruction loss can be determined based on paired images.

[0030] By processing clear thermal infrared images with low-light brightness attenuation, Gaussian dust noise, Gaussian blur, multiplicative noise, or Poisson noise, and by preprocessing and degrading real device data, degraded training samples that match low-light dust environments can be generated, thus providing a data foundation for the dust removal enhancement network to learn the mapping relationship from degraded images to enhanced images.

[0031] Furthermore, the multiple convolutional branches with different receptive fields mentioned in S2 include 3×3 convolutional branches, 5×5 convolutional branches, and 7×7 convolutional branches. Each convolutional branch includes two convolutional layers with the same kernel size and a ReLU activation function, and padding is used to ensure that the output size is consistent with the input size.

[0032] Preferably, the 3×3 convolutional branch, 5×5 convolutional branch, and 7×7 convolutional branch are set in parallel to extract detail features, structural features, and background correlation features from low-light dust degradation images, respectively.

[0033] By setting up 3×3, 5×5, and 7×7 convolutional branches, and ensuring that each convolutional branch includes a convolutional layer and a ReLU activation function, multi-scale features of low-light dust degradation images can be extracted from different receptive fields, thereby enhancing the dust removal enhancement network's ability to express target details and image structure.

[0034] Furthermore, S3, which involves fusing the multi-scale features output from the convolutional branches of the multiple receptive fields and outputting an enhanced thermal infrared image through the reconstructed output layer, includes: stitching the multi-scale features output from the convolutional branches of the multiple receptive fields along the channel dimension and performing feature fusion through the feature fusion layer; sequentially passing the fused multi-scale features through a 3×3 convolutional layer with 96 input channels and 64 output channels, and another 3×3 convolutional layer with 64 input channels and 1 output channel, and outputting an enhanced thermal infrared image with values ​​normalized to [0,1] through the Sigmoid function.

[0035] Preferably, the feature fusion layer is used to integrate multi-scale features after channel dimension stitching, and the reconstruction output layer is used to reconstruct the integrated features into a single-channel enhanced thermal infrared image.

[0036] By stitching multi-scale features together along the channel dimension and normalizing the numerical range of the output to [0,1] through a feature fusion layer and a reconstruction output layer, the enhanced thermal infrared image can be transformed into a unified enhancement result, thereby improving the target identifiability in low-light dust degradation images.

[0037] Further, the target detection network in S4 is a YOLO detection network; the cascading of the dust removal enhancement network and the target detection network in S4 includes: performing channel copying on the single-channel enhanced thermal infrared image output by the dust removal enhancement network to obtain a 3-channel input tensor, and inputting the 3-channel input tensor into the YOLO detection network.

[0038] Preferably, the channel replication is used to adapt the single-channel enhanced thermal infrared image to the input channel requirements of the YOLO detection network.

[0039] By copying the single-channel enhanced thermal infrared image output by the dust removal enhancement network and inputting the resulting 3-channel input tensor into the YOLO detection network, the enhanced thermal infrared image can be matched with the input format of the target detection network, thereby realizing the continuous transfer of image enhancement results to the target detection process.

[0040] Furthermore, the reconstruction loss in S4 is the MSE reconstruction loss, which is determined based on the pixel difference between the enhanced thermal infrared image and the clear thermal infrared image. The detection loss is the target detection loss output by the target detection network. The joint loss function is composed of the MSE reconstruction loss and the target detection loss.

[0041] Preferably, the MSE reconstruction loss is used to constrain the pixel consistency between the enhanced thermal infrared image and the clear thermal infrared image, and the target detection loss is used to constrain the detection prediction results of the target detection network.

[0042] By using the MSE reconstruction loss and the target detection loss to form a joint loss function, we can simultaneously constrain the reconstruction quality of the enhanced thermal infrared image and the detection results of the target detection network, thereby keeping the training process of the dust removal enhancement network related to the target detection requirements.

[0043] Furthermore, the joint optimization of the dust removal enhancement network and the target detection network by backpropagation as described in S5 includes: using the AdamW optimizer to include the trainable parameters of the dust removal enhancement network and the target detection network in the optimization range, and sequentially completing enhancement forward inference, detection forward inference, joint loss calculation, backpropagation and parameter update in each training batch.

[0044] Preferably, the detection loss is transmitted to the dust removal enhancement network during backpropagation, so that the parameter update of the dust removal enhancement network is constrained by the detection result of the target detection network.

[0045] By employing the AdamW optimizer to incorporate the trainable parameters of the dust removal enhancement network and the target detection network into the optimization scope, and completing forward inference, joint loss calculation, backpropagation, and parameter updates in each training batch, synchronous optimization of the dust removal enhancement network and the target detection network can be achieved, thereby making the enhanced thermal infrared image more suitable for the target detection process.

[0046] Further, S6 describes performing enhanced detection end-to-end inference on the thermal infrared image or video frame to be detected, including: performing grayscale conversion, size normalization to 256×256, floating-point tensor conversion, and [0,1] interval normalization on the input thermal infrared image or video frame; adding batch dimension and channel dimension before inputting it into the trained dust removal enhancement network to obtain the enhanced image tensor; converting the enhanced image tensor into a single-channel image; then converting it into a 3-channel BGR format before inputting it into the trained target detection network; and outputting the target bounding box coordinates, category ID, confidence score, and visualized detection results.

[0047] Preferably, the target detection result includes target bounding box coordinates, category ID, confidence score, and visual detection result.

[0048] By preprocessing, dust removal, enhancement, and target detection of the thermal infrared image or video frame to be detected, the end-to-end processing of enhanced detection can be completed continuously in the same inference flow, thereby outputting the target bounding box coordinates, category ID, confidence score, and visualized detection results.

[0049] The method described in this embodiment can be executed by a processor calling a computer program, which can be stored in a computer storage medium. When the computer program is executed by the processor, the above-mentioned method for enhancing the detection of thermal infrared security targets in oilfields under low-light dust environments can be implemented.

[0050] Implementation Method 2 In this embodiment, the oilfield thermal infrared security target enhancement detection method under low-light dusty environments forms a complete processing flow, encompassing multi-source data processing, dust removal enhancement network construction, end-to-end joint training, end-to-end inference for enhanced detection, and display and storage of detection results. This method is designed for thermal infrared security scenarios such as remote well sites and gathering stations in oilfields, and is used to enhance the detection of personnel, vehicles, and equipment in low-light, dusty environments at night. The method uses training thermal infrared data and thermal infrared images or video frames to be detected as input. By constructing training samples containing clear thermal infrared images and low-light dusty degraded images, the dust removal enhancement network and the target detection network form a continuous data transfer relationship during joint training, and the target detection results are output during the inference phase.

[0051] In this embodiment, when executing the method, a processing architecture consisting of a data processing layer, an image enhancement layer, a joint training layer, an inference detection layer, and a visualization interaction layer can be adopted. For details of the architecture, please refer to [link to specific architecture]. Figure 2 The data processing layer generates, loads, and standardizes preprocessing multi-source data for oilfield thermal infrared security scenarios, providing multi-scenario and multi-type data source support for model training. The image enhancement layer uses a multi-scale convolutional fusion-based dust removal enhancement network to perform end-to-end enhancement of low-light dust-degraded thermal infrared images, restoring target detail features and outputting enhanced thermal infrared images. The joint training layer enables end-to-end joint training of the dust removal enhancement network and the target detection network, achieving collaborative optimization of the two networks through a joint loss function, and outputting dust removal enhancement network weights and detection network weights adapted to oilfield security scenarios. The inference detection layer performs end-to-end inference for enhanced detection on the input thermal infrared static images or dynamic video streams, outputting target detection results and visualized detection results. The visualization interaction layer can build an integrated GUI interface based on a multi-tab, multi-threaded architecture, realizing visualized operations and result display for data preparation, model training, and inference demonstration.

[0052] In the data processing, thermal infrared data includes synthetic simulation data and real device data. Synthetic simulation data includes at least one of OpenCV physical simulation data and noise-injected mathematical simulation data. Real device data includes at least one of Anruida thermal infrared camera real data and FLIR public thermal infrared datasets. By using these data sources, both synthetic simulation data and real device data can be covered, and thermal infrared images from different sources can be uniformly generated, loaded, and standardized preprocessed to construct training samples containing clear thermal infrared images and low-light dust degradation images.

[0053] In the process of synthesizing simulation data, for low-illuminance dusty environments in oilfields, two methods can be used: OpenCV physical simulation for data generation and noise-injected mathematical simulation for data generation. When using OpenCV physical simulation, the clear thermal infrared image is first subjected to low-illuminance brightness attenuation, and then Gaussian dust noise and Gaussian blur are superimposed to generate a low-illuminance dusty degradation image. The low-illuminance brightness attenuation involves the brightness attenuation coefficient, brightness offset, the clear thermal infrared image, and the low-illuminance processed image; the dust degradation involves Gaussian noise, Gaussian kernel size, blur standard deviation, and the final output low-illuminance dusty degradation image. When using noise-injected mathematical simulation for data generation, multiplicative noise is used to simulate low-illuminance attenuation, and Poisson noise is used to simulate the particle scattering effect of dust, adapting to different degrees of harshness in the environment. Among them, the low-light multiplicative attenuation involves a random attenuation coefficient, the value of which ranges from [0.2, 0.5]; the dust Poisson noise degradation involves a Poisson distribution parameter and a pixel value truncation function, the pixel value truncation function being used to limit the output to an effective pixel range of 0-255.

[0054] During the processing of real equipment data, real thermal infrared images from Anruida thermal infrared cameras or the corresponding FLIR public thermal infrared dataset are loaded. These thermal infrared images undergo grayscale conversion and size normalization preprocessing, and corresponding low-light dust degradation images are generated using a degradation algorithm. For real equipment data, standardized samples with the same format as the synthetic simulation data are output after processing, allowing for direct integration into subsequent model training processes. Through this method, clear thermal infrared images and low-light dust degradation images maintain a correspondence, forming paired training samples.

[0055] In the construction of the dust removal enhancement network, the network includes multiple convolutional branches with different receptive fields, a feature fusion layer, and a reconstruction output layer. Its network structure is described in [reference needed]. Figure 3The multiple convolutional branches with different receptive fields include 3×3, 5×5, and 7×7 convolutional branches. Each convolutional branch includes two convolutional layers with the same kernel size and a ReLU activation function, and padding is used to ensure that the output size is consistent with the input size. The 3×3, 5×5, and 7×7 convolutional branches are set in parallel and used to extract multi-scale features from low-light dust degradation images under different receptive fields.

[0056] In the multi-scale feature fusion and reconstruction process, the multi-scale features output from the 3×3, 5×5, and 7×7 convolutional branches are concatenated along the channel dimension and then fused through a feature fusion layer. The fused multi-scale features are then sequentially passed through a 3×3 convolutional layer with 96 input channels and 64 output channels, and another 3×3 convolutional layer with 64 input channels and 1 output channel. Finally, an enhanced thermal infrared image with values ​​normalized to [0,1] is output using the Sigmoid function. The size of the enhanced thermal infrared image is the same as the input image size and serves as the input basis for the subsequent object detection network.

[0057] In the end-to-end joint training process, the dust removal enhancement network and the object detection network are cascaded to construct an end-to-end joint training framework. For details of the training framework, please refer to [link to training framework]. Figure 4 The target detection network can employ the YOLO detection network. After training the end-to-end joint training framework with input samples, the low-light dust degradation image is first input into the dust removal and enhancement network to obtain an enhanced thermal infrared image. Since the YOLO detection network requires 3-channel input, the single-channel enhanced thermal infrared image output by the dust removal and enhancement network is channel-copied to obtain a 3-channel input tensor. This 3-channel input tensor is then input into the YOLO detection network to obtain the detection prediction result, and the detection loss is determined based on the prediction result.

[0058] In the construction of the joint loss function, the reconstruction loss is the MSE reconstruction loss, which is determined based on the pixel difference between the enhanced thermal infrared image and the clear thermal infrared image. The detection loss is the target detection loss output by the target detection network. The joint loss function is composed of both the MSE reconstruction loss and the target detection loss. During backpropagation through the joint loss function, the gradient of the joint loss function is simultaneously transmitted to both the dust removal enhancement network and the target detection network, allowing the detection loss to participate in the parameter update of the dust removal enhancement network and simultaneously optimizing both networks.

[0059] During parameter optimization, the AdamW optimizer is used to incorporate the trainable parameters of the dust removal and enhancement network and the object detection network into the optimization scope, and corresponding learning rates are set. Each training batch sequentially completes enhancement forward inference, detection forward inference, joint loss calculation, backpropagation, and parameter update; traversing all training samples constitutes one training epoch; after completing the preset training epochs, the trained weights of the dust removal and enhancement network and the detection network are saved. During training, the loss value and training progress of each batch can be output in real time, and the training log on the visualization interface is updated synchronously.

[0060] During system initialization and parameter configuration, the hardware environment and deep learning framework are configured, and the computing devices for model training and inference are determined, prioritizing GPUs and using CPUs when GPUs are unavailable. Core operating parameters are set, including input image size, training batch size, number of training epochs, and learning rate, where the input image size can be set to 256×256. The core structure and weight parameters of the dust removal and object detection networks are initialized, and YOLO pre-trained weights are loaded.

[0061] During the generation, loading, and preprocessing of multi-source data, the data source type can be selected through a visual interface. Supported data source types include OpenCV physical simulation synthetic data, noise-injected mathematical simulation synthetic data, real data from Anruida thermal infrared cameras, and FLIR public thermal infrared datasets. If the synthetic simulation data type is selected, a corresponding number of clear thermal infrared simulation images are automatically generated, and low-light dust degradation images are generated using the corresponding degradation algorithm to construct paired training samples. If the real device data type is selected, thermal infrared images from the specified directory are automatically loaded, grayscale conversion and size normalization preprocessing are performed, and corresponding low-light dust degradation images are generated using the degradation algorithm to construct the training dataset. The preprocessed training dataset is then used as a data loader to enable batch data reading during the training process.

[0062] In the process of building the joint training framework, the dust removal enhancement network and the object detection network are cascaded to form an end-to-end joint training network architecture; a joint loss function is defined, which includes the MSE reconstruction loss of the dust removal enhancement network and the detection loss of the object detection network; the AdamW optimizer is initialized, and the trainable parameters of the dust removal enhancement network and the object detection network are included in the optimization range; a multi-threaded training task is constructed so that the training process can run in the background and output logs.

[0063] During end-to-end joint training and model optimization, a joint training task is initiated, and training data is input into the joint training framework in batches. The framework sequentially completes augmentation forward inference, detection forward inference, joint loss calculation, backpropagation, and parameter updates. During training, the loss value and training progress for each batch are output in real time, and the training log on the visualization interface is updated synchronously. After completing a preset number of training rounds, the trained weight files for the dust removal and augmentation networks and the detection network are automatically saved, and a training completion notification is displayed on the interface.

[0064] During the inference model loading and initialization process, the trained weights of the dust removal and enhancement network and the object detection network are loaded, and the dust removal and enhancement network and the object detection network are initialized respectively, then switched to the inference evaluation mode. An end-to-end inference engine is constructed, and the inference process is encapsulated to adapt to static images and dynamic video stream inputs. A multi-threaded inference task is constructed to enable the video inference process to run in the background and render results in real time.

[0065] During the end-to-end inference process for enhanced detection of thermal infrared images or video frames, the static image or video file to be detected is selected through a visual interface. If the input is a static image, the image is read and input into the end-to-end inference engine to complete the end-to-end inference of enhancement and detection, and the enhanced detection result image is displayed on the interface. If the input is a video file, a background inference thread is started, the video is read frame by frame and the inference processing is completed, the detection result image is rendered on the interface in real time, and the inference frame rate (FPS) is updated synchronously. After inference is completed, the interface operation buttons are automatically unlocked and an inference completion prompt is output.

[0066] During inference preprocessing, the input thermal infrared image or video frame undergoes grayscale conversion, size normalization to 256×256, floating-point tensor conversion, and [0,1] interval normalization. After adding batch and channel dimensions, it is input into the trained dust removal enhancement network to obtain the enhanced image tensor. The enhanced image tensor is converted into a single-channel image, then into a 3-channel BGR format, and input into the trained target detection network. The target detection network outputs target detection results, including target bounding box coordinates, category ID, confidence score, and visual detection results. In the visual detection results, target detection boxes can be drawn on the enhanced thermal infrared image, and the target category and confidence score can be labeled, while the inference frame rate (FPS) is calculated.

[0067] During the display and storage of detection results, a visual interface displays the results in real time. These results include enhanced images, target detection boxes, category labels, confidence scores, and inference frame rates (FPS). Detection result images can be stored locally to create a traceable record of oilfield security detection data. Logs for the training and inference processes are automatically recorded, facilitating subsequent troubleshooting and model optimization.

[0068] During the visualization interaction process, a multi-tab, multi-threaded architecture can be used to build the visualization GUI system, forming three functional areas: data preparation, model training, and inference demonstration. Multiple tabs are used to differentiate different processing flows, while multi-threading allows model training and video inference tasks to be executed in the background to maintain the responsiveness of the main interface. The visualization interface can output training logs, inference frame rates, and detection results in real time, and supports the selection and loading of local video or image files to adapt to the offline monitoring data processing needs of oilfield sites.

[0069] During the verification process, a specialized test dataset can be constructed for low-light dusty environments in oilfields. This dataset contains 1000 real-world thermal infrared images of oilfield scenes, covering three typical harsh oilfield scenarios: low-light nighttime conditions, light dust, and heavy dust. It labels four core security targets: personnel, vehicles, pumping units, and pipelines, with small targets (≤32×32 pixels) accounting for ≥60%. A traditional approach combining histogram equalization enhancement with YOLOv8n detection, and a separate training approach combining individual enhancement network training with YOLOv8n detection, can be selected as control groups. Comparative experiments can be conducted under the same hardware environment and dataset.

[0070] In target detection performance verification, the average detection accuracy (mAP@0.5) of this implementation scheme was 95.8%, compared to 82.3% for the traditional scheme and 88.7% for the separate training scheme, representing an improvement of 13.5 percentage points over the traditional scheme. In low-light scenes, the mAP@0.5 was 96.2%, compared to 84.1% for the traditional scheme and 90.2% for the separate training scheme, representing an improvement of 12.1 percentage points over the traditional scheme. In heavy dust scenes, the mAP@0.5 was 93.5%, compared to 76.8% for the traditional scheme and 83.4% for the separate training scheme, representing an improvement of 16.7 percentage points over the traditional scheme. The recall rate for small targets was 94.2%, compared to 78.5% for the traditional scheme and 85.1% for the separate training scheme, representing an improvement of 15.7 percentage points over the traditional scheme. The false positive rate was 2.8%, compared to 11.6% for the traditional scheme and 7.3% for the separate training scheme, representing a reduction of 8.8 percentage points over the traditional scheme.

[0071] In the real-time inference verification, the experimental hardware environments were a PC with an NVIDIA RTX 3060 GPU and a Jetson Xavier NX edge processor, with an input image resolution of 256×256. Under the RTX 3060 GPU environment, the end-to-end inference frame rate of this implementation scheme was 58.2 FPS, while the traditional scheme achieved 62.5 FPS. Under the Jetson Xavier NX edge processor environment, the end-to-end inference frame rate of this implementation scheme was 22.6 FPS, while the traditional scheme achieved 24.1 FPS.

[0072] In the verification of model generalization ability, for thermal infrared acquisition data of new oilfield well sites that were not involved in training, the mAP@0.5 of this implementation scheme was 92.7%, while the mAP@0.5 of the traditional scheme was 79.3%.

[0073] The above detailed description of the technical solution provided by the present invention is intended to highlight the advantages and benefits of the technical solution provided by the present invention. However, the above detailed embodiments are not intended to limit the scope of protection of the present invention. Any reasonable modifications and improvements to the present invention, recombination of embodiments, and equivalent substitutions based on the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0074] Those skilled in the art will understand that the above description is merely a preferred embodiment of the present invention, and the features described in the various embodiments and / or claims disclosed in the present invention can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in the disclosure of the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle scope of the present invention should be considered to fall within the protection scope of the present invention.

[0075] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.

Claims

1. A method for enhancing the detection of thermal infrared security targets in oilfields under low-illumination dusty environments, characterized in that, The method includes: S1. Acquire thermal infrared data of the oilfield thermal infrared security scene, and construct training samples containing clear thermal infrared images and low-light dust degradation images based on the thermal infrared data. S2. Construct a dust removal enhancement network, which includes multiple convolutional branches with different receptive fields, a feature fusion layer, and a reconstruction output layer. The network extracts multi-scale features from the low-light dust degradation image through the multiple convolutional branches with different receptive fields. S3. The multi-scale features output by the convolution branches with different receptive fields are fused and an enhanced thermal infrared image is output through the reconstruction output layer. S4. Concatenate the dust removal enhancement network and the target detection network to construct an end-to-end joint training framework. Input the training samples into the end-to-end joint training framework. Obtain an enhanced thermal infrared image based on the low-light dust degradation image. Determine the reconstruction loss based on the enhanced thermal infrared image and the clear thermal infrared image. Determine the detection loss based on the detection prediction results of the target detection network. S5. Construct a joint loss function based on the reconstruction loss and the detection loss, and perform joint optimization on the dust removal enhancement network and the target detection network through backpropagation, so that the detection loss participates in the parameter update of the dust removal enhancement network, and obtain the trained dust removal enhancement network weights and detection network weights. S6. Load the trained dust removal enhancement network weights and detection network weights, perform enhanced detection end-to-end inference on the thermal infrared image or video frame to be detected, and output the target detection result.

2. The method according to claim 1, characterized in that, The thermal infrared data mentioned in S1 includes synthetic simulation data and real device data. The synthetic simulation data includes at least one of OpenCV physical simulation synthetic data and noise-injected mathematical simulation synthetic data. The real device data includes at least one of Anruida thermal infrared camera real data and FLIR public thermal infrared dataset.

3. The method according to claim 1, characterized in that, S1 describes constructing training samples based on the thermal infrared data, including clear thermal infrared images and low-light dust degradation images, including: When the thermal infrared data is synthetic simulation data, the clear thermal infrared image is subjected to low-light brightness attenuation, and then Gaussian dust noise and Gaussian blur are superimposed to generate the low-light dust degradation image; or, the low-light attenuation is simulated by multiplicative noise and the particle scattering effect of dust is simulated by Poisson noise to generate the low-light dust degradation image. When the thermal infrared data is real device data, load the thermal infrared image corresponding to the real device data, perform grayscale conversion and size normalization preprocessing on the thermal infrared image, and generate the corresponding low-light dust degradation image through a degradation algorithm.

4. The method according to claim 1, characterized in that, The multiple convolutional branches with different receptive fields described in S2 include 3×3 convolutional branches, 5×5 convolutional branches, and 7×7 convolutional branches. Each convolutional branch includes two convolutional layers with the same kernel size and a ReLU activation function, and padding is used to ensure that the output size is consistent with the input size.

5. The method according to claim 1, characterized in that, S3 describes fusing the multi-scale features output from the convolutional branches with different receptive fields and outputting an enhanced thermal infrared image through the reconstructed output layer, including: The multi-scale features output by the convolutional branches with different receptive fields are concatenated in the channel dimension and then fused through the feature fusion layer. The fused multi-scale features are sequentially passed through a 3×3 convolutional layer with 96 input channels and 64 output channels, and then through another 3×3 convolutional layer with 64 input channels and 1 output channel. Finally, the enhanced thermal infrared image is output with the numerical range normalized to [0,1] by the Sigmoid function.

6. The method according to claim 1, characterized in that, The target detection network in S4 is a YOLO detection network; the cascading of the dust removal enhancement network and the target detection network in S4 includes: performing channel copying on the single-channel enhanced thermal infrared image output by the dust removal enhancement network to obtain a 3-channel input tensor, and inputting the 3-channel input tensor into the YOLO detection network.

7. The method according to claim 1, characterized in that, The reconstruction loss in S4 is the MSE reconstruction loss, which is determined based on the pixel difference between the enhanced thermal infrared image and the clear thermal infrared image. The detection loss is the target detection loss output by the target detection network. The joint loss function is composed of the MSE reconstruction loss and the target detection loss.

8. The method according to claim 1, characterized in that, S5 describes the joint optimization of the dust removal enhancement network and the target detection network through backpropagation, which includes: using the AdamW optimizer to include the trainable parameters of the dust removal enhancement network and the target detection network in the optimization range, and sequentially performing enhancement forward inference, detection forward inference, joint loss calculation, backpropagation and parameter update in each training batch.

9. The method according to claim 1, characterized in that, S6 describes performing end-to-end inference for enhanced detection on the thermal infrared image or video frame to be detected, including: performing grayscale conversion, size normalization to 256×256, floating-point tensor conversion, and [0,1] interval normalization on the input thermal infrared image or video frame; adding batch dimension and channel dimension before inputting it into the trained dust removal enhancement network to obtain the enhanced image tensor; converting the enhanced image tensor into a single-channel image; then converting it into a 3-channel BGR format before inputting it into the trained target detection network; and outputting the target bounding box coordinates, category ID, confidence score, and visualized detection results.

10. A computer storage medium having a computer program stored thereon, characterized in that, The computer program, when running on a processor, performs the method according to any one of claims 1 to 9.