Infrared image adaptive detection method and system facing complex background interference

By constructing the YOLO-CCA network model, combining the lightweight convolution module, attention mechanism and feature fusion structure, the loss function is optimized, and the classification and positioning in infrared image detection of substation equipment under complex background is solved, achieving high-precision and efficient detection effects.

CN120298679AActive Publication Date: 2025-07-11NANCHANG INST OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510787283.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-11
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

In complex background, in infrared image detection of substation devices, classification and positioning prediction are inconsistent, which affects detection accuracy.

Method used

A YOLO-CCA network model is constructed, combining lightweight convolution module, attention mechanism module and multi-level feature fusion structure, and dynamically adjustable exponential mass focus loss function and noise-robust focus loss function are introduced to optimize the loss function to achieve adaptive detection.

Benefits of technology

It improves the accuracy and efficiency of infrared image detection, enhances the robustness to complex backgrounds, reduces the leakage detection rate, and improves the detection accuracy of substation equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298679A_ABST
    Figure CN120298679A_ABST
Patent Text Reader

Abstract

The invention discloses a complex background interference-oriented infrared image adaptive detection method and system, and the method comprises the steps: constructing a YOLO-CCA network model which comprises a backbone network, a neck network and a head network; obtaining at least one historical power transformation equipment infrared image, and labeling the at least one historical power transformation equipment infrared image to obtain labeling information corresponding to the at least one historical power transformation equipment infrared image; inputting the at least one historical substation equipment infrared image and the annotation information into a YOLO-CCA network model for iterative training to obtain an adaptive detection model; and inputting the obtained real-time infrared image of the power transformation equipment into the self-adaptive detection model, and outputting a detection result corresponding to the real-time infrared image of the power transformation equipment. Cooperative optimization of positioning and classification tasks is realized by flexibly setting loss weights, and the problem of inconsistent classification and positioning prediction of power transformation equipment under a complex background is solved on the whole.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of infrared image detection, and particularly relates to an infrared image adaptive detection method and system for complex background interference. Background Art

[0002] Substations play important roles in the power grid, such as voltage level conversion, voltage regulation, current collection and distribution. Their operating conditions are directly related to the stability and reliability of the entire power grid. During long-term operation, substation equipment is vulnerable to multiple factors such as aging, overload, and environmental conditions, gradually resulting in latent faults. These faults are usually not easily detectable in the initial stage, but will gradually intensify over time and exhibit more obvious symptoms in the later stage. Among them, the abnormal increase in equipment temperature is regarded as one of the important warning signals of latent faults. Continuous high temperature will not only cause deterioration of the internal material properties of the equipment, but also weaken the insulation performance, thus seriously threatening the overall safety of the equipment and even possibly triggering serious equipment accidents. Therefore, ensuring the safe and stable operation of substation equipment not only has important technical significance, but also has profound social and economic value.

[0003] In the power system, continuous and accurate monitoring of the thermal state of equipment and timely diagnosis have become key technologies for preventing and identifying latent faults and ensuring the safe and stable operation of power equipment. Due to its non-contact operation, no need to disassemble the equipment, no interruption of operation, and high efficiency and accuracy, infrared imaging technology has been widely used in the field of condition monitoring of power equipment, showing great application potential.

[0004] In the past, the inspection work of substations mainly relied on manual use of handheld infrared imagers. This method is not only time-consuming but also has low detection efficiency. Especially in the environment where equipment is complex and intertwined, when detecting substation equipment, the detection accuracy of this traditional method is low and the missed detection rate is high, seriously affecting the detection quality and bringing potential risks to the operation of the power grid.

[0005] With the rapid development of unmanned aerial vehicle (UAV) technology, it has become increasingly common to use UAVs equipped with infrared imagers to perform daily inspection tasks. This method not only improves the inspection efficiency but also significantly enhances the monitoring quality of substation equipment. However, during the mission execution, UAVs will capture and transmit a large amount of infrared image data for background analysis. Given the large volume of these data and the diverse equipment states, relying on traditional manual methods to analyze these data one by one has become both inefficient and difficult to maintain.

[0006] Therefore, to make more effective use of the convenience brought by the UAV-borne infrared imaging technology, it has become necessary to develop a real-time infrared image recognition algorithm. The object detection algorithm based on deep learning is integrated into the UAV system equipped with an infrared imager, which becomes the key technical path to realize the real-time recognition of thermal defects in substation equipment. This method not only improves the intelligent and automated level of detection, but also provides an efficient, low-cost and flexible-response solution for the condition monitoring of substation equipment, thus further enhancing the safety and reliability of the power system. However, when detecting the infrared images of substation equipment in the face of interference from complex backgrounds, the phenomenon of inconsistent classification and positioning prediction easily occurs, thus affecting the detection accuracy. Summary of the Invention

[0007] The present invention provides an adaptive detection method and system for infrared images facing complex background interference, which is used to solve the technical problem of inconsistent classification and positioning prediction under complex backgrounds.

[0008] In the first aspect, the present invention provides an adaptive detection method for infrared images facing complex background interference, including: Construct a YOLO-CCA network model, where the YOLO-CCA network model includes a backbone network, a neck network and a head network. Among them, the expression of the object loss function in the head network is: , In the formula, is the object loss function, is the fusion weight coefficient of the dynamic adjustment exponential quality focal loss function, is the dynamic adjustment exponential quality focal loss function, is the fusion weight coefficient of the noise-robust focal loss function, is the noise-robust focal loss function; Obtain at least one historical infrared image of substation equipment, and annotate the at least one historical infrared image of substation equipment to obtain annotation information corresponding to the at least one historical infrared image of substation equipment; Input the at least one historical infrared image of substation equipment and the annotation information into the YOLO-CCA network model for iterative training to obtain an adaptive detection model; Input the obtained real-time infrared image of substation equipment into the adaptive detection model, and the adaptive detection model outputs a detection result corresponding to the real-time infrared image of substation equipment.

[0009] In the second aspect, the present invention provides an adaptive detection system for infrared images facing complex background interference, including: A building module configured to build a YOLO-CCA network model, the YOLO-CCA network model including a backbone network, a neck network, and a head network, wherein the expression of the object loss function in the head network is: , In the formula, is the object loss function, is the fusion weight coefficient of the dynamically adjusted exponential quality focal loss function, is the dynamically adjusted exponential quality focal loss function, is the fusion weight coefficient of the noise-robust focal loss function, is the noise-robust focal loss function; A labeling module configured to obtain at least one historical infrared image of a substation equipment and label the at least one historical infrared image of the substation equipment to obtain labeling information corresponding to the at least one historical infrared image of the substation equipment; A training module configured to input the at least one historical infrared image of the substation equipment and the labeling information into the YOLO-CCA network model for iterative training to obtain an adaptive detection model; An output module configured to input the obtained real-time infrared image of the substation equipment into the adaptive detection model, and the adaptive detection model outputs a detection result corresponding to the real-time infrared image of the substation equipment.

[0010] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the infrared image adaptive detection method for complex background interference according to any embodiment of the present invention.

[0011] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program instructions are executed by a processor, the processor is enabled to execute the steps of the infrared image adaptive detection method for complex background interference according to any embodiment of the present invention.

[0012] The adaptive infrared image detection method and system for complex background interference of the present application replace the C2f structure in YOLOv8 by introducing a lightweight convolution module, which significantly reduces the number of parameters and computational complexity while maintaining the feature extraction ability. Secondly, an attention mechanism module is designed to enhance the feature response of key regions, suppress background interference, and improve the thermal feature contrast and detection performance. Then, a multi-level feature fusion structure is added to the neck network to strengthen the context dependence between different scales and improve the recognition accuracy of large and small targets. Finally, a dynamic adjustment exponential quality focal loss function (D-QFL) is introduced to adaptively enhance the regression accuracy of high-quality samples; at the same time, a noise-robust focal loss function (NR-Focal) is proposed to handle label interference and background noise through a confidence mechanism; finally, it is fused into QF-Robust Loss, and the collaborative optimization of localization and classification tasks is achieved by flexibly setting the loss weights, so as to solve the problem of inconsistent classification and localization prediction of substation equipment in complex backgrounds as a whole. Description of the Drawings

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0014] Figure 1 It is a flowchart of an adaptive infrared image detection method for complex background interference provided by an embodiment of the present invention; Figure 2 It is a schematic structural diagram of the YOLO-CCA network model of a specific embodiment provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of the C2MFPPC of a specific embodiment provided by an embodiment of the present invention; Figure 4 It is a schematic structural diagram of the MFPPC of a specific embodiment provided by an embodiment of the present invention; Figure 5 It is a schematic structural diagram of the CFC of a specific embodiment provided by an embodiment of the present invention; Figure 6 It is a schematic structural diagram of the FWA of a specific embodiment provided by an embodiment of the present invention; Figure 7 It is a schematic structural diagram of the AMFF attention mechanism of a specific embodiment provided by an embodiment of the present invention; Figure 8 It is a block diagram of the structure of an adaptive infrared image detection system for complex background interference provided by an embodiment of the present invention; Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific Embodiments

[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0016] Please refer to Figure 1 , which shows a flowchart of an infrared image adaptive detection method for complex background interference in the present application.

[0017] As Figure 1 shown, the infrared image adaptive detection method for complex background interference specifically includes the following steps: Step S101, construct a YOLO-CCA network model, where the YOLO-CCA network model includes a backbone network, a neck network, and a head network.

[0018] In this step, as Figure 2 shown, the backbone network: The input end of the backbone network is an infrared electrical equipment image with a size of 640×640 and 3 channels. In the backbone network, first, the input image is reduced to a size of 160×160 and the number of channels is 128 through two Conv convolution module operations. Secondly, it successively goes through 3 C2MFPPC module operations, 1 Conv convolution module operation, 6 C2MFPPC module operations, 1 convolution module operation, 3 C2MFPPC module operations, 1 Conv convolution module operation, and 3 C2MFPPC module operations. Finally, the last three C2MFPPC modules are respectively connected to the neck network to prepare for further feature fusion. Among them, the convolution kernel size of the Conv convolution module is 3×3, and the stride is 2, which is used to extract edge and relatively low-level feature information of the image, and at the same time controls the output size by adjusting the convolution kernel size, stride, and padding. C2MFPPC consists of a series of Conv convolution modules and multiple feature parallel pooling convolution (MFPPC, Multiple feature parallel pooling convolution). The C2MFPPC module is a lightweight module that can reduce the number of parameters of the entire model. At the same time, the convolution layer of the MFPPC module is used to extract local features, and the pooling layer compresses spatial information and retains the key features of the image.

[0019] Neck Network: The neck network consists of a CFC attention mechanism module, an SPPF spatial pyramid module, an Upsample upsampling module, a Concat stacking module, a Conv convolution module, and an Adaptive Multi-Scale Feature Fusion Network (AMFF). The CFC module can help the model identify and enhance the most critical features while suppressing unimportant information. SPPF is used to improve the ability of convolutional neural networks to process inputs of different sizes. The Upsample module is not only used to enlarge the size of the feature map but also to accurately reconstruct the lost details. The Concat module is mainly used to stack the number of channels of features. The AMFF module integrates feature maps from different levels, enabling the final feature representation to contain both rich detailed information and high-level semantic information, thereby improving the model's object recognition accuracy for different scales.

[0020] Head Network: The head network consists of 3 Detect detection heads. Detect includes a Conv convolution module, 2 Conv2d convolutions, and a QF-Robust Loss function, which are used to output the category to which each image belongs and the corresponding detection accuracy.

[0021] C2MFPPC Module (as shown in Figure 3 ): First, the input feature undergoes an operation of the Conv convolution module. Subsequently, a split module is connected for channel splitting to obtain two output features. Then, the two features respectively undergo operations of the Conv convolution module and 2 MFPPC module operations, and the features obtained from the two operations are stacked in channels. Finally, after an operation of the Conv convolution module, the final output result is obtained.

[0022] MFPPC Module (as shown in Figure 4 ): First, the input feature undergoes a split channel splitting operation. Then, the channel-split features are successively subjected to 4 max pooling operations. After each max pooling, there are respectively operations of the Conv convolution module, BN normalization, and Mish activation function, thereby obtaining four groups of output features with the number of channels all being C / 4. Then, these four groups of features are stacked in channels to obtain a new feature. Finally, another operation of the Conv convolution module is performed to obtain the output result.

[0023] CFC Module (as shown in Figure 5 ): First, the input feature successively undergoes operations of the Conv convolution module and 2 FWA module operations, thereby obtaining three features. Then, the three features are stacked in channels by Concat to obtain a new feature. Finally, the new feature undergoes an operation of the Conv convolution module to obtain the output result.

[0024] The FWA module (as shown in Figure 6 ): The input feature undergoes a split channel operation once to obtain two features W1 and W2 with the number of channels both being C / 2; Feature W1 is successively subjected to a Conv convolution module, BN, and SiLU activation function operations to obtain feature W3; Feature W2 is subjected to the same Conv convolution module and BN operations twice to obtain features W4 and W5, and then feature W2 undergoes a BN operation once to obtain feature W6; Features W4, W5, and W6 are subjected to residual connection to obtain feature W7, and then feature W7 undergoes attention weighting and SiLU activation function operations to obtain feature W8; Features W3 and W8 are subjected to channel stacking and Conv convolution module operations to obtain the final output result W9.

[0025] The adaptive multi-scale feature fusion network (as shown in Figure 7 ): In the adaptive multi-scale feature fusion network, the input feature Input undergoes a Conv convolution module operation once to obtain feature Z1; Feature Z1 is respectively subjected to average pooling operations with convolution kernel sizes of 3, 5, and 7, and at the same time, three groups of Conv convolution module operations are performed to obtain features Z2, Z3, and Z4; Features Z2, Z3, and Z4 are subjected to channel stacking to obtain feature Z5, and together with the feature Z1 obtained by the first convolution, a residual network is formed to obtain feature Z6; Features Z6 and Z1 are subjected to channel stacking to obtain feature Z7; Feature Z7 is successively subjected to a Conv convolution module, BN, and SELU activation function operations to obtain the final output result Z8.

[0026] To solve the situation that it is difficult for the localization branch to accurately capture the target quality in complex backgrounds, an adaptive modulation mechanism D-QFL based on prediction error is introduced, which can fully reflect the dynamic characteristics of the prediction error, so that higher reinforcement is given to those samples with smaller prediction errors (i.e., high quality), making the network more inclined to learn the information of accurate localization, thereby improving the convergence and accuracy of the localization task.

[0027] By dynamically adjusting the exponent for adaptive weight allocation of the prediction error, strengthening the supervision of high-quality (predictions close to the true target) samples, and prompting the model to make more accurate predictions of the position and boundary of the target. To achieve the improvement of the target localization quality. For background or low-quality samples (s = 0), taking as the benchmark; for high-quality samples (s > 0), the prediction error is used as the modulation basis, so as to balance the contributions among samples of various categories and achieve the purpose of dynamic sample adjustment. The expression of the dynamic adjustment exponent quality focal loss function is: , ​ , , wherein, is the logit value output by the YOLO-CCA network model, that is, the original unactivated prediction value, is the sigmoid activation function, which is used to map the logit value to the interval [0,1], where , is the true confidence label, and its value range is [0,1], is the binary cross-entropy loss function, which measures and the true label the gap between them, is the adaptive modulation index, is the initial modulation index, set to 1-2, is the hyperparameter for adjusting the dynamic range, with a default value of 1, represents the binary cross-entropy loss when the true label =0, is the modulation index when the true label represented by =0. When s = 0, it is usually a background or non-target sample, and the model predicts its probability to be close to 0. Therefore, calculates the error when the model predicts the background class. When s = 0, it means that the sample belongs to the background class. At this time, is used to adjust the loss weight of the background sample. By adjusting the contribution of the background sample in training can be affected, enabling the model to more flexibly learn the characteristics of the background sample.

[0028] In complex substation image scenarios, there are often problems such as occlusion, blurring, and noise interference, which lead to a decline in classification performance. At the same time, missed detection (i.e., misjudging positive samples as negative samples) is also an important factor affecting detection performance. For this reason, in addition to introducing the label confidence and the regularization modulation term, NR-Focal also designs an additional weight mechanism for positive samples to amplify the penalty for positive samples with low prediction probabilities (prone to missed detection), thereby reducing the missed detection rate.

[0029] The additional factor for positive samples, when their prediction probabilities are low (i.e., the missed detection rate is high), the loss value increases significantly, thereby prompting the model to continuously improve the detection probability of positive samples to achieve the purpose of reducing the missed detection rate. By introducing the label confidence and the regularization term , the interference of noise labels is effectively reduced, and the model training is stabilized to achieve the purpose of classification robustness. The expression of the noise-robust focal loss function is: , In the formula, is the label confidence, , reflecting whether the label is accurate or a pseudo-label, is the true class label, , 1 represents the positive class, and 0 represents the negative class, is the basic weight parameter of the positive class samples, defaulting to 1, used to balance the proportion of positive and negative samples, is the adjustment parameter for reducing the missed detection rate, defaulting to 1, is the predicted probability after sigmoid activation, where , is the part where the predicted probability is consistent with the true label, where , if , then , if , then , is a constant, taking the value of , is the weight hyperparameter of the regularization term, and z is the logit value output by the model, that is, the original predicted value without activation.

[0030] It should be noted that the objective loss function in the head network, by fusing D-QFL and NR-Focal, achieves a balance between the two tasks of localization and classification, ensuring that the model accurately locates the target, reducing missed detections and noise interference, and improving the classification accuracy. In complex backgrounds, considering both the localization quality and the classification reliability, the overall performance of the model is comprehensively optimized to improve the problems of accurate localization but misclassification and missed detection of positive samples. The expression of the objective loss function is: , In the formula, is the objective loss function, is the fusion weight coefficient of the dynamic adjustment exponential quality focal loss function, is the dynamic adjustment exponential quality focal loss function, is the fusion weight coefficient of the noise-robust focal loss function, is the noise-robust focal loss function; Step S102, obtain at least one historical infrared image of a substation equipment, and annotate the at least one historical infrared image of the substation equipment to obtain annotation information corresponding to the at least one historical infrared image of the substation equipment.

[0031] In this step, before model training, an infrared image dataset of substation electrical equipment needs to be constructed and processed. This dataset is collected by drones in the actual operating environment, with a total of 1000 infrared substation equipment images collected, covering 13 common types of infrared electrical equipment in substations, including arresters, circuit breakers, busbars, bushings, clamps, conservators, current transformers, disconnectors, disconnector 1, disconnector 2, heat sinks, insulators, and transformers, etc. To ensure that the model has good detection accuracy and generalization ability, especially the stability performance in complex scenarios, the dataset needs to undergo unified annotation, enhancement, and division processing.

[0032] First, use the annotation tool LabelImg to perform rectangular box annotation on each type of electrical equipment in the image, and adopt the English labels of 13 types of electrical equipment (including equipment names such as arrester, breaker, busbar, bushing, clamp, conservator, current-transformer, disconnector, disconnector1, disconnector2, heat-sink, insulator, transformer, etc.) to ensure the accuracy of annotation. Ensure that the label information is standardized and consistent. After annotation, generate YOLO format annotation text files for subsequent model training.

[0033] Secondly, considering the dependence of deep learning on large-scale data, and the infrared images of substations are limited by the acquisition conditions, there is a problem of limited data volume. If directly used for training, it is easy to cause overfitting and affect the generalization performance of the model. Therefore, the present invention adopts a series of data augmentation strategies. Through operations such as random rotation, horizontal and vertical flipping, brightness adjustment, cropping and filling of images, it simulates the image differences caused by changes in angle, position, posture, and illumination during the actual shooting process of the equipment, and significantly enhances the robustness of the model to the target under spatial scale and direction changes. In addition, to ensure the consistency of data input, all infrared images are subjected to size normalization processing, uniformly adjusted to a preset input resolution (such as 640×640 pixels), and size consistency checks are performed.

[0034] Finally, the processed dataset is divided into a training set and a validation set in a ratio of 8:2. Among them, the training set is used for optimizing model parameters, and the validation set is used for evaluating the final model effect.

[0035] Step S103, input the at least one historical infrared image of substation equipment and the annotation information into the YOLO-CCA network model for iterative training to obtain an adaptive detection model.

[0036] In this step, the training process of the model can be roughly divided into three stages: image feature extraction, feature fusion, and classification detection. First, in the image feature extraction stage, the backbone network will identify and extract key features from the input image, such as edges, textures, and object shapes. Then, in the feature fusion stage, feature information at different levels and scales is integrated to enhance the model's ability to recognize objects of various sizes. Finally, in the classification detection stage, accurate classification and positioning of the target are achieved by predicting the category corresponding to each candidate box and the position of the bounding box. During training, the number of training epochs is set to 300, the training batch size is 8, and the learning rate is 0.01. After training, the model will generate the corresponding weight file to prepare for subsequent verification work.

[0037] As an important part of the entire object detection process, model verification aims to evaluate the performance of the model on new data. The verification process includes using the trained model for forward inference to generate prediction results with detection boxes, and calculating various evaluation metrics based on this, such as mean average precision (mAP), precision (P), recall (R), and frames per second (FPS), to measure the classification and positioning performance of the model.

[0038] Step S104: Input the obtained real-time infrared image of the substation equipment into the adaptive detection model, and the adaptive detection model outputs the detection result corresponding to the real-time infrared image of the substation equipment.

[0039] In summary, the method of this application, the proposed YOLO-CCA framework, through a series of improvements, has achieved a double improvement in accuracy and efficiency in infrared image object detection. First, by introducing more lightweight modules in the backbone network, the framework effectively reduces the number of model parameters and computational complexity. While maintaining the feature extraction ability, it greatly compresses the model size, not only improving the running efficiency but also making it more suitable for deployment on resource-constrained devices, meeting the edge computing requirements in the actual substation inspection scenario.

[0040] To overcome problems such as low contrast and strong background interference in infrared images, the present invention introduces an attention mechanism module to enhance the response to key regions while suppressing interference from irrelevant regions, thereby improving the contrast of thermal features and detection accuracy. In addition, the framework further adds a multi-level feature fusion mechanism to the network structure, enabling the model to better understand the context information between different scales, significantly improving the recognition effect of small and large targets, and enhancing the overall detection robustness.

[0041] In the design of the loss function, the Dynamic-Adjusted Quality Focal Loss (D-QFL) and the Noise-Robust Focal Loss (NR-Focal) are proposed and fused into the QF-Robust Loss to achieve the joint optimization of classification and localization tasks. The D-QFL can dynamically adjust the regression loss according to the prediction error, giving stronger learning incentives to high-quality samples and improving the convergence and accuracy of the localization branch; while the NR-Focal strengthens the attention to low-confidence positive samples by introducing the label confidence and regularization modulation mechanism, effectively reducing the missed detection rate and enhancing the model's recognition ability in complex backgrounds such as blurring and occlusion. Finally, this fused loss mechanism establishes a flexible collaborative relationship between the classification and localization tasks, solving the inconsistent problems in traditional object detection such as accurate localization but incorrect classification, and high classification confidence but localization deviation, and improving the model's prediction consistency, anti-interference ability, and detection accuracy as a whole. YOLO-CCA thus demonstrates strong practicality and robustness in complex power scenarios and can more reliably complete the high-precision detection task of infrared substation equipment.

[0042] Please refer to Figure 8 , which shows the structural block diagram of an infrared image adaptive detection system for complex background interference according to the present application.

[0043] As Figure 8 shown, the infrared image adaptive detection system 200 includes a construction module 210, an annotation module 220, a training module 230, and an output module 240.

[0044] Among them, the construction module 210 is configured to construct a YOLO-CCA network model, and the YOLO-CCA network model includes a backbone network, a neck network, and a head network; the annotation module 220 is configured to obtain at least one historical infrared image of a substation equipment and annotate the at least one historical infrared image of the substation equipment to obtain annotation information corresponding to the at least one historical infrared image of the substation equipment; the training module 230 is configured to input the at least one historical infrared image of the substation equipment and the annotation information into the YOLO-CCA network model for iterative training to obtain an adaptive detection model, where the expression of the target loss function in the adaptive detection model is: , in the formula, is the target loss function, is the fusion weight coefficient of the dynamic-adjusted quality focal loss function, is the dynamic-adjusted quality focal loss function, is the fusion weight coefficient of the noise-robust focal loss function, It is a noise-robust focal loss function; an output module 240, configured to input the acquired infrared image of the real-time power transmission and transformation equipment into the adaptive detection model, and the adaptive detection model outputs a detection result corresponding to the infrared image of the real-time power transmission and transformation equipment.

[0045] It should be understood that Figure 8 The various modules described in Figure 1 correspond to the respective steps in the method described in the reference Figure 8 Therefore, the operations, features, and corresponding technical effects described above for the method also apply to

[0046] In some other embodiments, the embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored. When the program instructions are executed by a processor, the processor is caused to execute the infrared image adaptive detection method for complex background interference in any of the above method embodiments; As an implementation manner, the computer-readable storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as: Construct a YOLO-CCA network model, and the YOLO-CCA network model includes a backbone network, a neck network, and a head network; Obtain at least one historical infrared image of the power transmission and transformation equipment, and annotate the at least one historical infrared image of the power transmission and transformation equipment to obtain annotation information corresponding to the at least one historical infrared image of the power transmission and transformation equipment; Input the at least one historical infrared image of the power transmission and transformation equipment and the annotation information into the YOLO-CCA network model for iterative training to obtain an adaptive detection model, where the expression of the target loss function in the adaptive detection model is: , In the formula, is the target loss function, is the fusion weight coefficient of the dynamic adjustment exponential quality focal loss function, is the dynamic adjustment exponential quality focal loss function, is the fusion weight coefficient of the noise-robust focal loss function, is the noise-robust focal loss function; Input the acquired infrared image of the real-time power transmission and transformation equipment into the adaptive detection model, and the adaptive detection model outputs a detection result corresponding to the infrared image of the real-time power transmission and transformation equipment.

[0047] A computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area may store an operating system and application programs required for at least one function; the storage data area may store data created according to the use of the infrared image adaptive detection system for complex background interference, etc. In addition, the computer-readable storage medium may include high-speed random access memory, and may also include a memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the computer-readable storage medium may optionally include a memory remotely set relative to the processor, and these remote memories may be connected to the infrared image adaptive detection system for complex background interference through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0048] Figure 9 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 9 shown, the device includes: a processor 310 and a memory 320. The electronic device may further include: an input device 330 and an output device 340. The processor 310, the memory 320, the input device 330, and the output device 340 may be connected through a bus or other means, Figure 9 taking connection through a bus as an example. The memory 320 is the above-mentioned computer-readable storage medium. The processor 310 executes various functional applications and data processing of the server by running non-volatile software programs, instructions, and modules stored in the memory 320, that is, implements the infrared image adaptive detection method for complex background interference in the above method embodiment. The input device 330 may receive input digital or character information, and generate key signal inputs related to user settings and function controls of the infrared image adaptive detection system for complex background interference. The output device 340 may include a display device such as a display screen.

[0049] The above electronic device may execute the method provided by the embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference may be made to the method provided by the embodiment of the present invention.

[0050] As an implementation manner, the above electronic device is applied to an infrared image adaptive detection system for complex background interference and is used for a client, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Construct a YOLO-CCA network model, and the YOLO-CCA network model includes a backbone network, a neck network, and a head network; Obtain at least one historical infrared image of a substation equipment, and annotate the at least one historical infrared image of the substation equipment to obtain annotation information corresponding to the at least one historical infrared image of the substation equipment; Input the at least one historical infrared image of the substation equipment and the annotation information into the YOLO-CCA network model for iterative training to obtain an adaptive detection model, wherein the expression of the target loss function in the adaptive detection model is: , In the formula, is the target loss function, is the fusion weight coefficient of the dynamically adjusted exponential quality focal loss function, is the dynamically adjusted exponential quality focal loss function, is the fusion weight coefficient of the noise-robust focal loss function, is the noise-robust focal loss function; Input the obtained real-time infrared image of the substation equipment into the adaptive detection model, and the adaptive detection model outputs a detection result corresponding to the real-time infrared image of the substation equipment.

[0051] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on such an understanding, the above technical solution essentially or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0052] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. An infrared image adaptive detection method for complex background interference, characterized in that Including: Construct a YOLO-CCA network model, where the YOLO-CCA network model includes a backbone network, a neck network, and a head network. Among them, the expression of the object loss function in the head network is: , In the formula, is the target loss function, is the fusion weight coefficient of the dynamically adjusted exponential quality focal loss function, is the dynamically adjusted exponential quality focal loss function, is the fusion weight coefficient of the noise-robust focal loss function, is the noise-robust focal loss function; Obtain at least one historical infrared image of a substation equipment, and annotate the at least one historical infrared image of the substation equipment to obtain annotation information corresponding to the at least one historical infrared image of the substation equipment; Input the at least one historical infrared image of the substation equipment and the annotation information into the YOLO-CCA network model for iterative training to obtain an adaptive detection model; Input the obtained real-time infrared image of the substation equipment into the adaptive detection model, and the adaptive detection model outputs a detection result corresponding to the real-time infrared image of the substation equipment.

2. The adaptive detection method for infrared images facing complex background interference according to claim 1, characterized in that, The expression of the dynamic adjustment exponential quality focal loss function is: , , , In the formula, is the logit value output by the YOLO-CCA network model, that is, the original unactivated prediction value, is the sigmoid activation function, which is used to map the logit value to the interval [0,1], where , is the true confidence label, and its value range is [0,1], is the binary cross-entropy loss function, which measures and the true label the gap between them, is the adaptive modulation index, is the initial modulation index, set to 1-2, is a hyperparameter for adjusting the dynamic range, with a default value of 1, represents the binary cross-entropy loss when the true label =0, is the modulation index when the true label =0.

3. An infrared image adaptive detection method for complex background interference according to claim 1, characterized in that, The expression of the noise-robust focal loss function is: , In the formula, is the label confidence, , which reflects whether the label is accurate or a false label, is the true class label, , where 1 represents the positive class and 0 represents the negative class, is the basic weight parameter of the positive class samples, with a default value of 1, used to balance the proportion of positive and negative samples, is the adjustment parameter for reducing the missed detection rate, with a default value of 1, is the predicted probability after sigmoid activation, where , is the part where the predicted probability is consistent with the true label, where , if , then , if , then , is a constant, taking the value of , is the weight hyperparameter of the regularization term, and z is the logit value output by the model, that is, the original unactivated predicted value.

4. An infrared image adaptive detection method for complex background interference according to claim 1, characterized in that The backbone network includes: a Conv convolution module and a C2MFPPC module; In the backbone network, the input image is reduced to a size of 160×160 with 128 channels after two operations of the Conv convolution module. Secondly, it sequentially passes through 3 operations of the C2MFPPC module, 1 operation of the Conv convolution module, 6 operations of the C2MFPPC module, 1 operation of the Conv convolution module, 3 operations of the C2MFPPC module, 1 operation of the Conv convolution module, and 3 operations of the C2MFPPC module; The C2MFPPC module includes a Conv convolution module, a split module, and an MFPPC module. Among them, the input feature passes through one operation of the Conv convolution module. Subsequently, a split module is connected for channel splitting to obtain two output features. The two features respectively pass through one operation of the Conv convolution module and 2 operations of the MFPPC module, and the features obtained from the two operations are stacked in channels. Finally, after one operation of the Conv convolution module, the final output result is obtained.

5. The adaptive detection method for infrared images facing complex background interference according to claim 4, wherein In the MFPPC module, the input feature undergoes one spilt channel splitting operation; Perform 4 max-pooling operations on the channel-split features in sequence. After each max-pooling, operations of the Conv convolution module, BN regularization, and Mish activation function are respectively performed to obtain four groups of output features with the number of channels all being C / 4; Stack the four groups of output features in channels to obtain a new feature; Perform sequential operations of the Conv convolution module on the new feature to obtain an output result.

6. The adaptive infrared image detection method for complex background interference according to claim 1, characterized in that The neck network includes: a CFC attention mechanism module, an SPPF spatial pyramid module, an Upsample upsampling module, a Concat stacking module, a Conv convolution module, and an adaptive multi-scale feature fusion network; The CFC attention mechanism module includes a Conv convolution module, an FWA module, and a Concat channel stacking module. Among them, the input feature is successively operated by the Conv convolution module and two FWA modules to obtain three features; the three features are channel-stacked by the Concat channel stacking module to obtain a new feature; and the new feature is operated by the Conv convolution module model to obtain the output result.

7. An adaptive infrared image detection method for complex background interference according to claim 6, characterized in that, In the FWA module, the input feature undergoes a split channel splitting operation to obtain two features W1 and W2 with both channel numbers of C / 2. The feature W1 is successively operated by the Conv convolution module, BN, and the SiLU activation function to obtain the feature W3. The feature W2 undergoes the same Conv convolution module and BN operations twice to obtain the features W4 and W5, and then the feature W2 undergoes a BN operation once to obtain the feature W6. The features W4, W5, and W6 are residually connected to obtain the feature W7, and then the feature W7 undergoes attention weighting and the SiLU activation function operation to obtain the feature W8. The features W3 and W8 are channel-stacked and operated by the Conv convolution module to obtain the final output result W9.

8. An infrared image adaptive detection method for complex background interference according to claim 6, characterized in that, In the adaptive multi-scale feature fusion network, the input feature Input undergoes a Conv convolution module operation once to obtain the feature Z1. The feature Z1 undergoes average pooling operations with convolution kernel sizes of 3, 5, and 7 respectively, and at the same time undergoes three groups of Conv convolution module operations to obtain the features Z2, Z3, and Z4. The features Z2, Z3, and Z4 are channel-stacked to obtain the feature Z5, and together with the feature Z1 obtained by the first convolution, a residual network is formed to obtain the feature Z6. The features Z6 and Z1 are channel-stacked to obtain the feature Z7. The feature Z7 is successively operated by the Conv convolution module, BN, and the SELU activation function to obtain the final output result Z8.

9. An infrared image adaptive detection method for complex background interference according to claim 1, characterized in that The head network includes a Detect detection head, and the Detect detection head includes a Conv convolution module, a Conv2d convolution module, and the target loss function.

10. An infrared image adaptive detection system for complex background interference, characterized in that, It includes: A construction module configured to construct a YOLO-CCA network model. The YOLO-CCA network model includes a backbone network, a neck network, and a head network. Among them, the expression of the target loss function in the head network is: , In the formula, is the target loss function, is the fusion weight coefficient of the dynamically adjusted exponential quality focal loss function, is the dynamically adjusted exponential quality focal loss function, is the fusion weight coefficient of the noise-robust focal loss function, is the noise-robust focal loss function; A labeling module configured to obtain at least one historical infrared image of a substation equipment, and label the at least one historical infrared image of the substation equipment to obtain labeling information corresponding to the at least one historical infrared image of the substation equipment. A training module configured to input the at least one historical infrared image of the substation equipment and the labeling information into the YOLO-CCA network model for iterative training to obtain an adaptive detection model. An output module configured to input the obtained real-time infrared image of the substation equipment into the adaptive detection model, and the adaptive detection model outputs a detection result corresponding to the real-time infrared image of the substation equipment.

Citation Information

Patent Citations

  • Video action detection method based on lightweight spatio-temporal joint network

    CN117475350A

  • Power transmission line infrared target detection method and system based on improved Yolov8

    CN118071676A

  • Power equipment infrared image detection method and device and medium

    CN119785071A

  • Target detection method for image collected by AR wearable device based on improved YOLOv8

    CN119942059A

  • Method for detecting infrared ship target based on improved yolov7

    US20250078541A1