Infrared image detection system and method based on improved yolk and Laplacian pyramids

The integration of YOLO and Laplacian Pyramid networks addresses infrared target detection challenges by enhancing image details and low-frequency information, improving detection accuracy and speed for small and distant targets.

CN120318533APending Publication Date: 2025-07-15WU HAN XUAN YUAN ZHI JIA KE JI YOU XIAN GONG SI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510526742.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Infrared object detection technology is difficult to effectively detect small targets and distant pedestrians in low contrast and fuzzy environments, and there are noise interference, temperature changes, equipment differences, multi-objective recognition and real-time challenges.

Method used

The improved YOLOv8 and Laplace pyramid networks are adopted to enhance image details and low-frequency information through the Detail Processing Module (DPM), combined with the Pyramid Enhancement Network (PENet) for multi-scale feature fusion and noise suppression, improving the object detection recall and accuracy.

Benefits of technology

It effectively solves the detection problems of small and medium-sized targets and distant pedestrians in low-contrast and blurred infrared images, improves the recall and detection rate of infrared target detection, reduces the false detection rate, and meets real-time processing requirements within a wide temperature range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318533A_ABST
    Figure CN120318533A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of infrared target detection algorithms, and discloses an infrared image detection system and method based on improved yolk and Laplacian Pyramid.The infrared image detection system comprises the steps that firstly, an infrared image is obtained, user equipment receives a first infrared image shot by infrared equipment, the first infrared image is uploaded to an algorithm module, and the algorithm module sends the first infrared image to the user equipment; the algorithm module generates a second infrared image; secondly, the generated second infrared image is uploaded to a detail processing module (DPM), and the detail processing module senses and processes details of the second infrared image; analyzing a target environment in the second infrared image by using the detail processing module (DPM); the problem of low-contrast and fuzzy infrared image target detection is solved, the problem that it is difficult to extract infrared target feature information from small targets and distant pedestrians is effectively solved, and the recall rate and the detection rate of infrared target detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of infrared target detection algorithms, and specifically relates to an infrared image detection system and method based on improved YOLO and Laplacian pyramid. Background Art

[0002] Infrared target detection technology has been widely applied in many fields, such as security monitoring, autonomous driving, military reconnaissance, etc. However, it also has some defects and challenges, mainly reflected in the following aspects: 1. Low contrast and blurring problems, 2. Noise interference, 3. Influence of temperature changes on detection, 4. Resolution limitations, 5. Target occlusion and multi-target recognition, 6. Differences in equipment and environment, 7. Difficulty in data annotation, 8. Real-time performance and computational complexity, 9. Cross-domain adaptability problems, 10. Difficulty in detecting small targets, which makes it difficult to be effectively detected during the process of infrared target detection.

[0003] Although infrared target detection technology has shown great potential in many practical applications, it still faces many challenges and defects. From low contrast, noise interference to cross-device adaptability, real-time performance and other issues, continuous optimization and improvement are required in terms of algorithms, hardware, and data. Future development directions may include more efficient deep learning algorithms, more powerful sensor hardware, and the application of technologies such as data augmentation and domain adaptation to improve the accuracy, robustness, and real-time performance of infrared target detection. Summary of the Invention

[0004] The purpose of the present invention is to address the above problems. The present invention provides an infrared image detection system and method based on improved YOLO and Laplacian pyramid, which has the advantage of being able to adapt to object detection under different infrared conditions.

[0005] To achieve the above purpose, the present invention provides the following technical solution: An infrared image detection system based on improved YOLO and Laplacian pyramid, comprising:

[0006] An image acquisition module, configured to acquire the first infrared image collected;

[0007] An algorithm module, configured to process the first infrared image to generate a second infrared image;

[0008] A detail processing module, configured to sense and process the details of the second infrared image, including: analyzing the target environment in the second infrared image, distinguishing the targets and the background in the second infrared image, and obtaining the edge information of the targets in the second infrared image to generate a third infrared image;

[0009] A low-frequency enhancement module, configured to capture and enhance the low-frequency information in the third infrared image;

[0010] A detection module for detecting and analyzing the target of the processed third infrared image.

[0011] As a more preferred technical solution of the present invention, the algorithm module decomposes the first infrared image into multiple resolution components, enhances the infrared images and low-frequency information of the multiple resolution components, and forms the second infrared image.

[0012] As a more preferred technical solution of the present invention, the detail processing module is divided into a context branch and an edge branch to identify and enhance the target information in the second infrared image.

[0013] As a more preferred technical solution of the present invention, during the process of the second infrared image being processed by the context branch, context information is obtained, and the environment around the target in the second infrared image is understood by capturing long-range dependencies, forming the third infrared image.

[0014] As a more preferred technical solution of the present invention, the edge branch receives the third infrared image, identifies the contour and edge features of the target in the third infrared image, and enhances the texture information of the target components.

[0015] As a more preferred technical solution of the present invention, the low-frequency enhancement module receives the third infrared image with enhanced target edge information, performs adaptive average pooling according to the size of the third infrared image, and enables the target features to be independently processed in the form of channel separation.

[0016] As a more preferred technical solution of the present invention, the access management function node pushes an algorithm set to the associated device, and the algorithm set includes a Laplacian pyramid network with multi-scale feature fusion and a single-stage detection model YOLOv8 based on CNN.

[0017] As a more preferred technical solution of the present invention, the Laplacian pyramid decomposition can be replaced by wavelet transform, and the number of frequency bands is adjusted to 3 or 5 levels.

[0018] As a more preferred technical solution of the present invention, the edge detection algorithm of the edge branch is the Sobel operator or the Canny operator.

[0019] A method implemented by an infrared image detection system based on improved YOLO and Laplacian pyramid, including:

[0020] Step S1: Receive the first infrared image captured by the infrared device, upload the first infrared image, and generate the second infrared image;

[0021] Step S2: Upload the generated second infrared image and sense and process the details of the second infrared image;

[0022] Step S3: Analyze the target environment in the second infrared image, distinguish the target and background in the second infrared image, and form a third infrared image;

[0023] Step S4: Analyze the target and background in the third infrared image, identify the gradients of the computer images in different directions to obtain the edge information of the target, and enhance the target information;

[0024] Step S5: Analyze the enhanced target information, upload the target information, and capture and enhance the low-frequency information in the third infrared image;

[0025] Step S6: Receive the third infrared image processed in Step S5, and upload a detector to judge and analyze the target information.

[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0027] 1. By combining the Pyramid Enhancement Network (PENET) and YOLOv8, the problem of target detection in low-contrast and blurred infrared images is solved, effectively solving the problem of difficult extraction of infrared target feature information such as small targets and distant pedestrians, and improving the recall rate and detection rate of infrared target detection.

[0028] 2. By combining the Pyramid Enhancement Network (PENET) and YOLOv8, PENet decomposes the image into multiple resolution components through the Laplacian pyramid to enhance image details and low-frequency information. It includes a Detail Processing Module (DPM) for enhancing image details through context branches and edge branches, and a Low-Frequency Enhancement Module to capture low-frequency semantics and reduce high-frequency noise, reducing the probability of false detection in target detection, further improving the detection speed and accuracy, and having high practical value for weak-contrast scenarios of infrared images. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a method logic diagram of the network structure algorithm of the present invention;

[0030] Figure 2 It is an improved network structure diagram of the present invention based on improved YOLOv8 and Laplacian pyramid;

[0031] Figure 3 It is a detailed design diagram of the DPM network algorithm structure of the present invention;

[0032] Figure 4 It is a detailed design diagram of the LEF network algorithm structure of the present invention. Detailed implementation mode

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] The present invention discloses an infrared image detection system based on improved YOLO and Laplacian pyramid, including:

[0035] An image acquisition module for acquiring the first infrared image collected;

[0036] An algorithm module for processing the first infrared image to generate a second infrared image;

[0037] A detail processing module for perceiving and processing the details of the second infrared image, including: analyzing the target environment in the second infrared image, distinguishing the targets and backgrounds in the second infrared image, and obtaining the edge information of the targets in the second infrared image to generate a third infrared image;

[0038] A low-frequency enhancement module for capturing and enhancing the low-frequency information in the third infrared image;

[0039] A detection module for detecting and analyzing the targets in the processed third infrared image.

[0040] An enhancement unit for analyzing the targets and backgrounds in the third infrared image, identifying the gradients of computer images in different directions to obtain the edge information of the targets, enhancing the target information, analyzing the enhanced target information, and uploading the target information to the low-frequency enhancement module.

[0041] Furthermore, the algorithm module decomposes the first infrared image into multiple resolution components, and enhances the infrared images and low-frequency information of the multiple resolution components to form the second infrared image.

[0042] Specifically, an improved frequency-domain weighted enhancement algorithm is adopted: frequency-domain filters are designed for each decomposition level. Among them, the high-frequency components adopt an exponential gain function G(f)=A·e^(-Bf), and the low-frequency components adopt an S-shaped gain function G(f)=C / (1+e^(-D(f-E))). Adaptive enhancement is achieved through the dynamic adjustment of parameters A-E.

[0043] Further, the detail processing module (DPM) is divided into a context branch and an edge branch to identify and enhance the target information in the second infrared image.

[0044] In the embodiment, the context branch adopts a Transformer architecture, including 4 encoder layers, and each layer is provided with 8-head attention mechanism. The edge branch adopts a double-layer Sobel operator. The first layer calculates the gradients in the 45° and 135° directions, and the second layer calculates the gradients in the 0° and 90° directions to form a multi-directional edge feature map.

[0045] Further, during the process of the second infrared image being processed by the context branch, context information is obtained, and the environment around the target in the second infrared image is understood by capturing long-range dependencies to form the third infrared image.

[0046] Specifically, the long-range dependencies are established through an improved sparse attention mechanism. The maximum correlation distance is set to 1 / 4 of the feature map size, reducing the computational complexity from O(n²) to O(n logn). After testing, the processing speed is increased by 41% on a 640×512 infrared image.

[0047] Further, the edge branch receives the third infrared image, and identifies the contour and edge features of the target in the third infrared image, and enhances the texture information of the target component;

[0048] The edge branch includes two Sobel operators.

[0049] The improved Sobel operator adopts a 5×5 convolution kernel, and the weight matrix is designed as:

[0050] Horizontal direction:

[0051] [[-1,-2,0,2,1],

[0052] [-2,-4,0,4,2],

[0053] [-3,-6,0,6,3],

[0054] [-2,-4,0,4,2],

[0055] [-1,-2,0,2,1]]

[0056] The vertical direction is transposed. Through experimental verification, the response intensity of the improved operator to weak edges is increased by 37%.

[0057] Further, the low-frequency enhancement module receives the third infrared image after the target edge information is enhanced, performs adaptive average pooling according to the size of the third infrared image, and at the same time, in the form of channel separation, enables the target features to be independently processed.

[0058] Specifically, the channel separation adopts a four-way parallel structure, each way corresponding to a different pooling scale. When feature recombination is performed, a channel attention mechanism is used, and the SE module automatically learns the fusion weights of each channel. Tests on the public dataset KAIST show that this design improves the mAP by 5.2%.

[0059] Further, the access management function node pushes an algorithm set to the associated device, and the algorithm set includes a Laplacian pyramid network for multi-scale feature fusion and a single-stage detection model YOLOv8 based on CNN.

[0060] The system supports dynamic algorithm loading. When it detects that the GPU video memory is lower than 2GB, it automatically switches to the lightweight PENet-Lite version. In this version, the number of pyramid levels is reduced from 4 levels to 3 levels, and the number of channels in the DPM module is compressed by 50%. While maintaining 91% accuracy, the inference speed is increased by 2.3 times.

[0061] Specifically, an improved object detection model under infrared conditions is used, which combines a Pyramid Enhancement Network (PENet) and YOLOv8. Among them, PENet decomposes the image into multiple resolution components through a Laplacian pyramid to enhance image details and low-frequency information.

[0062] This model has three improvements based on YOLOv8: (1) Deformable convolutions are introduced in the backbone network to enhance the adaptability to deformed targets; (2) Cross-layer skip connections are added to the feature pyramid; (3) The Focal-EIoU loss function is used to improve the detection of dense small targets. When tested on the FLIR dataset, the false detection rate is reduced by 18%.

[0063] The object detection model also includes a Detail Processing Module (DPM) for enhancing image details through context branches and edge branches, and a low-frequency enhancement module for capturing low-frequency semantics and reducing high-frequency noise.

[0064] A hierarchical training strategy is adopted among the modules: First, freeze the YOLOv8 backbone and train the PENet network for 20 epochs, and then jointly fine-tune all network parameters for 10 epochs. This strategy improves the model convergence speed by 60%, and finally the mAP reaches 76.3%.

[0065] Among them, the Pyramid Enhancement Network is a key component of this design, which is used to enhance the model's detection ability for targets of different scales.

[0066] The pyramid enhancement network mainly includes the following key points:

[0067] Multi-scale feature pyramid: The pyramid enhancement network uses multiple feature pyramids of different scales, which contain feature maps from different levels. This can detect targets of different sizes simultaneously, effectively detecting both small-sized and large-sized objects.

[0068] The feature pyramid construction adopts an improved BiFPN structure, adding bidirectional paths from top to bottom and from bottom to top. Each path contains 3 convolutional layers, and depthwise separable convolutions are used to reduce the number of parameters. The overall computational cost is reduced by 37% compared to the original FPN.

[0069] Feature fusion: The pyramid enhancement network combines feature maps from different scales through feature fusion. This helps improve the model's accuracy in target localization and detection because information at different scales is effectively integrated.

[0070] The fusion process adopts attention-guided weighted fusion, calculating spatial attention weights for feature maps at each scale. The formula is: W = σ(Conv3×3([F1,F2])), where σ represents the Sigmoid function. This method increases the feature weights in important regions by 2 - 5 times.

[0071] Upsampling and downsampling: The pyramid enhancement network also includes upsampling and downsampling operations to further adjust the scale of the feature pyramid. Upsampling is used to increase the resolution to better capture the detailed information of small targets, while downsampling is used to reduce the resolution to better capture the global information of large targets.

[0072] The upsampling adopts improved sub-pixel convolution, adding an edge-guided constraint term based on the traditional ESPCN. The downsampling adopts a dual-path structure with max pooling and average pooling in parallel, and the output result selects the optimal feature through a gating mechanism.

[0073] Attention mechanism: The pyramid enhancement network also introduces an attention mechanism to enable the model to focus on the most important features, thereby further improving the detection performance. This helps reduce the cases of false detection and missed detection.

[0074] The attention module adopts a hybrid-domain design, containing two sub-modules: channel attention and spatial attention. The channel attention is improved using the SE module, and the spatial attention adopts the coordinate attention mechanism. After fine-tuning on the ImageNet pre-trained model, the classification error rate is reduced by 1.8%.

[0075] In summary, the pyramid enhancement network is one of the key innovations based on the infrared target YOLOv8 and the Laplacian pyramid improved network. Through techniques such as multi-scale feature pyramids, feature fusion, upsampling, downsampling, and attention mechanisms, the performance of the model based on the infrared target YOLOv8 and the Laplacian pyramid improved network in object detection tasks has been improved, enabling it to better handle objects of different sizes and scales.

[0076] Actual deployment tests show that within the operating temperature range of -40°C to +85°C, the detection latency of this algorithm on the HiSilicon Hi3559A chip is stable at 23.5 ± 1.2 ms, meeting the real-time processing requirements. For images with a resolution of 320×256, the power consumption is less than 1.2W.

[0077] Figure 2 Shown is the structure of DPM, including the context branch (CB) and the edge branch (EB):

[0078] The context branch captures context information by capturing long-range dependencies and globally enhances components.

[0079] The edge branch uses Sobel operators in two different directions to calculate the image gradient to obtain edges and enhance the texture of components.

[0080] The main features and functions of the low-frequency enhancement module in the infrared target YOLOv8 and Laplacian pyramid improved network are summarized as follows:

[0081] Adaptive average pooling: LEF uses adaptive average pooling operations of different sizes to intercept low-frequency components. This means that LEF can dynamically adapt to low-frequency information of different scales and semantics to ensure maximum capture of key details in the image.

[0082] Low-frequency information capture: The main task of LEF is to capture and enhance low-frequency information in the image, which contains the main semantics and key details of the image. By using a low-pass filter to filter features, LEF only allows information below the cut-off frequency to pass through, thus enhancing the low-frequency components.

[0083] Multi-scale processing: Considering the multi-scale structure of Inception, LEF applies adaptive average pooling at different sizes to adapt to low-frequency information of different semantics and scales. This helps improve the model's understanding and capture of image details.

[0084] Among them, Inception assembles multiple convolution or pooling operations into a network module. When designing a neural network, the entire network structure is assembled in units of modules. The Inception structure designs a sparse network structure, but can generate dense data, which can not only improve the performance of the neural network, but also ensure the use efficiency of computing resources.

[0085] Channel separation: LEF divides the feature f into four parts, namely {f1, f2, f3, f4}. Through the way of channel separation, each part can be processed independently to further enhance the low-frequency information.

[0086] Figure 3 The detailed information of the low-frequency enhancement module is shown. LEF consists of adaptive average pooling of different sizes and is used to intercept the low-frequency components.

[0087] Considering the multi-scale structure of Inception, we use adaptive average pooling with sizes of 1×1, 2×2, 3×3, and 6×6, and use upsampling at the end of each scale to restore the original size of the feature. The average pooling with different kernel sizes forms a low-pass filter. We divide f into four parts, namely {f1, f2, f3, f4}, by channel separation. Each part is processed using pooling of different sizes.

[0088] Furthermore, the Laplacian pyramid decomposition can be replaced by wavelet transform, and the number of frequency bands is adjusted to 3 or 5 levels. When using Haar wavelet transform, the number of decomposition layers needs to be adjusted: 3-level decomposition corresponds to 8 subbands, and 5-level decomposition corresponds to 32 subbands. Experiments show that 5-level decomposition improves the detection accuracy of small targets by 2.7% while maintaining the same computational complexity, but the memory occupancy increases by 45%.

[0089] Furthermore, the Sobel operator of the edge branch is replaced by the Canny operator. When replaced by the Canny operator, a double-threshold processing module needs to be added. The high threshold is set to the 70% quantile of the gradient amplitude, and the low threshold is 30% quantile. In a smoke interference environment, this replacement scheme improves the edge detection accuracy by 12%, but the computational time increases by 18%.

[0090] The present invention discloses an infrared image detection system based on improved YOLO and Laplacian pyramid, which can solve the problems of target detection in low-contrast and blurred infrared images, effectively solve the problems of difficult extraction of infrared target feature information such as small targets and distant pedestrians, and improve the recall rate and detection rate of infrared target detection.

[0091] YOLOv8 belongs to the latest generation of the YOLO (You Only Look Once) series. It continues the core idea of the YOLO series - quickly completing the position recognition and classification of multiple targets in an image through a single forward pass, while further optimizing the model speed, accuracy, and ease of use.

[0092] As shown in the present invention Figure 1 The present invention also discloses an infrared image detection method based on improved YOLO and Laplacian pyramid, including:

[0093] Step S1: Receive the first infrared image captured by the infrared device, and upload the first infrared image to generate a second infrared image;

[0094] Specifically, the above-mentioned user device receives the above-mentioned first infrared image captured by the infrared device, uploads the first infrared image to the above-mentioned algorithm module, decomposes the first infrared image into multiple resolution components, and enhances the infrared images and low-frequency information of the multiple resolution components to generate the above-mentioned second infrared image. Through the above technical solution, better detail retention of various information in the first infrared image, enhancement of the overall structure, improvement of the contrast, reduction of the influence of environmental noise, and cross-scale information fusion are achieved, and different processing requirements can be adapted.

[0095] Step S2: Upload the generated second infrared image, and sense and process the details of the second infrared image;

[0096] Specifically, the above-mentioned Detail Processing Module (DPM for short) is a key component for improving network object detection based on infrared target YOLOv8 and Laplacian pyramid. By uploading the second infrared image to the above-mentioned detail processing module, context branches and edge branches are used to enhance the image details, thereby enhancing the model's ability to sense and process the detailed information of the target. Among them, the context branch is responsible for obtaining the context information in the infrared image and understanding the environment around the target in the infrared image by capturing long-range dependencies. The edge branch uses two Sobel operators to calculate the gradients of the image in different directions in the infrared image, thereby obtaining the edge information of the target.

[0097] Step S3: Analyze the target environment in the second infrared image, distinguish the target and the background in the second infrared image, and form a third infrared image;

[0098] Specifically, the above-mentioned detail processing module is also used to perform a context branch on the above-mentioned second infrared image to obtain context information, understand the environment around the target by capturing long-range dependencies, so as to better understand the relationship between the target and its surrounding environment, and further improve the accuracy of target detection. Through the above technical solution, introducing context information can enable the model to better distinguish the difference between the target and the background.

[0099] Step S4: Analyze the target and background in the third infrared image, identify the gradients of the computer images in different directions to obtain the edge information of the target, and enhance the target information.

[0100] Specifically, the above-mentioned detail processing module is used to perform an edge branch on the third graphic information to obtain the edge information of the target. By obtaining the edge contour of the target in the third graphic information, the target area is enhanced, thereby ensuring that the model can better identify the contour and edge features of the target. On the one hand, the texture information of the target components is enhanced, and on the other hand, the detail recognition and detection of the target are improved by enhancing the edge information. At the same time, the comprehensive effect of the infrared image through DPM is to enhance each component of the target, including the enhancement of context information and edge information. This can enable the model to more accurately capture the detail features of the target, thereby improving the target detection performance.

[0101] Step S5: Analyze the enhanced target information, upload the target information, and capture and enhance the low-frequency information in the third infrared image.

[0102] Specifically, the low-frequency enhancement module is used to enhance the low-frequency information of the third infrared image processed by the detail processing module. By capturing the low-frequency information in the third infrared image after step S4 and enhancing the captured low-frequency information, it is ensured that the semantics and key information in the low-frequency information can be effectively recognized, and the recognition accuracy of the target information is improved.

[0103] Among them, the low-frequency enhancement module (Low-Frequency Enhancement Filter, abbreviated as LEF) is used to capture and enhance the low-frequency information in the image. These low-frequency information usually contain most of the semantics and key information of the image, which is convenient for the detector to improve the prediction of the infrared image.

[0104] Step S6: Receive the third infrared image processed in step S5, and upload the detector to judge and analyze the target information.

[0105] Specifically, since the target information to be analyzed on each infrared image is different, but each infrared image entering the detector needs to be processed through the above steps, and then recognized and detected in the imported detector. Furthermore, the detector is used to judge and identify the target information in the third infrared image, which can effectively solve the problem of difficult extraction of infrared target feature information such as small targets and distant pedestrians, and at the same time improve the recall rate and detection rate of infrared target detection.

[0106] Specifically, the basic principles of the improved network based on the infrared target YOLOv8 and the Laplacian pyramid can be divided into several key points:

[0107] Pyramid Enhancement Network (PENet): The Laplacian pyramid is used to decompose the image into different resolution components to enhance details and low-frequency information.

[0108] In this embodiment, the Laplacian pyramid decomposition adopts a four-level decomposition structure (L0 - L3), where the L0 level is the original resolution and the L3 level is the lowest resolution. Each level of decomposition generates a low-frequency component through Gaussian filtering and downsampling, and then performs a difference calculation with the high-frequency residue of the upper level to form a multi-scale feature representation. Compared with traditional single-scale feature extraction, this decomposition method can effectively retain the multi-band feature information in the image.

[0109] Detail Processing Module (DPM): It includes a context branch and an edge branch, which are specifically used to enhance the details of the image.

[0110] In this embodiment, the Laplacian pyramid decomposition adopts a four-level decomposition structure (L0 - L3), where the L0 level is the original resolution and the L3 level is the lowest resolution. Each level of decomposition generates a low-frequency component through Gaussian filtering and downsampling, and then performs a difference calculation with the high-frequency residue of the upper level to form a multi-scale feature representation. Compared with traditional single-scale feature extraction, this decomposition method can effectively retain the multi-band feature information in the image.

[0111] Low-frequency Enhancement Module: It is used to capture low-frequency semantic information while reducing high-frequency noise.

[0112] The LEF module adopts a multi-core parallel computing architecture in hardware implementation, and each pooling kernel corresponds to an independent computing unit. Through experimental verification, the 1×1 pooling kernel is used for global feature smoothing, the 6×6 pooling kernel focuses on local structure preservation, and the superposition of multi-scale pooling results can effectively suppress high-frequency noise by 23.6%.

[0113] It shows how the input image is decomposed into different levels (L0 to L3) through the Laplacian pyramid and processed by the PENet, and finally the image quality is improved for object detection. The Detail Processing Module (DPM) and the Low-frequency Enhancement Module in the figure work together to enhance the image.

[0114] It should be noted that, in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0115] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An infrared image detection system based on improved YOLO and Laplacian pyramid, characterized in that: An image acquisition module for acquiring the first infrared image collected; An algorithm module for processing the first infrared image to generate a second infrared image; A detail processing module for perceiving and processing the details of the second infrared image, including: analyzing the target environment in the second infrared image, distinguishing the target and the background in the second infrared image, and obtaining the edge information of the target in the second infrared image to generate a third infrared image; A low-frequency enhancement module for capturing and enhancing the low-frequency information in the third infrared image; A detection module for detecting and analyzing the target of the processed third infrared image.

2. The infrared image detection system based on improved YOLO and Laplacian pyramid according to claim 1, wherein: The algorithm module decomposes the first infrared image into multiple resolution components, and enhances the infrared images and low-frequency information of the multiple resolution components to form the second infrared image.

3. The infrared image detection system based on improved YOLO and Laplacian pyramid according to claim 1, wherein: The detail processing module is divided into a context branch and an edge branch to identify and enhance the target information in the second infrared image.

4. The infrared image detection system based on improved YOLO and Laplacian pyramid according to claim 3, characterized in that: During the process of the second infrared image being processed by the context branch, context information is obtained, and the environment around the target in the second infrared image is understood by capturing long-range dependencies to form the third infrared image.

5. The infrared image detection system based on improved YOLO and Laplacian pyramid according to claim 4, wherein: The edge branch receives the third infrared image, and identifies the contour and edge features of the target in the third infrared image, and enhances the texture information of the target components.

6. The infrared image detection system based on improved YOLO and Laplacian pyramid according to claim 1, characterized in that: The low-frequency enhancement module receives the third infrared image with enhanced target edge information, performs adaptive average pooling according to the size of the third infrared image, and enables the target features to be independently processed in the form of channel separation.

7. The infrared image detection system based on improved YOLO and Laplacian pyramid according to claim 1, characterized in that: The access management function node pushes an algorithm set to the associated device, and the algorithm set includes a Laplacian pyramid network with multi-scale feature fusion and a single-stage detection model YOLOv8 based on CNN.

8. The infrared image detection system based on improved YOLO and Laplacian pyramid according to claim 1, characterized in that: The Laplacian pyramid decomposition can be replaced by wavelet transform, and the number of frequency bands is adjusted to 3 or 5 levels.

9. The infrared image detection system based on improved YOLO and Laplacian pyramid according to claim 5, wherein: The edge detection algorithm of the edge branch is the Sobel operator or the Canny operator.

10. A method implemented by using the infrared image detection system based on the improved YOLO and Laplacian pyramid as described in claim 1, characterized in that, Including: Step S1: Receive the first infrared image captured by the infrared device, and upload the first infrared image to generate a second infrared image; Step S2: Upload the generated second infrared image, and perceive and process the details of the second infrared image; Step S3: Analyze the target environment in the second infrared image, distinguish the target and the background in the second infrared image, and form a third infrared image; Step S4: Analyze the target and the background in the third infrared image, identify the gradients of the computer images in different directions to obtain the edge information of the target, and enhance the target information; Step S5: Analyze the enhanced target information, and upload the target information to capture and enhance the low-frequency information in the third infrared image; Step S6: Receive the third infrared image processed in Step S5, and upload the detector to judge and analyze the target information.