Infrared image target detection method based on edge embedded device
By improving the YOLOv8 algorithm architecture and model conversion quantization technology, the problem of excessive storage and memory requirements of infrared image object detection on edge embedded devices is solved, efficient and real-time infrared object detection is achieved, and small object detection capabilities are improved.
Patent Information
- Application Number
- CN202510366652.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-18
AI Technical Summary
When deploying to edge embedded devices, existing infrared image object detection methods are difficult to meet the resource limitations of the hardware platform due to excessive storage space and memory requirements, resulting in problems such as sudden frame rate drop, power consumption loss and device overheating.
The improved YOLOv8 algorithm architecture is adopted, and the spatial pyramid pooling module with a large separable nuclear attention mechanism is combined with a lightweight bidirectional cross-scale feature fusion network and a dynamic focus weighted cross-match loss function, and the resource limitations of edge embedded devices are adapted to the resource limitations of edge embedded devices through model conversion and quantization technology.
While maintaining real-time detection, it effectively solves the problem of missed detection of small targets, improves detection performance and efficiency, and is successfully deployed on edge embedded devices, suitable for resource-constrained environments.
Smart Images

Figure CN120339982A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection for autonomous driving, and particularly to an infrared image object detection method based on an edge embedded device. Background Art
[0002] Autonomous driving is an important branch in the field of artificial intelligence applications and an important part of future intelligent transportation systems. Infrared image object detection in vehicle-mounted scenarios can enhance the visual perception ability of vehicles at night, in bad weather, or in complex environments, improve the reliability of autonomous driving systems, and provide strong technical support for the development of intelligent transportation.
[0003] Currently, the research on infrared image object detection mainly includes methods based on region of interest search and methods based on threshold segmentation. The disadvantages of these two types of methods are that there is a large amount of redundant calculation, the detection efficiency is low, so the detection effect is not good in the case of complex backgrounds. At the same time, the existing technology does not consider the applicability of the algorithm to the hardware platform, and there will be a situation where the storage space required by the model trained by the algorithm and the memory called during operation are both very large. For example, edge embedded devices (such as Raspberry Pi, Jetson Nano) usually carry low-power CPUs or weakened GPUs, and generally their computing power is less than 3% of the mainstream desktop GPU (such as Jetson Nano is about 0.25 TOPS compared with RTX 3090 which is about 40 TOPS), and the memory is often less than 4GB, and there is a lack of efficient heat dissipation design; in such a resource-constrained scenario, the existing infrared image object detection methods are very likely to cause problems such as a sudden drop in frame rate, out-of-control power consumption, and overheating of the device, and it is difficult to meet the requirements for deployment to the hardware platform. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention proposes an infrared image object detection method based on an edge embedded device to solve the technical problem that when the existing infrared image object detection methods are deployed to edge embedded devices, due to the large storage space required and the memory called during operation, it is difficult to meet the deployment to the hardware platform when resources are limited.
[0005] The technical solution adopted by the present invention is as follows:
[0006] In a first aspect, an infrared image object detection method based on an edge embedded device is provided, including the following steps:
[0007] Based on the edge embedded device, allocate the memory space required when loading the infrared image object detection model; the infrared image object detection model uses an improved YOLOv8 algorithm architecture;
[0008] After sequentially performing model conversion and model quantization on the infrared image object detection model, load the model;
[0009] According to the inference acceleration judgment result, use the infrared image target detection model after model conversion and model quantization to perform infrared target detection on the acquired input image.
[0010] Furthermore, the edge embedded device adopts a heterogeneous computing architecture, including a dual-core central processing unit, memory, bus, dedicated acceleration unit, and peripheral support module.
[0011] Furthermore, the improvement of the YOLOv8 algorithm architecture includes: setting a spatial pyramid pooling module integrating the large separable kernel attention mechanism in the backbone network, setting a lightweight bidirectional cross-scale feature fusion network in the neck network, and using a dynamic focus weighted intersection over union loss function as the loss function.
[0012] Furthermore, setting the spatial pyramid pooling module integrating the large separable kernel attention mechanism in the backbone network includes:
[0013] Decompose the traditional two-dimensional large convolution kernel into a one-dimensional convolution sequence in the horizontal and vertical directions. The large separable kernel attention mechanism first extracts long-range spatial dependence relationships through one-dimensional convolution in the horizontal direction and one-dimensional convolution in the vertical direction to generate a spatial attention weight map;
[0014] Introduce a channel attention guidance mechanism, generate a channel weight vector through global average pooling, and perform channel-by-channel weighted addition with the spatial attention map.
[0015] Furthermore, setting the lightweight bidirectional cross-scale feature fusion network in the neck network includes:
[0016] Establish a bidirectional interaction path between the pixel-scale feature layers output by the backbone network, allowing high-level semantic information to be transmitted from top to bottom to the shallower layers, and low-level detailed features to enhance the high-level localization accuracy from bottom to top;
[0017] Aiming at the importance difference of features at different scales, introduce learnable normalization weights to dynamically adjust the fusion coefficients;
[0018] Remove the single-hop connections with low contribution, retain the key feature streams through cross-layer residual connections, and perform lightweight processing on the bidirectional cross-scale feature fusion network.
[0019] Furthermore, the dynamic focus weighted intersection over union loss function includes:
[0020] Dynamically adjust the penalty term weight according to the target size, and assign a higher regression priority to small targets;
[0021] Modulate the gradient normalization through the gradient magnitude clipping method;
[0022] Optimize the coupled aspect ratio constraint term through a decoupled shape optimization function.
[0023] Furthermore, the model conversion includes:
[0024] Convert the infrared image target detection model into a format model, simplify the model through dynamic dimension adaptation, and generate a static model by embedding the normalization operation of the input image into the model computation graph.
[0025] Furthermore, the model quantization includes: Parse the static model through a toolchain and generate an intermediate representation, statistically analyze the activation value distribution of each layer, dynamically generate quantization parameters, and enable a layer fusion strategy to generate a quantization model.
[0026] Furthermore, the judgment conditions for the inference acceleration include: When the ambient light brightness is greater than the set brightness threshold, do not perform infrared image target detection; when the ambient light brightness is less than or equal to the set brightness threshold, perform infrared image target detection; or
[0027] When the interval without a target in the acquired input image is greater than the set time threshold, do not perform infrared image target detection; when the interval without a target in the acquired input image is less than or equal to the set time threshold, perform infrared image target detection.
[0028] In a second aspect, an electronic device is provided, including: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the infrared image target detection method based on an edge embedded device described in the first method.
[0029] From the above technical solutions, the beneficial technical effects of the present invention are as follows:
[0030] 1. While maintaining the real-time detection ability of the model, it effectively solves the problem of missed detection of small targets in infrared target detection. Through model transfer and model quantization, this detection method is successfully loaded and deployed on edge embedded devices, improving the detection performance and efficiency of lightweight target detection models.
[0031] 2. The high-purity features output by the spatial pyramid pooling module with a large separable kernel attention mechanism, after being cross-level fused by the lightweight bidirectional cross-scale feature fusion network, the high-level semantic information reversely guides the underlying network to dynamically adjust the attention distribution, strengthening the focusing ability on thermal radiation characteristics; the two work together to effectively separate noise and features, solve complex problems that cannot be handled by a single module in the infrared scene, and significantly improve the detection ability for pixel-level small targets and occluded targets. Description of the Drawings
[0032] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0033] Figure 1 Schematic flow chart of the infrared image target detection method according to an embodiment of the present invention;
[0034] Figure 2 Schematic network structure diagram of the spatial pyramid pooling module integrating the large separable kernel attention mechanism according to an embodiment of the present invention;
[0035] Figure 3 Schematic network structure diagram of the lightweight bidirectional cross-scale feature fusion network according to an embodiment of the present invention. Specific embodiments
[0036] The following will describe in detail the embodiments of the technical solutions of the present invention with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, so they are only examples and cannot be used to limit the protection scope of the present invention.
[0037] It should be noted that unless otherwise specified, the technical terms or scientific terms used in this application should have the ordinary meanings understood by those skilled in the art to which the present invention belongs.
[0038] Embodiment
[0039] In engineering practice, the in-vehicle infrared image target detection model is generally deployed on edge embedded devices. The edge embedded device adopted in this embodiment uses a heterogeneous computing architecture, including a dual-core central processing unit (CPU), memory, bus, a dedicated acceleration unit (TPU), and a peripheral support module, as follows:
[0040] The dual-core RISC CPU, as a general-purpose processor, is responsible for task scheduling and data processing, and the core AI acceleration capability is realized through the TPU module.
[0041] As a dedicated acceleration unit, the TPU has a large number of computing cores optimized for matrix multiplication and addition operations. Through a parallel computing architecture, it significantly improves the efficiency of neural network operations such as convolution and fully connected layers. The acceleration of the neural network model by the TPU is mainly achieved through the collaborative optimization of hardware and software. At the hardware level, the TPU adopts a multi-core parallel computing design, supports single instruction multiple data stream (SIMD) operations, and can perform operations on multiple data channels simultaneously. For example, when processing the convolutional layer, the TPU will split the input feature map and weight matrix into multiple sub-blocks and distribute them to different computing cores for parallel processing, greatly shortening the computing time. In addition, the on-chip cache (SRAM) built in the TPU reduces external memory access through a data reuse mechanism, and combined with weight compression techniques (such as sparse models), further reduces the data transmission overhead.
[0042] The peripheral support module can support multiple camera inputs (such as MIPI-CSI interfaces) and video outputs (such as HDMI), and expands interfaces such as Gigabit Ethernet and USB 3.0 to meet the data transmission requirements in complex scenarios.
[0043] Based on the above edge embedded device, the infrared image target detection method provided in this embodiment includes the following steps:
[0044] S1. Based on the edge embedded device, allocate the memory space required when loading the infrared image target detection model
[0045] In a specific implementation, there is no limitation on the implementation method of allocating the memory space required for loading the infrared image target detection model, and any existing implementation method can be used, such as: directly loading the infrared image target detection model into the dedicated memory area of the TPU through memory mapping technology, and static allocation (Static Allocation) or dynamic allocation (Dynamic Allocation) can be used.
[0046] The infrared image target detection model provided in this embodiment is an improvement on the traditional YOLOv8 algorithm architecture to adapt to the application scenario of autonomous driving infrared target detection under the condition of limited hardware resources. The improvement of the YOLOv8 algorithm architecture includes the following three parts: setting a spatial pyramid pooling module integrating a large separable kernel attention mechanism in the backbone network, setting a lightweight bidirectional cross-scale feature fusion network in the neck network, and using a dynamic focus weighted intersection over union loss function as the loss function.
[0047] The specific improvement method is as follows:
[0048] 1. Set a spatial pyramid pooling module integrating a large separable kernel attention mechanism in the backbone network
[0049] The low contrast characteristics of infrared images and complex thermal noise interference lead to significant limitations in the spatial pyramid pooling module in the traditional YOLOv8 algorithm architecture during feature extraction. Although the pooling operation with fixed scales in the spatial pyramid pooling module is difficult to adaptively enhance the significant regions of infrared targets, especially the edge features of small targets and occluded targets have insufficient responses.
[0050] To solve the above problems, this paper proposes a spatial pyramid pooling module integrated with a large separable kernel attention mechanism. This module can enhance the response intensity of the target thermal radiation region dynamically by expanding the receptive field and suppress background noise interference simultaneously. Its innovative design is reflected in the following two aspects:
[0051] Firstly, the traditional two-dimensional large convolution kernel is decomposed into a one-dimensional convolution sequence in the horizontal and vertical directions. Specifically, given the input feature map The large separable kernel attention mechanism first performs one-dimensional convolution in the horizontal direction and one-dimensional convolution in the vertical direction to extract long-range spatial dependence relationships and generate a spatial attention weight map as shown in formula (1)
[0052] A = σ(K v *(K h *X)) (1)
[0053] In the above formula, A represents the spatial attention weight map, σ is the Sigmoid activation function, * represents the convolution operation, K v represents the one-dimensional convolution in the vertical direction, K h represents the one-dimensional convolution in the horizontal direction, and X represents the input feature map.
[0054] This decomposition strategy avoids the huge computational cost of directly using large kernel convolutions (the number of parameters is reduced from O(k 2 ) to O(2k)), while retaining the ability to capture the thermal diffusion characteristics of infrared targets.
[0055] Secondly, a channel attention guidance mechanism is introduced. The channel weight vector is generated through global average pooling and weighted with the spatial attention map channel by channel to further focus on the highly discriminative feature channels. Finally, the enhanced feature map X ′ is obtained, as shown in formula (2)
[0056]
[0057] In the above formula, X ′ represents the enhanced feature map, X represents the input feature map, is the element-wise multiplication, ⊙ represents the broadcast multiplication, A represents the spatial attention weight map, and w represents the channel weight vector.
[0058] The separable kernel attention mechanism can focus more precisely on the regions with the most information in the feature map. In addition, this mechanism is particularly suitable for infrared images because it dynamically assigns larger weights to prominent target feature regions, effectively enhancing the detection ability for small targets or partially occluded targets. The separable kernel attention mechanism not only improves the robustness of the model to background noise and occlusion but also optimizes the target localization accuracy, greatly enhancing the detection performance of infrared images.
[0059] Based on the above separable kernel attention mechanism, the spatial pyramid pooling module integrated with the separable kernel attention mechanism realizes multi-scale feature enhancement by cascading the spatial pyramid pooling module and the separable kernel attention mechanism. Its network structure is as Figure 2 shown. By using its large receptive field mechanism to spatially decouple background noise and suppressing interference signals through dynamic channel weights, it can effectively solve problems such as blurred target edges and background noise interference in infrared images, laying a high-quality feature expression foundation for subsequent feature fusion and loss function optimization.
[0060] 2. Set up a lightweight bidirectional cross-scale feature fusion network in the neck network
[0061] The traditional YOLOv8 algorithm architecture realizes multi-scale feature fusion through the Path Aggregation Network (PANet), but its unidirectional information flow and fixed-weight fusion strategy still face significant challenges in the infrared target detection scenario. The original PANet structure transmits features through top-down (FPN) and bottom-up (PAN) paths, but the large receptive field characteristics of high-level features (32-pixel scale) easily cause the detailed information of small targets (such as distant pedestrians) to decay during the cross-scale transmission process, resulting in insufficient target response intensity in the shallow feature map.
[0062] To solve the above technical problems, in this embodiment, a lightweight bidirectional cross-scale feature fusion network is set up in the neck network, constructing a bidirectional interaction path and a learnable weight fusion mechanism to achieve dynamic balance between shallow details and deep semantics. Its network structure is as Figure 3 shown. This feature fusion network optimizes the feature fusion efficiency through the following three stages:
[0063] The first stage: Cross-scale bidirectional connection
[0064] Establish a bidirectional interaction path among the 8, 16, and 32-pixel scale feature layers output by the backbone network, allowing high-level semantic information (such as target categories) to be transmitted from top to bottom to the shallows, while low-level detailed features (such as edge textures) enhance the high-level localization accuracy from bottom to top, realizing multi-level feature interaction. This process is shown in formula (3).
[0065]
[0066] In the above formula, represents the feature fusion function, ↑ and ↓ represent the upsampling and downsampling operations respectively, represents the input, represents the output.
[0067] The second stage: learnable feature weights
[0068] Regarding the importance difference of features at different scales, a learnable normalization weight w i (constrained to be non - negative by ReLU activation) is introduced to dynamically adjust the fusion coefficient. For the input feature map The fused feature P out is expressed as shown in formula (4).
[0069]
[0070] In the above formula, is the scale alignment operation (such as upsampling or downsampling), w represents the weight, and in a specific implementation, ∈ = 0.0001 is set to prevent numerical instability.
[0071] The third stage: redundant connection pruning
[0072] Removing the single - hop connections with low contribution in PANet (such as the 32→16 pixel cross - layer jump) reduces the computational amount, and at the same time, the key feature stream is retained through the cross - layer residual connection, which lightens the two - way cross - scale feature fusion network.
[0073] Through the processing of the above three stages, the lightweight two - way cross - scale feature fusion network constructs a fusion mechanism with two - way complementarity of shallow and deep features: the top - down path strengthens the semantic consistency of shallow features, the bottom - up path enhances the localization accuracy of deep features, and the cross - layer connection maintains the detail integrity of the original resolution. Compared with the one - way fusion PANet, the lightweight two - way cross - scale feature fusion network establishes the association effect between pixel - level thermal radiation features and semantic - level targets through weight adaptive allocation and multi - path information interaction, fuses feature information of different resolutions to enhance the multi - scale expression ability, and significantly improves the feature response intensity of small targets in infrared images.
[0074] As described above, a spatial pyramid pooling module incorporating a fused large separable kernel attention mechanism is set in the backbone network, and a lightweight bidirectional cross-scale feature fusion network is set in the neck network. The deep interaction between the two forms a synergistic effect of "forward enhancement - reverse optimization": the high-purity features output by the spatial pyramid pooling module with the large separable kernel attention mechanism are cross-level fused by the lightweight bidirectional cross-scale feature fusion network, and the high-level semantic information reversely guides the underlying network to dynamically adjust the attention distribution, enhancing the focusing ability on the thermal radiation characteristics; this synergistic mechanism realizes the effective separation of noise and features, solves the complex problems that cannot be handled by a single module in the infrared scenario, significantly improves the detection ability for pixel-level small targets and occluded targets, and there is no additional requirement for hardware resources and computing power compared with the traditional YOLOv8 algorithm architecture.
[0075] 3. The loss function adopts a dynamic focusing weighted intersection over union loss function
[0076] The traditional YOLOv8 algorithm architecture comprehensively considers the consistency of the overlapping area, the distance between the center points, and the aspect ratio through the CIoU loss function, achieving a comprehensive constraint on the bounding box regression. However, it still faces challenges in the infrared target detection scenario, such as insufficient adaptability to dynamic scenes and gradient stability. Specifically, the fixed aspect ratio penalty term of CIoU is prone to overfitting small targets (such as pedestrians in the hot spot area) under complex thermal radiation interference, and its linear distance normalization strategy is difficult to adapt to the characteristics of the drastic scale changes of targets in infrared images (such as the scale difference between a nearby vehicle and a distant pedestrian can reach more than 10 times).
[0077] To solve the above technical problems, this embodiment adopts an improved scheme of the dynamic focusing weighted intersection over union loss function in the YOLOv8 algorithm architecture, introducing a dynamic focusing mechanism and a gradient modulation strategy to achieve adaptive allocation of loss weights and precise regulation of the regression process. The improvements of the dynamic focusing weighted intersection over union loss function compared with the existing loss functions include the following three aspects
[0078] The first aspect is the dynamic scale-sensitive weight
[0079] To address the problem of the imbalance in the sensitivity of CIoU to multi-scale targets, the dynamic focusing weighted intersection over union loss function dynamically adjusts the penalty term weight according to the target size, giving higher regression priority to small targets. Define the area of the target box as S gt , then the normalized scale factor w s The calculation formula is shown in (5):
[0080]
[0081] In the above formula, S base is the average target area in the dataset.
[0082] Embed this scale factor into the center point distance penalty term as shown in Equation (6):
[0083] L distance = ρ 2 ·(1 + w s ) (6)
[0084] In the above formula, L distance represents the center point distance, and ρ represents the center point distance penalty term.
[0085] This design enables the model to automatically enhance the gradient signal strength of small targets during training and alleviate the positioning deviation of distant targets.
[0086] Second aspect, gradient normalization modulation
[0087] To suppress the abnormal gradient fluctuations caused by thermal noise in infrared images, the dynamic focusing weighted intersection over union loss function introduces a gradient amplitude clipping mechanism based on the IoU value. Define the current IoU as v, then the gradient gain factor is as shown in Equation (7):
[0088]
[0089] Apply strong gradient guidance to the prediction boxes with low overlap degree ν < 0.5 through a piecewise linear function, while reducing the update amplitude for the prediction boxes with high overlap degree (ν ≥ 0.5) to avoid overfitting.
[0090] Third aspect, decoupling of shape sensitivity factors
[0091] Improve the coupled aspect ratio constraint term in the traditional CIoU and propose a decoupled shape optimization function. Calculate the difference degrees of the width ratio and the height ratio separately:
[0092]
[0093] where λ is a learnable parameter, and the ratio offset direction is constrained by a hyperbolic function; w p represents the width of the prediction box, w gt represents the width of the ground truth box, h p represents the height of the prediction box, and h gt represents the height of the ground truth box.
[0094] The above improvements enable the infrared image target detection model to maintain the aspect ratio consistency while allowing elastic adaptation to the irregular shapes of infrared targets (such as the stretching effect of pedestrian thermal shadows).
[0095] Through the improvements in the above three aspects, the dynamic focusing weighted intersection over union loss function constructs a loss function system with scale adaptability, gradient robustness, and shape decoupling. In particular, it shows stronger anti-interference ability in high-noise night scenes, provides a more refined regression guidance strategy for infrared target detection, and breaks through the limitations of traditional geometric constraints in complex thermal imaging environments.
[0096] The dynamic focusing weighted intersection over union loss function effectively improves the positioning accuracy of infrared targets, especially small-scale and blurred targets, through the dynamic focusing strategy and gradient optimization of anchor box quality perception. Together with the aforementioned spatial pyramid pooling module integrating the large separable kernel attention mechanism and the lightweight bidirectional cross-scale feature fusion network, the three jointly constitute a collaborative optimization system for complex infrared scenes.
[0097] Combined with the improvements in the above three aspects, the improved YOLOv8 algorithm architecture neural network is trained using the existing infrared image dataset, and the trained neural network is the infrared image target detection model provided in this embodiment.
[0098] To comprehensively evaluate the performance of the infrared image target detection model provided in this embodiment, it is compared and analyzed with several existing target detection models. The experimental results are shown in Table 1. The comparative experimental data listed in the table details the average precision, number of parameters (Params / MB), etc. obtained by different models when detecting targets in the infrared image dataset. The infrared image target detection model provided in the embodiment achieves an overall detection accuracy of 76.3% in mAP@0.5, which is better than other models; it shows a lightweight advantage in terms of the number of parameters, only 3.13MB, far less than other models. The excellent detection performance, combined with the lightweight structure of the model, highlights its great potential for deployment in resource-constrained edge embedded devices.
[0099] Table 1 Comparison test results of different models
[0100]
[0101] S2. After sequentially performing model conversion and model quantization on the infrared image target detection model, load the model
[0102] When deploying and loading the vehicle-mounted infrared image target detection model into the above-mentioned edge embedded device, to overcome the defect that both the storage space required by the model and the memory called during operation are very large in the prior art, in this embodiment, the infrared image target detection model is processed by sequentially adopting model conversion and model quantization during the model deployment stage.
[0103] For model conversion, in some embodiments, the model conversion includes the following operations: converting the infrared image object detection model described in step S1 into a static ONNX model. In a specific implementation, the export.py in the infrared image object detection model is modified to adapt to the converted model, and the modified export.py is used to convert the infrared image object detection model into a format model, where the input size is preferably 640×640, and the model is simplified through dynamic dimension adaptation. During the dynamic dimension adaptation process, the consistency of the output nodes is verified in the following manner: the detection box coordinates, class probabilities, and confidence levels output by the original model are aligned with the tensor dimensions of the format model; for example, when the input size is 640×640 pixels, the tensor dimension is [1,84,8400]). In some embodiments, for the operator compatibility issue of edge embedded devices, redundant operators are pruned and optimized to ensure that the model structure meets the deployment requirements; the normalization operation of the input image (pixel value / 255.0) is embedded into the model computation graph to reduce the computational overhead of edge-side preprocessing, and finally a static model is generated to adapt to the hardware characteristics of edge embedded devices. At the same time, the static model can support different deep learning frameworks, which helps to reduce the difficulty of deploying the model on edge embedded devices.
[0104] For model quantization, in combination with the hardware resources provided by edge embedded devices, to better reduce the model's computational resource requirements, in a specific implementation, the operations for model quantization of the static model include: selecting the TPU-MLIR toolchain to complete the model quantization process, parsing the static model through model_transform.py and generating an intermediate representation MLIR file; using run_calibration.py to statistically analyze the activation value distribution of each layer and dynamically generate quantization parameters, where the quantization parameters include scaling factors and zeros; executing model_deploy.py and enabling layer fusion strategies such as merging Conv-BN-ReLU to generate a quantized model optimized from the static model; the quantized model can improve the efficiency of the next-step inference acceleration.
[0105] After applying the above technical solutions of this embodiment to engineering practice, performance tests are conducted on the quantized model, and the test data is shown in Table 2:
[0106] Table 2 Comparison test results before and after model quantization
[0107]
[0108] The above data shows that the average precision of the quantized model for object detection, mAP@0.5 and mAP@[0.5:0.95], are 74.9% and 39.9% respectively, which are basically the same as before quantization. However, the model size is reduced from 3.1M to 1.9M, a decrease of 38.7%, achieving the purpose of lightweighting. And the inference frame rate reaches 27.23FPS, meeting the real-time requirement (>25FPS) of the infrared detection system for object detection. This further proves the strong applicability of the technical solution of this embodiment to edge embedded devices, which can achieve real-time infrared object detection while maintaining high precision and is applicable to resource-constrained environments.
[0109] S3. According to the inference acceleration judgment result, use the infrared image object detection model after model conversion and model quantization to perform infrared object detection on the acquired input image.
[0110] In this embodiment, the judgment conditions for inference acceleration include: when the ambient brightness is greater than the set brightness threshold, no infrared image object detection is performed; when the ambient brightness is less than or equal to the set brightness threshold, infrared image object detection is performed; or, when the interval without a target in the acquired input image is greater than the set time threshold, no infrared image object detection is performed; when the interval without a target in the acquired input image is less than or equal to the set time threshold, infrared image object detection is performed. In a specific implementation manner, the brightness threshold only needs to be set so that the set brightness value can distinguish between day and night environments; the setting of the time threshold is matched with the hardware response speed of the edge embedded device. Through inference acceleration judgment, the memory call can be reduced when infrared object detection is not required.
[0111] In some embodiments, when the inference acceleration judgment result is that infrared image object detection is required, the infrared image object detection model after model conversion and model quantization loaded in the edge embedded device will be called to perform infrared object detection on the acquired input image, and the data output by the forward calculation will be converted into the final detection result, which includes the target category and the position coordinates of the target in the picture.
[0112] In a specific implementation manner, when this step is executed, the start and stop operations of the model inference can be achieved by sending instructions through computer serial communication. During the accelerated operation of the model inference, the edge embedded device adopted in this embodiment will adopt a CPU-TPU collaborative pipeline mode. The CPU is responsible for data preprocessing (such as image normalization, format conversion, etc.) and result postprocessing (such as non-maximum suppression), while the TPU dynamically batches the object detection process of the infrared image object detection model.
[0113] The infrared image target detection method based on edge embedded devices provided in this embodiment effectively solves the problem of missed detection of small targets in infrared target detection while maintaining the real-time detection of the model. This detection method has been successfully deployed on edge embedded devices, improving the detection performance and efficiency of lightweight target detection models.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.
Claims
1. An infrared image target detection method based on edge embedded devices, characterized in that, It includes the following steps: Based on the edge embedded device, allocate the memory space required when loading the infrared image target detection model; The infrared image target detection model uses an improved YOLOv8 algorithm architecture; After sequentially performing model conversion and model quantization on the infrared image target detection model, load the model; According to the inference acceleration judgment result, use the infrared image target detection model after model conversion and model quantization to perform infrared target detection on the acquired input image.
2. The infrared image target detection method based on an edge embedded device according to claim 1, wherein The edge embedded device adopts a heterogeneous computing architecture, including a dual-core central processing unit, memory, bus, dedicated acceleration unit, and peripheral support module.
3. The infrared image target detection method based on an edge embedded device according to claim 1, wherein The improvement of the YOLOv8 algorithm architecture includes: setting a spatial pyramid pooling module integrating a large separable kernel attention mechanism in the backbone network, setting a lightweight bidirectional cross-scale feature fusion network in the neck network, and using a dynamic focusing weighted intersection over union loss function as the loss function.
4. The infrared image target detection method based on an edge embedded device according to claim 3, characterized in that, The setting of the spatial pyramid pooling module integrating a large separable kernel attention mechanism in the backbone network includes: Decompose the traditional two-dimensional large convolution kernel into a one-dimensional convolution sequence in the horizontal and vertical directions. The large separable kernel attention mechanism first extracts long-range spatial dependence through one-dimensional convolution in the horizontal direction and one-dimensional convolution in the vertical direction to generate a spatial attention weight map; Introduce a channel attention guidance mechanism, generate a channel weight vector through global average pooling, and perform channel-by-channel weighting with the spatial attention map.
5. The infrared image target detection method based on an edge embedded device according to claim 3, characterized in that The setting of the lightweight bidirectional cross-scale feature fusion network in the neck network includes: Establish a bidirectional interaction path between the pixel-scale feature layers output by the backbone network, allowing high-level semantic information to be transmitted from top to bottom to the shallow layer, and low-level detailed features to enhance the high-level localization accuracy from bottom to top; Aiming at the importance difference of features at different scales, introduce learnable normalization weights to dynamically adjust the fusion coefficient; Remove the single-hop connections with low contribution, retain the key feature stream through cross-layer residual connections, and lightweight the bidirectional cross-scale feature fusion network.
6. The infrared image target detection method based on an edge embedded device according to claim 3, wherein The dynamic focusing weighted intersection over union loss function includes: Dynamically adjust the penalty term weight according to the target size, giving higher regression priority to small targets; Modulate the gradient normalization through the gradient magnitude clipping method; Optimize the coupled aspect ratio constraint term through a decoupled shape optimization function.
7. The infrared image target detection method based on an edge embedded device according to claim 1, wherein The model conversion includes: Convert the infrared image target detection model into a format model, simplify the model through dynamic dimension adaptation, and generate a static model by embedding the normalization operation of the input image into the model computation graph.
8. The infrared image target detection method based on an edge embedded device according to claim 7, characterized in that, The model quantization includes: parsing the static model through a toolchain to generate an intermediate representation, statistically analyzing the activation value distribution of each layer, dynamically generating quantization parameters, and enabling a layer fusion strategy to generate a quantized model.
9. The infrared image target detection method based on an edge embedded device according to claim 1, characterized in that The judgment conditions for inference acceleration include: when the external environment brightness is greater than the set brightness threshold, do not perform infrared image target detection; when the external environment brightness is less than or equal to the set brightness threshold, perform infrared image target detection; or When the interval during which the target does not appear in the acquired input image is greater than the set time threshold, infrared image target detection is not performed. When the interval during which the target does not appear in the acquired input image is less than or equal to the set time threshold, infrared image target detection is performed.
10. An electronic device, characterized in that, Including: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the infrared image target detection method based on an edge embedded device according to any one of claims 1-9.
Citation Information
Cited By
Target detection method and device based on infrared image, equipment and storage medium
CN120953593A
DUMYOLO-based infrared and visible light fusion cross-modal target detection system and method
CN121033608A