Target detection joint optimization method for infrared single lens calculation imaging

The joint optimization of image restoration and target detection in infrared single-lens imaging systems addresses the challenge of reduced performance by integrating these tasks through an end-to-end algorithm, resulting in improved detection accuracy and reduced inference time.

CN120318498APending Publication Date: 2025-07-15TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478022.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In infrared single-lens imaging systems, image quality degradation leads to limited object detection performance. The existing methods lack a collaborative optimization mechanism between image recovery and object detection, which increases the computational complexity and is not suitable for high-speed and low-power scenarios.

Method used

The end-to-end joint optimization algorithm is used to deeply integrate image recovery and object detection tasks. Through feature sharing and joint optimization, a joint optimization loss function is designed, combined with image recovery and object detection network, a model compression and acceleration strategy is realized, and finally deployed on the AI chip.

Benefits of technology

It significantly reduces inference time, improves object detection performance, and realizes a lightweight infrared single-lens imaging system, suitable for high-speed and low-power application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318498A_ABST
    Figure CN120318498A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection joint optimization method for infrared single-lens calculation imaging, and relates to the technical field of infrared calculation imaging. The method comprises the following steps: firstly, acquiring and marking infrared target data, and generating an infrared single-lens simulation data set; then, a reconstruction-identification integrated network and a joint optimization loss function thereof are constructed, model training parameters are initialized and trained, and whether the model training parameters meet target detection precision requirements is detected; in the reasoning stage, whether the reasoning time reaches a preset requirement is observed, and if not, the model design, training and verification processes are repeated until a network model meeting the precision and frame frequency requirements is obtained; and finally, through a model compression and edge acceleration strategy, deploying the reconstruction-identification integrated model meeting the requirements to the AI chip. According to the target detection joint optimization method for infrared single-lens calculation imaging, the target detection precision is improved in an infrared single-lens system, and the reasoning time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of infrared computational imaging, and in particular to a target detection joint optimization method for infrared single-lens computational imaging. Background Art

[0002] Infrared optical imaging technology is an important branch of modern imaging and is widely used in many fields such as remote sensing, security monitoring, and medical diagnosis. Compared with visible light imaging, infrared imaging can maintain stable imaging performance in low light and complex environments, which makes it an irreplaceable advantage in many application scenarios. However, with the increasing demand for lightweight and low power consumption of equipment, traditional infrared optical imaging systems face the challenge of declining image quality during design. Especially in the process of single-lens lightweight design, due to the limited complexity of the lens assembly, the imaging quality is significantly reduced, which in turn has a negative impact on the performance of the subsequent target detection model, limiting the effectiveness of this technology in practical applications.

[0003] In recent years, computational optical imaging methods based on deep learning have brought new directions to solving these problems. Computational optical imaging uses information coding technology and mathematical modeling to deeply explore light field information. In addition to obtaining two-dimensional spatial light intensity distribution, it can also extract high-dimensional light field features such as phase and polarization. Unlike traditional optical imaging systems that rely on physical lens components to record information, computational optical imaging uses dedicated hardware for optical encoding and decodes light field data computationally to achieve more accurate target information extraction. However, in the infrared single-lens imaging scenario, how to achieve computational imaging with limited resources and closely integrate it with target detection tasks is still a very challenging technical problem.

[0004] Most current methods use image restoration technology to improve the quality of low-quality images before target detection, but this method separates image restoration from target detection tasks and lacks a collaborative optimization mechanism. Moreover, integrating a dedicated image restoration algorithm increases computational complexity and latency, and is not suitable for high-speed, low-power application scenarios. More importantly, how to build a positive correlation between image restoration and target detection tasks within the architecture to ensure that the two can promote each other when working together is still a research problem that needs to be solved urgently. Summary of the invention

[0005] The purpose of this invention is to propose a joint optimization method for target detection for infrared single-lens computational imaging, which uses an end-to-end joint optimization algorithm to deeply fuse image restoration and target detection tasks, reduce reasoning time, improve performance, and realize the complete technical process from algorithm design to model deployment through model compression and acceleration strategies.

[0006] To achieve the above object, the present invention proposes a joint optimization method for object detection for infrared single-lens computational imaging, and the specific steps are as follows:

[0007] Step S1, generate a simulation data set, including data annotation, image block division, convolution processing, weighted fusion, and noise addition;

[0008] Step S2, network design and definition of loss function, design an independent image restoration network and object detection network, define a feature sharing layer, combine the image restoration and object detection tasks, design a jointly optimized loss function, and set the loss penalty weight;

[0009] Step S3, initialize the model training parameters, including the random noise range, the penalty weight of the jointly optimized loss function, and the number of training epochs;

[0010] Step S4, jointly optimize the training, and monitor the training log, observe the changes in the object detection performance curve and the jointly optimized loss function curve; if the final detection accuracy does not meet the requirements, repeat Step S2, Step S3, and Step S4 until the accuracy requirements are met;

[0011] Step S5, inference time verification, perform inference on the model that meets the detection accuracy, and check whether the inference time requirements are met; if not, repeat Steps S2 to S5 until the inference time requirements are met;

[0012] Step S6, model compression and deployment, perform operator configuration, sensitivity pruning, and mixed-precision quantization on the model that meets the detection accuracy and inference time requirements, and finally deploy it to the AI chip.

[0013] Preferably, in Step S1, the specific steps for generating the simulation data set are as follows:

[0014] Step S11, obtain infrared image data, and complete the annotation of the target category and coordinates;

[0015] Step S12, center-crop the infrared image to 640×480 pixels, and divide it into 8×6 sub-blocks of 80×80 pixels;

[0016] Step S13, use the point spread function PSF calibrated by the single-lens camera to fill each sub-block and apply block convolution for processing;

[0017] Step S14, based on a preset weight function, perform weighted fusion on the convolved sub-blocks to restore the overall image;

[0018] Step S15, add noise with a mean of 0 and a variance of scale×3e -4Random noise, where scale is a hyperparameter in training, to generate a single-lens imaging simulation dataset ultimately for object detection.

[0019] Preferably, the formula for generating the simulation dataset is as follows:

[0020]

[0021] Among them, I out is the finally output simulation image, is the i,j sub-block segmented from the image, PSF i,j is the point spread function corresponding to the i,j sub-block, w i,j is the weight of each image block, η is Gaussian noise, is the convolution operation.

[0022] Preferably, in step S2, the specific steps of network design and defining the loss function are as follows:

[0023] Step S21: Design an image restoration network, adopting the Unet architecture, consisting of an encoder and a decoder;

[0024] Step S22: Design an object detection network, adopting the YOLO architecture, and connecting by adjusting the input channels and the feature sharing layer;

[0025] Step S23: Use the shared feature layer of the image restoration network to provide key feature inputs for the object detection network, and improve the overall performance through joint optimization.

[0026] Preferably, in step S21, the encoder of the image restoration network consists of 4 convolutional blocks, each convolutional block includes two groups of convolution and LeakyReLU activation functions, the convolution kernel size is 3×3, the strides are 1×1 and 2×2 respectively, and the number of channels is 4, 8, 16, 32 in sequence.

[0027] Preferably, in step S21, the decoder of the image restoration network consists of 4 upsampling blocks, each upsampling block includes a transposed convolution ConvTrans with a stride of 2×2, stacking and convolution.

[0028] Preferably, in step S22, the object detection network adopts the YOLO architecture and removes the first convolutional layer of the YOLO architecture backbone, and adjusts the input channels to adapt to the shared features.

[0029] Preferably, in step S2, the feature sharing layer is part of the convolutional layers of the encoder in the image restoration network.

[0030] Preferably, in step S2, the formula for the joint optimization loss function is as follows:

[0031]

[0032] Among them, is the overall joint optimization loss function, hyp res , hyp box , hyp obj and hyp cls are the weight coefficients of the loss function, is the mean square error between the restoration result and the actual result; is the bounding box regression loss; is the object confidence loss; is the classification loss; bs is the batch size, that is, the size of the number of samples selected for each training.

[0033] Preferably, the bounding box regression loss measures the regression error by calculating the intersection over union IoU between the predicted box and the ground truth box, and the loss function is calculated by 1 - IoU. The larger the IoU, the smaller the loss; the object confidence loss measures the credibility of the detection result by calculating the IoU between the predicted object and the ground truth object; the classification loss uses binary cross - entropy loss BCE to optimize the class prediction.

[0034] Therefore, the present invention proposes a joint optimization method for object detection for infrared single - lens computational imaging, and its beneficial effects are as follows:

[0035] (1) The joint optimization method for object detection for infrared single - lens computational imaging proposed by the present invention adopts an end - to - end joint optimization algorithm, deeply fuses the image restoration and object detection tasks, effectively reduces the redundant information between the two through feature sharing and collaborative optimization, enabling the two to enhance each other; during the model training process, the image reconstruction and object detection networks jointly participate in the joint optimization; during inference, only the feature sharing layer and the object detection network need to be inferred, thus significantly reducing the inference time and bringing a certain degree of performance improvement.

[0036] (2) Compared with the traditional multi - lens imaging system, the joint optimization method for object detection for infrared single - lens computational imaging proposed by the present invention has a smaller volume. Although the inference time increases slightly, the object detection performance is further improved; compared with the method of separately processing the single - lens image restoration and object detection tasks, the joint optimization algorithm of the present invention significantly reduces the overall inference time and significantly improves the detection performance.

[0037] (3) The joint optimization method for target detection in infrared single-lens computational imaging proposed by the present invention, through model compression and acceleration strategies such as operator configuration, sensitivity pruning, and mixed-precision quantization, deploys the integrated reconstruction and target detection model obtained by training onto the RK3588 chip, successfully realizing the complete technical process from algorithm design, model training to model deployment.

[0038] The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0039] Figure 1 It is a schematic flow chart of a joint optimization method for target detection in infrared single-lens computational imaging provided by the present invention;

[0040] Figure 2 It is a schematic flow chart of the production of simulation data of an infrared single lens for target detection provided by the present invention;

[0041] Figure 3 It is a schematic diagram of a joint optimization framework provided by the present invention;

[0042] Figure 4 It is a schematic diagram for comparing the detection effects of different imaging methods provided in the embodiments of the present invention; among them, Figure 4 (a) in is a schematic diagram of traditional imaging, Figure 4 (c) in is a schematic diagram of the detected imaging after restoration, Figure 4 (d) in is a schematic diagram of joint optimization imaging. Detailed Embodiments

[0043] To make the technical solutions, advantages, and objectives of the present invention clearer, the technical solutions of the embodiments of the present invention will be described clearly and completely below. The described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts belong to the protection scope of this application.

[0044] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the art to which the present invention pertains.

[0045] Embodiment 1

[0046] As Figure 1As shown in the figure, the present invention provides a joint optimization method for target detection for infrared single-lens computational imaging. The method uses an infrared detector with a spectral range of 8-12 μm, a focal length of 70 mm, an F value of 1.0, and integrates an uncooled infrared detector with a resolution of 640×480 pixels as the imaging system. It can be applied to the joint optimization of imaging systems with a high degree of image degradation such as infrared single-lens imaging and target detection tasks. The specific steps are as follows:

[0047] S1. As shown in Figure 2 the figure, obtain infrared image data, and complete the annotation of the target category and coordinates. Subsequently, perform convolution processing on the segmented infrared images using the point spread function calibrated by the single-lens camera. Then, based on a preset weight function, perform weighted fusion on the convolved image blocks to restore the overall image. Finally, add a certain range of random noise to obtain the final single-lens simulation image for target detection. The specific steps are as follows:

[0048] S11: Obtain infrared image data and complete the annotation of the target category and coordinates;

[0049] S12: Center-crop the image to 640×480 pixels and divide it into 8×6 sub-blocks of 80×80 pixels;

[0050] S13: Load the PSF calibrated by the single-lens camera, fill each sub-block, and apply block convolution;

[0051] S14: Based on a preset weight function, perform weighted fusion on the convolved sub-blocks to restore the overall image;

[0052] S15: Add Gaussian noise with a mean of 0 and a variance of scale×3e -4 to the image, where scale is a hyperparameter in the training to generate the final single-lens imaging simulation dataset for target detection, that is:

[0053]

[0054] where I out is the final output simulation image, is the i,j-th sub-block segmented from the clear image, with a size of 80×80 pixels; PSF i,j is the point spread function (Point Spread Function, PSF) corresponding to the i,j-th sub-block, representing the optical characteristics of the camera system. This PSF is used to simulate the blurring effect that may occur during the imaging process; represents the convolution operation, which is used to apply the PSF to the image block to simulate the optical characteristics of the single lens; w i,jis the weight of each image block, determined by a preset weight function, ensuring that each sub-block contributes differently to the overall image during the restoration process; η is the added Gaussian noise, satisfying In this embodiment, σ is scale × 3e -4 .

[0055] S2: As Figure 3 shown, design an independent image restoration network and a target detection network, and define a feature sharing layer for information transfer and task collaboration between the two. Combining the image restoration and target detection tasks, design a jointly optimized loss function and set the loss penalty weight to coordinate the optimization objectives between the two. The specific steps are as follows:

[0056] S21. The data processing part in the figure is the visualization of step S1.

[0057] S22. The image restoration network adopts the Unet structure, consisting of an encoder and a decoder. It fuses multi-scale context information through encoding and decoding operations to enhance image details and improve the restoration quality;

[0058] The image restoration network adopts the Unet structure, consisting of an encoder and a decoder. The encoder is composed of 4 cascaded convolutional blocks, and each convolutional block includes 2 groups of convolution and the LeakyReLU activation function. The convolution kernel size is 3×3, the strides are 1×1 and 2×2 respectively, and the padding is 1 for all. The encoder controls the size of the output feature map through different strides, and the number of channels of the 4 convolutional blocks are 4, 8, 16, and 32 respectively. The decoder is composed of 4 upsampling blocks, and each upsampling block includes a transposed convolution (ConvTrans) with a stride of 2×2, stacking, and convolution operations. The transposed convolution realizes the upsampling of the feature map. Each time upsampling is performed, the decoder fuses the channels with the corresponding scale feature map in the encoder, effectively avoiding the problem of gradient disappearance, and finally outputs the result through a convolutional layer.

[0059] S23. For target detection, the YOLO architecture is adopted to perform target classification, localization, and confidence prediction using its fast and efficient characteristics;

[0060] Three convolutional blocks in the Unet encoder are used as shared features and input into the target detection network. The target detection network adopts the YOLO architecture. To adapt to the input of the shared features, the first convolutional layer of the YOLO network backbone is removed, and the input channels are adjusted to ensure that the shared features can be correctly input.

[0061] S24. The feature sharing layer of the image restoration network not only provides restoration features but also serves as the input of the YOLO network, providing key features for the target detection network; through the joint optimization of multiple loss functions, the shared feature layer can learn more feature information helpful for target detection.

[0062] During model training, the image restoration and object detection networks are trained through joint optimization; during inference, only through the shared layer, without passing through the complete restoration network, thus significantly reducing the inference time of object detection.

[0063] The formula for the loss function of joint optimization is as follows:

[0064]

[0065] Among them, is the overall joint optimization loss function, hyp res , hyp box , hyp obj and hyp cls are the weight coefficients of the loss function, used to adjust the contribution of each loss term; The loss is used to calculate the mean square error between the restoration result and the actual result; the bounding box regression loss measures the regression error by calculating the intersection over union between the predicted box and the ground truth box; the object confidence loss measures the credibility of the detection result by calculating the IoU between the predicted object and the ground truth object, thereby improving the quality of the predicted box; the classification loss uses binary cross-entropy loss to optimize the class prediction to ensure that the network correctly classifies the object; the final loss function is multiplied by the batch size bs to ensure that the loss of each batch of data meets the standards of the overall training process.

[0066] In this embodiment, the hyperparameters hyp res , hyp box , hyp obj and hyp cls are 0.01, 0.05, 1.0, and 0.5 respectively, and bs is 16;

[0067] S3: Initialize the model training parameters, including the random noise range, the penalty weight of the joint optimization loss function, and the number of training epochs, etc.;

[0068] S4: Start joint optimization training, and monitor the training log, observe the changes in the object detection performance curve and the joint optimization loss function curve, and provide a basis for the next parameter adjustment. If the final detection accuracy does not meet the requirements, repeat steps S2, S3, and S4 until the accuracy requirements are met;

[0069] S5: Perform inference on the model that meets the detection accuracy, and check whether it meets the inference time requirements. If not, repeat steps S2 to S5 until the inference time requirements are met;

[0070] S6: Compress the model that meets the requirements of detection accuracy and inference time and perform edge acceleration, and finally deploy it to the AI chip; the model compression and deployment specifically include: operator configuration, sensitivity pruning, and mixed-precision quantization. The integrated reconstruction and object detection model obtained through training was deployed to the RK3588 chip, successfully realizing the complete technical process from algorithm design, model training to model deployment.

[0071] As shown in Table 1, it shows the comparison between the joint optimization method and the traditional imaging and detection methods after single-lens restoration. Specifically, compared with traditional imaging, the detection accuracy of the joint optimization method has been improved, especially in terms of mAP@0.5 and mAP@0.5:0.95. Although the inference speed is inferior to that of traditional imaging methods that do not require a restoration algorithm, the volume and weight of traditional imaging systems are large and cannot be compared with the lightweight and thin single-lens imaging system. The single-lens imaging system is suitable for more practical application scenarios due to its lightweight and compactness.

[0072] Compared with the "restore-then-detect" method of single-lens imaging, the joint optimization method shows obvious advantages in both detection accuracy and inference time. The joint optimization method can optimize both the image restoration and object detection tasks simultaneously, thus improving the detection accuracy. During inference, it only needs to pass through the feature sharing layer of the restoration network instead of the complete restoration network, thereby reducing the inference time and achieving a better balance between the performance and efficiency of the system.

[0073] Table 1 Comparison between Joint Optimization and Traditional Imaging and Detection after Restoration

[0074] Precision Recall mAP@.5 mAP@05:.95 Para speed Traditional imaging 0.826 0.742 0.802 0.506 7235389 4.53ms Detection after restoration 0.855 0.698 0.784 0.493 7370674 34.63ms Joint optimization 0.87 0.742 0.812 0.525 7032331 6.12ms

[0075] As Figure 4 shown, it shows the comparison of the effects of different imaging methods in object detection. Compared with traditional imaging methods, the imaging quality of single-lens imaging decreases due to the weakened ability to correct lens aberration. After adopting the joint optimization method, the object detection performance of single-lens imaging does not show a significant decline; compared with the single-lens imaging method of "detect after restoration", the joint optimization method has comparable detection accuracy but has an obvious advantage in inference time.

[0076] Therefore, the present invention provides a joint optimization method for object detection for infrared single-lens computational imaging, which deeply integrates the image restoration and object detection tasks using an end-to-end joint optimization algorithm, reduces the inference time, improves the performance, and realizes the complete technical process from algorithm design, model training to model deployment on the RK3588 chip through strategies such as operator configuration, sensitivity pruning, and mixed-precision quantization.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements do not enable the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A joint optimization method for target detection in infrared single-lens computational imaging, characterized in that, The specific steps are as follows: Step S1: Generate a simulation dataset, including data annotation, image chunking, convolution processing, weighted fusion, and noise addition; Step S2: Network design and define the loss function. Design an independent image restoration network and object detection network, define a feature sharing layer, combine the image restoration and object detection tasks, design a jointly optimized loss function, and set the loss penalty weight. Among them, the specific steps of network design and define the loss function are as follows: Step S21: Design an image restoration network, using the Unet architecture, consisting of an encoder and a decoder. Among them, the feature sharing layer is a partial convolutional layer of the encoder in the image restoration network; Step S22: Design an object detection network, using the YOLO architecture, and connect it to the feature sharing layer by adjusting the input channels; Step S23: Use the shared feature layer of the image restoration network to provide key feature inputs for the object detection network, and improve the overall performance through joint optimization; Step S3: Initialize the model training parameters, including the random noise range, the penalty weight of the jointly optimized loss function, and the number of training epochs; Step S4: Jointly optimize the training, monitor the training log, and observe the changes in the object detection performance curve and the jointly optimized loss function curve. If the final detection accuracy does not meet the requirements, repeat Step S2, Step S3, and Step S4 until the accuracy requirements are met; Step S5: Inference time verification. Perform inference on the model that meets the detection accuracy, and check whether the inference time requirements are met. If not, repeat Step S2 to Step S5 until the inference time requirements are met; Step S6: Model compression and deployment. Perform operator configuration, sensitivity pruning, and mixed-precision quantization on the model that meets the detection accuracy and inference time requirements, and finally deploy it to the AI chip.

2. The object detection joint optimization method for infrared single-lens computational imaging according to claim 1, characterized in that In Step S1, the specific steps for generating the simulation dataset are as follows: Step S11: Obtain infrared image data and complete the annotation of the target category and coordinates; Step S12: Center-crop the infrared image and divide it into sub-blocks of 80×80 pixels; Step S13: Use the point spread function PSF calibrated by a single-lens camera to fill each sub-block and apply block convolution for processing; Step S14: Based on a preset weight function, perform weighted fusion on the convolved sub-blocks to restore the overall image; Step S15: Add random noise to the image to generate the final single-lens imaging simulation dataset for object detection.

3. The object detection joint optimization method for infrared single-lens computational imaging according to claim 2, wherein The formula for generating the simulation dataset is as follows: Among them, I out is the finally output simulation image, is the i,j sub-block segmented from the image, and PSF i,j is the point spread function corresponding to the i,j sub-block, w i,j is the weight of each image block, η is Gaussian noise, is the convolution operation.

4. The object detection joint optimization method for infrared single-lens computational imaging according to claim 1, wherein In Step S21, the encoder of the image restoration network consists of 4 convolutional blocks. Each convolutional block includes two groups of convolution and the LeakyReLU activation function. The convolution kernel size is 3×3, the strides are 1×1 and 2×2 respectively, and the number of channels is 4, 8, 16, and 32 in sequence.

5. The joint optimization method for target detection for infrared single-lens computational imaging according to claim 1, wherein In Step S21, the decoder of the image restoration network consists of 4 upsampling blocks. Each upsampling block includes a transposed convolution ConvTrans with a stride of 2×2, stacking, and convolution.

6. The joint optimization method for target detection for infrared single-lens computational imaging according to claim 1, characterized in that In Step S22, the object detection network uses the YOLO architecture and removes the first convolutional layer of the YOLO architecture backbone, and adjusts the input channels to adapt to the shared features.

7. A target detection joint optimization method for infrared single-lens computational imaging according to claim 1, characterized in that In step S2, the formula of the jointly optimized loss function is as follows: Among them, is the overall joint optimization loss function, hyp res , hyp box , hyp obj and hyp cls are the weight coefficients of the loss function, is the mean square error between the recovery result and the actual result; is the bounding box regression loss; is the object confidence loss; is the classification loss; bs is the batch size, that is, the size of the number of samples selected for each training.

8. The joint optimization method for target detection for infrared single-lens computational imaging according to claim 7, wherein Bounding box regression loss The regression error is measured by calculating the intersection over union (IoU) between the predicted box and the ground truth box. The loss function is calculated by 1 - IoU. The larger the IoU, the smaller the loss; Object confidence loss By calculating the IoU between the predicted object and the ground truth object, the credibility of the detection result is measured; Classification loss The binary cross-entropy loss (BCE) is used to optimize the class prediction.