Intelligent wavefront restoration algorithm full-process acceleration method for extended target

By building a lightweight neural network and optimizing the preprocessing and inference stages, the delay problem in the closed loop of wavefront recovery is solved, real-time imaging requirements for the expansion target are achieved, and the processing efficiency and accuracy of the wavefront recovery algorithm are improved.

CN120259483APending Publication Date: 2025-07-04INST OF OPTICS & ELECTRONICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510364213.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing closed-loop process of wavefront recovery is too long, making it difficult to meet the needs of real-time imaging, especially in the wavefront recovery algorithm for expanding the target, which is relatively long inference time.

Method used

Using information fusion strategies based on sub-aperture slicing and channel dimension stacking, a lightweight neural network is built, and the network is compressed through deep learning inference framework and quantization technology to optimize the computing efficiency of preprocessing and network inference parts.

Benefits of technology

It greatly improves the processing efficiency of extended target wavefront recovery, realizes real-time closed-loop correction at kilohertz level, ensures the accuracy and wavefront recovery capabilities of the network model, and is suitable for real-time imaging of adaptive optical systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259483A_ABST
    Figure CN120259483A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent wavefront restoration algorithm full-process acceleration method for an extended target, and relates to the field of wavefront restoration, and the method comprises the following steps: S110, obtaining an extended target Shack-Hartmann image, and accelerating the preprocessing process of the extended target Shack-Hartmann image, the preprocessing process being to extract a feature image; s120, dividing the extended target Shack-Hartmann image into small blocks according to sub-apertures, and stacking the small blocks along the channel dimension to form a high-latitude tensor as the input of the lightweight neural network; constructing a lightweight neural network for obtaining a mapping relation between the feature image and a Zernike coefficient; and step S130, using a deep learning reasoning framework for the lightweight neural network, and compressing the lightweight neural network by using quantization. The method comprises an overall process of an intelligent wavefront restoration algorithm, and through optimization of each stage, it is ensured that the designed lightweight network can enhance key feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wavefront restoration, and in particular to a full-process acceleration method of an intelligent wavefront restoration algorithm for extended targets. Background Art

[0002] The wavefront restoration algorithm based on the Shack-Hartmann wavefront sensor (SHWFS) is an important research direction in the fields of astronomical high-resolution imaging and microscopic imaging. The most classic method in the traditional wavefront restoration algorithm is the slope-based reconstruction method, which uses the average slope of the sub-aperture to form a slope vector and multiply it by the reconstruction matrix to obtain the aberration mode coefficient. Subsequent studies focused on introducing additional feature information to reconstruct the wavefront, such as introducing secondary wavefront curvature information and more slopes in addition to the horizontal and vertical directions. Due to the limited number of reconstruction modes of the slope method, some studies use SHWFS images to iteratively calculate aberrations, but the number of iterations required is large, time-consuming, and difficult to converge.

[0003] With the development of artificial intelligence technology, researchers in the field of optics have begun to combine deep learning technology to solve various adaptive optics problems. Hu Lejia et al. proposed SH-Net, which uses 256 input The phase distribution diagram is directly obtained from the image of size 256. When the algorithm measures the speed, it includes the whole process of slope measurement, wavefront reconstruction and wavefront upsampling, which takes 40.2ms. Wu Yu et al. proposed the SH-CNN simple model for wavefront restoration of point targets, using Tensor RT acceleration. Only the time of the inference part is tested, which can reach sub-millisecond level. For the detection of extended target wavefront, De Bruijine et al. processed SHWFS images by combining blind convolution and deep learning in 2022, and the inference time reached 99ms. The SPT-Net network proposed by Wang Ning et al. uses the wavefront images at time T and time T-1 to predict the wavefront image at time T+1, and completes the deployment and application in the actual 3km laser atmospheric transmission closed-loop system. The SPT-net inference time is 1.3ms, and the slope calculation and voltage application time are about 0.5ms. Among the above methods, the speed of the entire process is difficult to meet the high frame rate requirements of the wavefront restoration closed-loop process, or some methods only estimate the time of neural network inference, and do not make a detailed analysis and optimization of the time situation of the entire process of original input image processing, and there is still room for improvement in the neural network inference time.

[0004] The Institute of Optoelectronic Technology of the Chinese Academy of Sciences proposed an extended object depth phase recovery wavefront reconstruction method based on a Shack-Hartmann wavefront sensor in previous work. On the basis of this research, aiming at the problem of too long time delay in the above-mentioned wavefront recovery closed-loop process, a full-process method for the deployment and acceleration of an intelligent wavefront recovery algorithm for extended objects was proposed, giving full play to the parallel advantages of the GPU, and an information fusion strategy based on sub-aperture segmentation and channel dimension stacking was proposed to construct a lightweight network to achieve the computational acceleration of the adaptive optical system. Summary of the Invention

[0005] (1) The technical problem to be solved by the present invention is:

[0006] Aiming at the problem of too long time delay in the above-mentioned wavefront recovery closed-loop process, a full-process method for the deployment and acceleration of an intelligent wavefront recovery algorithm for extended objects was proposed to meet the frame rate requirement of real-time closed-loop of the intelligent wavefront recovery algorithm.

[0007] (2) To achieve the above object, the technical solution adopted by the present invention is:

[0008] A full-process acceleration method for an intelligent wavefront recovery algorithm for extended objects, including:

[0009] Step S110: Obtain the Shack-Hartmann image of the extended object, and accelerate the preprocessing process of the Shack-Hartmann image of the extended object. The preprocessing process refers to extracting the feature image;

[0010] Step S120: Cut the Shack-Hartmann image of the extended object into small pieces according to sub-apertures and stack them along the channel dimension to form a high-dimensional tensor as the input of the lightweight neural network; construct a lightweight neural network for obtaining the mapping relationship between the feature image and the Zernike coefficients;

[0011] Step S130: Use a deep learning inference framework for the lightweight neural network, and use quantization to compress the lightweight neural network.

[0012] (3) Beneficial effects:

[0013] The present invention proposes a full-process method for the deployment and acceleration of an intelligent wavefront recovery algorithm for extended objects. This solution gives full play to the computing performance of the GPU, provides optimization solutions from the preprocessing part to the network inference part, greatly improves the processing efficiency of the entire wavefront recovery process with the extended object as the input and the Zernike coefficients as the output, and can effectively solve the real-time imaging problem of the adaptive optical system using the wavefront recovery algorithm based on deep learning.

[0014] The method includes the overall process of the intelligent wavefront restoration algorithm. Through optimization at each stage, it ensures that the designed lightweight network can enhance key feature extraction, suppress irrelevant information, reduce the size of the network model, and significantly improve the inference speed to reach the kHz-level real-time closed-loop correction. At the same time, it effectively guarantees the accuracy and wavefront restoration ability of the network model, which is of great significance for the deployment and application of the intelligent wavefront restoration algorithm for extended targets in actual AO systems. Description of the Drawings

[0015] Figure 1 It is a flowchart of the full-process acceleration method for the intelligent wavefront restoration algorithm for extended targets;

[0016] Figure 2 It is the lightweight neural network structure diagram in the embodiment of the present invention;

[0017] Figure 3 It is the intelligent wavefront restoration result for extended targets in the embodiment of the present invention. Detailed Embodiments

[0018] Hereinafter, the exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0019] The following will describe in detail the specific embodiments of the full-process method for the deployment and acceleration of the intelligent wavefront restoration algorithm for extended targets proposed by the present invention with reference to the accompanying drawings.

[0020] Figure 1 It is a flowchart of the full-process acceleration method for the intelligent wavefront restoration algorithm for extended targets of the present invention. In one embodiment, the Ubuntu 18.04 LTS operating system is adopted, the CPU model is ARM8 Processor rev 0 (v8l)x6, the GPU model is NVIDIA Tegra Xavier (NV GPU) / integrated, the compilation environment is VS Code (CPU side), and the CUDA toolkit 11.3.0 (GPU side).

[0021] Under the condition, a dataset consisting of the feature maps of 100,000 extended target Shack-Hartmann images and the corresponding 299th-order Zernike coefficients is generated, where D is the telescope aperture diameter. is the atmospheric coherence length. Table 1 shows the Shack-Hartmann camera parameters used in this embodiment.

[0022] Table 1. Shack-Hartmann Camera Parameters

[0023] Step S110: Obtain the extended target Shack-Hartmann image and accelerate the preprocessing process of the extended target Shack-Hartmann image. The preprocessing process refers to extracting the feature image. This step may include: when preprocessing the extended target Shack-Hartmann image affected by atmospheric turbulence, using the standard deviation of each sub-aperture as the contrast measurement benchmark, and selecting the sub-aperture with the largest contrast as the reference sub-aperture of the image. The process of extracting the feature image through the reference sub-aperture is well-known to those skilled in the art and will not be elaborated here; during the acceleration process, use the GPU to parallel process the computationally intensive preprocessing tasks, and write a CUDA program to reconstruct and optimize the preprocessing operator according to the architecture characteristics of the GPU. The preprocessing task refers to extracting the feature image, and the CUDA stream is used to concurrently execute multiple kernel functions to improve the calculation efficiency; during the process of calculating the standard deviation, a single thread simultaneously accesses the four pixel values of the sub-aperture through the vectorized memory access technology;

[0024] When extracting the feature image, the Fourier transform and inverse Fourier transform of each sub-aperture do not interfere with each other, and the method of concurrent kernel functions is used to improve the processing speed;

[0025] When accelerating the part of extracting the feature image from the extended target Shack-Hartmann image, it is necessary to write a normalization operator to process the image. At this time, the intermediate calculation results are stored in the shared memory, the edge padding optimization method is adopted to adjust the size of the shared memory, and the block reduction method is used for parallel calculation. To avoid the time-consuming problem of syncthreads waiting caused by only warp0 running in the last iteration, the last warp of the for loop is selected for calculation, thereby further optimizing the calculation efficiency.

[0026] Step S120: Cut the extended target Shack-Hartmann image with a size of, for example, 240 240 into small pieces according to the sub-apertures and stack them along the channel dimension to form a high-dimensional tensor as the input of the lightweight neural network; construct a lightweight neural network for obtaining the mapping relationship between the feature image and the Zernike coefficients. For the stacked multi-channel input, such as Figure 2As shown, a lightweight neural network is designed. Its core module adopts the grouped convolution technique. By dividing the conventional convolution operation into multiple smaller sub-convolution groups, the number of parameters and computational complexity of the model are significantly reduced, while the parallel computing efficiency of feature extraction is improved. Each group only processes a part of the input channels. Batch normalization is used to accelerate convergence, and Relu is used to alleviate the vanishing gradient problem. Combined with the max pooling layer, downsampling is performed on the feature map. Finally, through the fully connected layer, the neural network outputs high-order Zernike coefficients for wavefront reconstruction; in one embodiment, the high order can be 299 orders.

[0027] Step S130: Use a deep learning inference framework for the lightweight neural network and compress the lightweight neural network using quantization. This step may include: using the INT8 quantization technique in the Tensor RT framework to perform INT8 numerical quantization on the lightweight neural network. The compression process may include: by collecting the statistical information of the neural network on real data, estimating the dynamic range of the weights and activation outputs of each layer in the lightweight neural network, and then determining appropriate scaling factors to minimize the quantization error. In the INT8 quantization technique, the IInt8EntropyCalibrator2 entropy calibrator is selected to achieve an efficient mapping from FP32 to INT8.

[0028] For the above parameter conditions, this paper uses the CPU and GPU to test the duration of each computational part. To ensure that the measured time can truly reflect the time delay effect of this part of the process in the entire closed-loop system, the method in this paper preheats the CPU and GPU before statistically collecting the duration data. In addition, the single calculation time during operation may jitter, but the overall distribution of the duration is relatively concentrated. To reduce the contingency of the speed measurement value, 100 SHWFS images are used to perform single-frame speed measurement in sequence, record the total time, and obtain the average value. The final obtained time results are shown in Table 2:

[0029] Table 2. Algorithm Time Comparison

[0030] Figure 3 shows the root mean square error (RMSE) performance of the algorithm on 1 group of samples. Figure 3 The left figure in shows the reconstructed wavefront phase diagram, the middle figure shows the actual wavefront phase diagram, and the right figure shows the residual wavefront phase diagram. From Figure 3 it can be seen the effectiveness of the present invention, where PV is the peak-to-valley value and RMS is the root mean square.

[0031] The above are the specific embodiments disclosed by the present invention. The parts not elaborated in detail belong to the well-known technologies in the field. However, the protection scope of the present invention is not limited thereto. Any modification, equivalent replacement, improvement, etc. made by any person skilled in the art within the technical scope disclosed by the present invention shall be covered by the protection scope of the present invention.

Claims

1. An all - process acceleration method for an intelligent wavefront restoration algorithm for extended targets, characterized in that, Including: Step S110: Obtain an extended target Shack - Hartmann image and accelerate the pre - processing process of the extended target Shack - Hartmann image. The pre - processing process refers to extracting a feature image; Step S120: Cut the extended target Shack - Hartmann image into small pieces according to sub - apertures and stack them along the channel dimension to form a high - dimensional tensor as the input of a lightweight neural network; construct a lightweight neural network for obtaining the mapping relationship between the feature image and Zernike coefficients; Step S130: Use a deep - learning inference framework for the lightweight neural network and compress the lightweight neural network using quantization.

2. The full - process acceleration method of the intelligent wavefront restoration algorithm for extended targets according to claim 1, wherein When processing the extended target Shack - Hartmann image, use the standard deviation of each sub - aperture as the contrast measurement benchmark, and select the sub - aperture with the largest contrast as the reference sub - aperture of the image.

3. The full - process acceleration method for the intelligent wave - front restoration algorithm for extended targets according to claim 1, characterized in that, During the acceleration process, use the GPU to parallel - process computationally intensive pre - processing tasks, and write a CUDA program to reconstruct and optimize the pre - processing operator according to the architecture characteristics of the GPU.

4. The full - process acceleration method for the intelligent wave - front restoration algorithm for extended targets according to claim 1, wherein, When accelerating the extended target Shack - Hartmann image, write a normalization operator to process the image. At this time, the intermediate calculation results are stored in shared memory, adopt an edge - padding optimization method to adjust the size of the shared memory, and use a block - reduction method for parallel calculation.

5. The full - process acceleration method for the intelligent wavefront restoration algorithm for extended targets according to claim 2, wherein When calculating the standard deviation, a single thread uses vectorized memory - access technology to simultaneously access four pixel values of the sub - aperture; when extracting the feature image, the Fourier transform and inverse Fourier transform of each sub - aperture do not interfere with each other, and the method of concurrent kernel functions is used to improve the processing speed.

6. The full - process acceleration method of the intelligent wavefront restoration algorithm for extended targets according to claim 1, characterized in that: The core module of the lightweight neural network divides the conventional convolution operation into multiple smaller sub - convolution groups, each group only processes part of the input channels, and combines with a pooling layer to perform downsampling on the feature image. The lightweight neural network outputs high - order Zernike coefficients for wavefront reconstruction.

7. The full - process acceleration method of the intelligent wavefront restoration algorithm for extended targets according to claim 6, characterized in that: The high - order Zernike coefficients are 299 - order Zernike coefficients.

8. The full - process acceleration method for the intelligent wavefront restoration algorithm for extended targets according to claim 1, characterized in that: Step S13 includes: Using the INT8 quantization technology in the Tensor RT framework to perform INT8 numerical quantization on the lightweight neural network.

9. The full - process acceleration method for the intelligent wavefront restoration algorithm for an extended target according to claim 8, characterized in that: The compression process includes: By collecting the statistical information of the lightweight neural network on real data, estimating the dynamic range of the weights and activation outputs of each layer in the lightweight neural network, and then determining appropriate scaling factors to minimize the quantization error.

10. The full - process acceleration method of the intelligent wavefront restoration algorithm for an extended target according to claim 8, characterized in that: In the INT8 quantization technology, select the IInt8EntropyCalibrator2 entropy calibrator to achieve the mapping from FP32 to INT8.