Optical lens automatic focusing system and method

By combining the light field camera and polarizer, collecting and processing polarized light field information, building a deep reconstruction network and performing ray tracing simulation, the automatic focus accuracy and speed problems of the existing technology in complex light and low contrast scenarios are solved, and a more efficient automatic focus effect is achieved.

CN120010086AActive Publication Date: 2025-05-16SHENZHEN YONGTAI PHOTOELECTRIC CO LTD
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
CN202510482490.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
2045-04-17

Smart Images

  • Figure CN120010086A_ABST
    Figure CN120010086A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of optical lens automatic focusing, in particular to an optical lens automatic focusing system and method. The method comprises the following steps: carrying out polarized light field acquisition through a light field camera to obtain a polarized multi-view image set; performing light field feature extraction on the polarization multi-view image set to obtain a polarization phase gradient spectrum; constructing a polarization sensing depth reconstruction network according to the polarization phase gradient spectrum to obtain a preliminary depth map; performing ray tracing simulation processing according to the initial depth map to obtain an enhanced depth map; performing depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map; acquiring optical system parameters; performing scene layering on the optimized depth map to obtain a layered scene; and performing Fresnel diffraction simulation according to the optical system parameters and the layered scene to obtain a Fresnel diffraction image set. According to the invention, the focusing precision, speed and robustness are improved through the automatic focusing technology of the optical lens.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automatic focusing of optical lenses, and in particular to an automatic focusing system and method for an optical lens. Background Art

[0002] The goal of autofocus (AF) technology is to automatically adjust the focus of the camera lens to make the captured image clearer. Mainstream autofocus technologies are mainly divided into two categories: Phase Detection Autofocus (PDAF): The principle is to set up dedicated phase detection pixels on the image sensor (or use a separate phase detection sensor). These pixels are divided into pairs, and each pair of pixels receives light from different parts of the lens. By comparing the phase difference between the images received by these pairs of pixels, it can be determined whether the focus is in front of or behind the target, as well as the direction and distance in which the lens needs to be moved.

[0003] Contrast Detection Autofocus (CDAF): The principle is to determine the focus by analyzing the contrast of the image captured by the image sensor. When the image contrast is the highest, the image is considered to be the clearest, that is, in focus. The camera finds the lens position with the highest contrast by continuously moving the lens and comparing the image contrast.

[0004] However, the existing autofocus technology is limited in focus accuracy and speed in complex lighting environments, low-contrast scenes, and when shooting fast-moving objects. Complex lighting environments cause the image sensor to be overexposed, the contrast is reduced, CDAF has difficulty finding the contrast peak, and the phase detection pixels of PDAF will also be saturated. Low-contrast scenes lack texture and details, CDAF has difficulty finding contrast changes, and PDAF's phase detection pixels are difficult to distinguish. Fast-moving objects make it difficult to predict the object's movement trajectory and the focus position cannot be predicted. Summary of the invention

[0005] Based on this, it is necessary to provide an optical lens automatic focusing system and method to solve at least one of the above technical problems.

[0006] To achieve the above object, an optical lens automatic focusing method comprises the following steps: Step S1: collecting polarized light fields through a light field camera to obtain a polarized multi-view image set; extracting light field features from the polarized multi-view image set to obtain a polarized phase gradient spectrum; Step S2: constructing a polarization-aware depth reconstruction network according to the polarization phase gradient spectrum to obtain a preliminary depth map; performing ray tracing simulation processing according to the preliminary depth map to obtain an enhanced depth map; performing depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map; Step S3: Obtain optical system parameters; perform scene stratification on the optimized depth map to obtain a stratified scene; perform Fresnel diffraction simulation according to the optical system parameters and the stratified scene to obtain a Fresnel diffraction image set; perform focal plane prediction on the Fresnel diffraction image set to obtain a focal plane confidence map; Step S4: calculating the lens displacement according to the focal plane confidence map to obtain the lens displacement; correcting the lens displacement according to the lens displacement, and generating a lens control instruction to obtain a corrected lens displacement and a lens control instruction; Step S5: moving the lens according to the lens control instruction, and determining the focus state to obtain the focus state; performing focus iteration optimization according to the focus state to obtain the optimized lens control instruction.

[0007] The present invention realizes the comprehensive capture of scene light field information by combining light field camera and polarizer, and provides richer and more reliable information for subsequent depth reconstruction and focal plane prediction, especially in complex lighting and low contrast scenes, which can significantly improve the accuracy and robustness of autofocus. By constructing a polarization-aware depth reconstruction network, and combining ray tracing simulation and data enhancement technology, the accurate reconstruction of depth information is realized, and the generalization ability of the network is improved, laying the foundation for subsequent focal plane prediction. By stratifying the scene and simulating Fresnel diffraction of the optimized depth map, the imaging effect on different depth layers can be simulated, and the accurate prediction of the focal plane position can be realized by combining light field information, thereby improving the accuracy of autofocus. By calculating and correcting the lens displacement and generating lens control instructions, the precise control of the lens motor is realized, so that the lens is moved to the predicted optimal focal plane position, and fast autofocus is realized. Through focus state judgment and iterative optimization, the gradual adjustment of the lens position is realized, the focus accuracy and efficiency are further improved, the clarity of the image is ensured, and the entire autofocus process is completed. Therefore, the present invention provides an optical lens automatic focusing method, which comprehensively utilizes polarization information, wave optics principles, light field information and deep learning technology to construct a new automatic focusing framework. This framework not only overcomes the limitations of traditional methods in complex lighting, low contrast and fast-moving object scenes, but also improves focusing accuracy, speed and robustness.

[0008] Preferably, step S1 comprises the following steps: Step S11: collecting polarized light fields through a light field camera to obtain a polarized multi-view image set; Step S12: performing viewpoint image Fourier transform on the polarization multi-view image set to obtain a Fourier transform spectrum; Step S13: performing phase gradient calculation on the Fourier transform spectrum to obtain a phase gradient map; Step S14: performing polarization phase gradient calculation on the phase gradient image to obtain a polarization phase gradient image; Step S15: generating a multi-dimensional light field feature vector for the polarization phase gradient image to obtain a polarization phase gradient spectrum.

[0009] The present invention realizes a more comprehensive and detailed sampling of the scene light field by combining a light field camera and a linear polarizer. By collecting multi-view images in multiple polarization directions, the polarization characteristics of the light in the scene can be obtained, which provides richer information for subsequent depth reconstruction and focal plane prediction. Compared with using only traditional cameras, this polarized light field acquisition can better capture the depth and structural information of the scene, especially in complex lighting environments and low-contrast scenes, and can more effectively extract scene features, thereby improving the accuracy and robustness of subsequent autofocus. By converting the viewpoint image from the spatial domain to the frequency domain, the frequency components in the image can be better analyzed. The Fourier transform spectrum provides the energy distribution of the image at different frequencies, which is crucial for subsequent phase gradient calculation and feature extraction. Frequency domain analysis can highlight the texture and detail information in the image, especially when processing low-contrast scenes, the frequency domain information can more effectively help extract the features of the image, thereby improving the accuracy of subsequent autofocus. By calculating the phase gradient of the Fourier transform spectrum, the edge and structural information of the image can be extracted. The phase gradient reflects the phase change between pixels in the image and can be used to detect the edge, texture and depth information of the image. Phase gradient images can highlight the details in the image, reduce the impact of illumination changes, and improve the accuracy of subsequent depth reconstruction and focal plane prediction. By fusing phase gradient information in different polarization directions, the polarization characteristics of the scene can be enhanced, thereby improving the accuracy of extracting depth and focus information. The polarization phase gradient image combines polarization information and phase gradient information, which can better reflect the polarization state of light in the scene, and provide more reliable features for subsequent depth reconstruction and focal plane prediction. In complex lighting environments and scenes with polarization characteristics, the polarization phase gradient can more effectively extract the characteristics of the scene, thereby improving the performance of autofocus. By converting the polarization phase gradient image into a multi-dimensional feature vector, the complex information of the light field can be compressed and encoded, which is convenient for subsequent depth reconstruction and focal plane prediction. The polarization phase gradient spectrum contains the polarization-related phase gradient information of the light field in the Fourier domain, which can comprehensively describe the characteristics of the light field and provide rich input features for the subsequent depth reconstruction network. The feature vector form is convenient for the processing of deep learning networks, which can improve the computational efficiency and the generalization ability of the model, thereby improving the performance of the entire autofocus system.

[0010] Preferably, step S11 includes the following steps: Step S111: placing an angle-calibrated linear polarizer in front of the light field camera, performing synchronous control of the light field camera and the polarizer, and obtaining synchronous control parameters; Step S112: collecting light field data in multiple polarization directions according to the synchronization control parameters to obtain an original light field image set; Step S113: correcting the effect of the polarizer on the light field on the original light field image set to obtain a corrected light field image set; Step S114: storing and marking the corrected light field image set to obtain a polarization multi-view image set.

[0011] The present invention ensures that the collected light field data has accurate polarization direction information by accurately calibrating the angle of the linear polarizer. The light field camera and the polarizer are synchronously controlled to ensure the synchronization of the data collected in different polarization directions in time, which is crucial for the subsequent fusion of multi-polarization direction information. The setting of the synchronous control parameters enables the light field camera and the polarizer to work efficiently and accurately, laying a solid foundation for the subsequent multi-polarization light field data collection, thereby improving the accuracy and robustness of the automatic focusing. By collecting multi-view images under different preset polarization directions, the light field information of the scene under different polarization states is obtained. The original light field image set contains rich polarization information, which provides a data basis for the subsequent polarization characteristic analysis and feature extraction. Collecting multi-view images can obtain the stereoscopic information of the scene, further enhancing the ability of subsequent depth reconstruction. The collection of this multi-polarization direction data enables the subsequent automatic focusing method to better adapt to complex lighting conditions and scenes with polarization characteristics, and improves the accuracy and reliability of focusing. By correcting the influence of the polarizer, the inhomogeneity brought by the polarizer to the light field data is eliminated, and the accuracy and reliability of the data are guaranteed. Correction can effectively reduce the errors caused by differences in polarizer transmittance and light loss, thereby improving the accuracy of subsequent calculations and analysis. The corrected light field image set enables subsequent processing to more accurately reflect the true polarization characteristics of the scene, further improving the performance of autofocus. Standardized data storage and labeling facilitates subsequent data management and processing. Data storage ensures that the collected light field data can be saved intact and can be used in subsequent steps. Data labeling, for example, polarization angle information is included in the file name, making it easy to identify and call data in different polarization directions. The polarization multi-view image set provides well-organized data for subsequent feature extraction, depth reconstruction, and focal plane prediction, improving the efficiency and maintainability of the entire autofocus system.

[0012] Preferably, step S2 comprises the following steps: Step S21: performing polarization phase gradient spectrum preprocessing on the polarization phase gradient spectrum to obtain a preprocessed polarization phase gradient spectrum; Step S22: constructing a polarization-aware depth reconstruction network according to the preprocessed polarization phase gradient spectrum to obtain a preliminary depth map; Step S23: performing ray tracing simulation according to the preliminary depth map and the preprocessed polarization phase gradient spectrum, and performing training data enhancement to obtain an enhanced depth map; Step S24: performing depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map.

[0013] The present invention ensures that the polarization phase gradient spectrum (PPGS) can be effectively utilized by the deep reconstruction network by preprocessing it, and improves the training efficiency and stability of the network. Preprocessing includes dimension adjustment, normalization and data type conversion, which can unify the format of input data, reduce the impact of numerical differences on training, and improve computational efficiency. Data enhancement, such as random cropping, rotation and adding noise, can increase the diversity of data, improve the generalization ability of the network, and thus improve the robustness of the autofocus system in different scenarios. A polarization-aware deep reconstruction network is constructed to achieve effective utilization of polarization information and improve the accuracy of deep reconstruction. The network structure includes an encoder, a polarization-aware module, a decoder and a loss function. The encoder extracts the multi-scale features of PPPGS, the polarization-aware module specifically processes polarization information, and the decoder maps the features to a depth map. The design of the polarization-aware module enables the network to learn the association between features in different polarization directions, thereby better capturing the depth information of the scene. Using the mean square error (MSE) as the loss function, the network can be effectively trained to accurately predict the depth map, thereby improving the accuracy of autofocus. Through ray tracing simulation and training data enhancement, the training data set is enriched, and the generalization ability and robustness of the depth reconstruction network are improved. Ray tracing simulation can generate images and depth maps similar to real scenes, which are used to simulate different lighting conditions, perspective changes and object movement. Using the data generated by ray tracing simulation to enhance the training data enables the network to perform depth reconstruction in various complex scenes, improving the adaptability of the autofocus system in different environments. By fusing the preliminary depth map and the enhanced depth map, the accuracy and reliability of the depth map are further improved. The fusion operation can combine the prior information of the preliminary depth map and the detailed information of the enhanced depth map to generate a higher quality depth map. The weighted average fusion method can dynamically adjust the weights according to the quality of the depth map to make the fusion result more accurate. The optimized depth map provides more reliable depth information for subsequent focal plane prediction, thereby improving the overall performance of autofocus.

[0014] Preferably, step S22 includes the following steps: Step S221: performing initial convolution layer feature extraction on the preprocessed polarization phase gradient spectrum to obtain a preliminary feature map; performing downsampling layer feature extraction based on the preliminary feature map to obtain a downsampled feature map; Step S222: performing polarization channel separation on the downsampled feature map to obtain separated polarization features; Step S223: introducing channel attention to the separated polarization features to obtain polarization features with channel attention; Step S224: performing polarization feature fusion on the polarization features with channel attention to obtain a fused feature map; Step S225: performing feature extraction based on the residual block encoder on the fused feature map to obtain encoder output features; Step S226: Perform decoder-based feature mapping on the encoder output features to obtain decoder output features; Step S227: Output the decoder output features as a depth map to obtain a preliminary depth map.

[0015] The present invention realizes effective feature extraction and multi-scale representation of PPPGS through the initial convolution layer and the downsampling layer. The initial convolution layer can extract preliminary features from the original data, laying the foundation for subsequent processing. The downsampling layer reduces the size of the feature map through convolution and pooling operations, and extracts higher-level abstract features, so that the network can capture the global information of the image, reduce the amount of calculation, and improve the robustness of the network. The multi-scale features enable the deep reconstruction network to better understand the depth information of the scene. Through channel separation, independent processing of information in different polarization directions is achieved, so that the network can better utilize polarization information. The separated polarization features enable the network to perform independent feature learning and processing for each polarization channel, thereby improving the sensitivity of the network to polarization information. This separation processing method can better capture the possible differences and associations between different polarization directions, thereby improving the accuracy of deep reconstruction. By introducing the channel attention mechanism, adaptive weighting of important polarization feature channels is achieved, and the performance of the network is improved. The channel attention mechanism enables the network to automatically learn the importance of features in different polarization directions, enhance important feature channels, and suppress unimportant feature channels, thereby improving the expressiveness of features and improving the accuracy of deep reconstruction. By fusing polarization features with channel attention, the information of different polarization channels is integrated, thereby enhancing the network's expressiveness. The fusion operation can integrate features of different polarization directions to form a more comprehensive feature representation, thereby improving the network's ability to understand scene information. The fused feature map contains information of different polarization directions, providing a richer feature representation for subsequent depth reconstruction. Using an encoder based on residual blocks, deeper features can be effectively extracted and the gradient vanishing problem can be alleviated, thereby improving the performance of the model. The design of residual blocks allows the network to train deeper and extract more complex features. The encoder can convert the fused feature map into a higher-level feature representation, providing more effective features for subsequent depth map generation. Through the decoder, the features extracted by the encoder are mapped to the depth map. The decoder gradually restores the spatial resolution of the feature map through deconvolution, upsampling, and skip connections. The decoder output features contain features for generating depth maps, providing necessary information for subsequent depth map output. By mapping the decoder output features to a single-channel depth map, a preliminary depth map is obtained. The preliminary depth map contains the depth information of each pixel in the scene, providing a basis for subsequent focal plane prediction.

[0016] Preferably, step S3 comprises the following steps: Step S31: performing scene stratification on the optimized depth map to obtain a layered scene; Step S32: acquiring optical system parameters; performing Fresnel diffraction simulation according to the layered scene to obtain a Fresnel diffraction image set; Step S33: extracting diffraction features from the Fresnel diffraction image set to obtain a diffraction feature map; Step S34: performing light field focusing evaluation on the layered scene according to the polarization phase gradient spectrum to obtain a light field clarity map; Step S35: performing focal plane confidence calculation on the diffraction characteristic map and the light field clarity map to obtain a focal plane confidence map.

[0017] The present invention decomposes a three-dimensional scene into multiple two-dimensional depth layers by performing scene stratification on the optimized depth map, which is convenient for subsequent Fresnel diffraction simulation and focal plane prediction. Scene stratification can simplify the complexity of subsequent processing, so that each depth layer can be analyzed independently. The construction of layered scenes makes it possible to more effectively evaluate the focusing conditions at different depth layers and improve the accuracy of focal plane prediction. Through Fresnel diffraction simulation, the propagation process of light waves at different depth layers is simulated, so that the imaging effects of different depth layers can be predicted. Fresnel diffraction simulation takes into account the wave characteristics of light waves and can more accurately simulate the diffraction phenomenon of light waves during propagation. The Fresnel diffraction image set includes diffraction images of different depth layers, which provides key data for subsequent feature extraction and focal plane prediction, thereby improving the accuracy of automatic focusing. By extracting the diffraction features of the Fresnel diffraction image, the clarity of images at different depth layers is quantified, providing a basis for subsequent focal plane confidence calculation. Diffraction features, such as image contrast, edge sharpness and frequency domain features, can reflect the clarity of the image. Extracting the diffraction feature map can effectively evaluate the imaging effect of different depth layers, provide important information for focal plane prediction, and thus improve the accuracy of autofocus. Combining the polarization phase gradient spectrum and the layered scene, the light field focus evaluation of each depth layer is realized, providing more comprehensive information for the calculation of the focal plane confidence. The light field focus evaluation combines the light field information and the depth information, and can more accurately evaluate the focusing degree of different depth layers. The light field clarity map contains the clarity evaluation results of each depth layer, which provides important supplementary information for the subsequent focal plane confidence calculation, thereby improving the robustness of autofocus. By fusing the diffraction feature map and the light field clarity map, the focal plane confidence of each depth layer is calculated, so that the optimal focal plane position can be accurately predicted. The focal plane confidence map can comprehensively consider the results of Fresnel diffraction simulation and light field focus evaluation, so as to more comprehensively and accurately evaluate the focusing degree of different depth layers. The focal plane confidence map contains the focus confidence information of different depth layers, which provides key information for the subsequent generation of lens control instructions, thereby improving the accuracy of autofocus.

[0018] Preferably, step S32 includes the following steps: Step S321: constructing an input light field according to the layered scene to obtain an input light field; Step S322: performing light field filling on the input light field to obtain a filled light field; Step S323: performing a light field Fourier transform on the filled light field to obtain a frequency domain light field; Step S324: Calculate the propagation function according to the optical system parameters and the frequency domain light field to obtain the propagation function; Step S325: performing frequency domain multiplication on the propagation function and the frequency domain light field to obtain the frequency domain light field after propagation; Step S326: performing an inverse Fourier transform on the propagated frequency domain light field to obtain a diffraction image; Step S327: performing Fresnel diffraction simulation for each depth layer according to the diffraction image and the layered scene to obtain a Fresnel diffraction image set.

[0019] The present invention converts the depth information of the layered scene into the complex amplitude distribution of the light wave by constructing an input light field, thereby providing a basis for the subsequent Fresnel diffraction simulation. The input light field can accurately describe the initial state of the light wave at different depth layers, thereby ensuring the accuracy of the diffraction simulation. The construction of the input light field enables the propagation process of the light wave to be simulated according to the depth information, thereby improving the accuracy of the automatic focusing. By filling the input light field, the edge effect in the subsequent Fourier transform and diffraction simulation is avoided, thereby ensuring the accuracy of the simulation. The filling operation can eliminate the influence of the periodic boundary conditions of the Fourier transform, thereby making the simulation results more reliable. The filled light field provides a more stable input for subsequent calculations, thereby improving the accuracy of the automatic focusing. Through the Fourier transform, the filled light field is converted from the spatial domain to the frequency domain, thereby facilitating the subsequent propagation function calculation and diffraction simulation. The frequency domain light field can describe the frequency components of the light wave, thereby simplifying the subsequent calculation process. The Fourier transform can calculate efficiently, thereby improving the efficiency of the simulation. By calculating the propagation function, the changes of light waves during propagation are described, providing key information for subsequent diffraction simulation. The propagation function can accurately describe the propagation characteristics of light waves at different depth layers, ensuring the accuracy of diffraction simulation. The calculation of the propagation function makes it possible to simulate the propagation process of light waves according to the optical system parameters and depth information, thereby improving the accuracy of autofocus. By multiplying the propagation function with the frequency domain light field in the frequency domain, the changes of light waves during propagation are simulated, and the frequency domain light field after propagation is obtained. Frequency domain multiplication can efficiently simulate the propagation process of light waves, and the simulation results are more accurate. The frequency domain light field after propagation contains the frequency domain information of the light wave after propagation, providing input for the subsequent inverse Fourier transform. Through the inverse Fourier transform, the frequency domain light field after propagation is converted from the frequency domain to the spatial domain, and a diffraction image is obtained, simulating the light intensity distribution of the light wave after propagation to the focal plane. The diffraction image can intuitively show the imaging effect of different depth layers, providing key data for subsequent feature extraction and focal plane prediction. The inverse Fourier transform can be calculated efficiently, improving the efficiency of the simulation. By performing Fresnel diffraction simulation on each depth layer in the layered scene, a Fresnel diffraction image set is obtained, which can simulate the imaging effects of different depth layers and provide key data for subsequent focal plane prediction. The Fresnel diffraction image set contains diffraction images of different depth layers and can fully reflect the light intensity distribution of different depth layers. The Fresnel diffraction simulation of each depth layer enables the clarity evaluation of different depth layers and improves the accuracy of focal plane prediction.

[0020] Preferably, step S4 comprises the following steps: Step S41: determining the optimal focal plane position on the focal plane confidence map to obtain the optimal focal plane depth; Step S42: obtaining camera parameters; performing lens displacement calculation according to the optimal focal plane depth and the camera parameters to obtain the lens displacement; Step S43: performing lens displacement correction according to the lens displacement amount and the optical system parameters to obtain a corrected lens displacement; Step S44: generating a lens control instruction for the corrected lens displacement to obtain a lens control instruction.

[0021] The present invention predicts the clearest focal plane position in a scene by determining the optimal focal plane depth, and provides a target for subsequent lens control. The processing of the focal plane confidence map can quickly and accurately find the depth value with the highest confidence, that is, the optimal focal plane depth. The determination of the optimal focal plane depth enables the lens to be adjusted to the clearest position, thereby improving the accuracy of automatic focusing. By calculating the lens displacement, the distance that the lens needs to move is determined, which provides a basis for subsequent lens control. The calculation of the lens displacement takes into account the focal length, sensor size and other parameters of the camera, thereby ensuring the accuracy of the calculation. The calculation of the lens displacement enables the lens to be moved to a position corresponding to the optimal focal plane depth, thereby achieving automatic focusing. Through lens displacement correction, the accuracy of lens displacement is improved, thereby improving the accuracy of focusing. The lens displacement correction takes into account optical system parameters such as aperture value and zoom information, thereby making the lens displacement more accurate. The corrected lens displacement can more accurately control the lens movement, thereby achieving more accurate automatic focusing. By generating a lens control instruction, the control of the lens motor is achieved, thereby controlling the lens to move to the target position. The lens control command converts the corrected lens displacement into a signal that the lens motor can understand, so as to drive the lens motor. The generation of the lens control command enables automatic control of the lens movement, thereby achieving automatic focusing.

[0022] Preferably, step S5 comprises the following steps: Step S51: moving the lens according to the lens control instruction and acquiring the current image to obtain the current image; Step S52: Evaluate the image clarity of the current image to obtain an image clarity value; Step S53: using a preset clarity threshold to determine the focus state of the image clarity value to obtain the focus state; Step S54: performing focus iterative optimization according to the focus state, the image clarity value, the focal plane confidence map and the corrected lens displacement to obtain an optimized lens control instruction.

[0023] The present invention realizes precise control of the lens position by moving the lens according to the lens control instruction, and collects the current image, providing a data basis for subsequent focus state judgment and iterative optimization. The lens movement can move the lens to the target position and collect the current image for evaluating the clarity of the image, thereby judging the focus state. Collecting the current image can obtain the scene information under the current lens position, providing a data basis for subsequent processing. By evaluating the image clarity of the current image, the clarity of the image is quantified, providing a basis for subsequent focus state judgment. The image clarity evaluation can effectively evaluate the clarity of the image and obtain the image clarity value for judging the focus state. The calculation of the image clarity value enables the clarity of the image to be quantified, providing a reference for subsequent iterative optimization. By comparing the image clarity value with the clarity threshold, the focus state is judged, providing a control condition for subsequent iterative optimization. The focus state judgment can determine whether the current image is already focused, and provides a control condition for subsequent iterative optimization. The focus state judgment enables the determination of whether further iterative optimization is required, thereby achieving automatic focusing. Through iterative focus optimization, the lens position is gradually adjusted, improving the accuracy and efficiency of focusing. Iterative focus optimization can gradually adjust the lens position according to the focus state, image clarity value, focal plane confidence map and corrected lens displacement. Iterative optimization can continuously adjust the lens position until the image clarity reaches the preset threshold or the preset number of iterations is reached, thereby improving the accuracy and efficiency of automatic focusing.

[0024] Preferably, the present invention further provides an optical lens automatic focusing system, which is used to execute the optical lens automatic focusing method as described above, and the optical lens automatic focusing system comprises: The light field feature extraction module is used to collect polarized light fields through a light field camera to obtain a polarized multi-view image set; extract light field features from the polarized multi-view image set to obtain a polarization phase gradient spectrum; The depth information reconstruction module is used to construct a polarization-aware depth reconstruction network according to the polarization phase gradient spectrum to obtain a preliminary depth map; perform ray tracing simulation processing according to the preliminary depth map to obtain an enhanced depth map; perform depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map; The focal plane prediction module is used to obtain optical system parameters; perform scene stratification on the optimized depth map to obtain a stratified scene; perform Fresnel diffraction simulation based on the optical system parameters and the stratified scene to obtain a Fresnel diffraction image set; perform focal plane prediction on the Fresnel diffraction image set to obtain a focal plane confidence map; The lens control instruction generation module is used to calculate the lens displacement according to the focal plane confidence map to obtain the lens displacement; perform lens displacement correction according to the lens displacement, and generate the lens control instruction to obtain the corrected lens displacement and the lens control instruction; The focus confirmation and optimization module is used to move the lens according to the lens control instruction, and to judge the focus state to obtain the focus state; and to perform focus iterative optimization according to the focus state to obtain the optimized lens control instruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A schematic diagram of the steps of an automatic focusing method for an optical lens; Figure 2 It is a schematic diagram of the detailed implementation steps of step S2 in the present invention.

[0026] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION

[0027] The technical method of the present invention is described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by technicians in this field without creative work are within the scope of protection of the present invention.

[0028] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.

[0029] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are used only to distinguish one unit from another unit. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0030] To achieve this, please refer to Figure 1 to Figure 2 , an optical lens automatic focusing method, comprising the following steps: Step S1: collecting polarized light fields through a light field camera to obtain a polarized multi-view image set; extracting light field features from the polarized multi-view image set to obtain a polarized phase gradient spectrum; Step S2: constructing a polarization-aware depth reconstruction network according to the polarization phase gradient spectrum to obtain a preliminary depth map; performing ray tracing simulation processing according to the preliminary depth map to obtain an enhanced depth map; performing depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map; Step S3: Obtain optical system parameters; perform scene stratification on the optimized depth map to obtain a stratified scene; perform Fresnel diffraction simulation according to the optical system parameters and the stratified scene to obtain a Fresnel diffraction image set; perform focal plane prediction on the Fresnel diffraction image set to obtain a focal plane confidence map; Step S4: calculating the lens displacement according to the focal plane confidence map to obtain the lens displacement; correcting the lens displacement according to the lens displacement, and generating a lens control instruction to obtain a corrected lens displacement and a lens control instruction; Step S5: moving the lens according to the lens control instruction, and determining the focus state to obtain the focus state; performing focus iteration optimization according to the focus state to obtain the optimized lens control instruction.

[0031] In the embodiment of the present invention, reference Figure 1 FIG. 1 is a schematic diagram of the process flow of the automatic focusing method of an optical lens according to the present invention. In this example, the automatic focusing method of an optical lens includes the following steps: Step S1: collecting polarized light fields through a light field camera to obtain a polarized multi-view image set; extracting light field features from the polarized multi-view image set to obtain a polarized phase gradient spectrum; In an embodiment of the present invention, the core is to collect and extract light field information and integrate polarization characteristics. First, a light field camera and a rotatable linear polarizer are combined to collect multi-view images in different polarization directions to form a polarization multi-view image set. Then, each viewpoint image is Fourier transformed to obtain a Fourier transform spectrum. Next, the phase gradient of the Fourier transform spectrum is calculated to obtain a phase gradient map, and the polarization-related phase gradient is further calculated to obtain a polarization phase gradient map. Finally, the polarization phase gradient map is subjected to multi-dimensional light field feature vector generation to obtain a polarization phase gradient spectrum, which contains the polarization-related phase gradient information of the light field in the Fourier domain as input for subsequent depth reconstruction.

[0032] Step S2: constructing a polarization-aware depth reconstruction network according to the polarization phase gradient spectrum to obtain a preliminary depth map; performing ray tracing simulation processing according to the preliminary depth map to obtain an enhanced depth map; performing depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map; In an embodiment of the present invention, a deep reconstruction network is mainly constructed, and ray tracing is used to enhance training data. First, the polarization phase gradient spectrum is preprocessed, including dimensionality adjustment, normalization, and data type conversion. Then, a polarization-aware deep reconstruction network is constructed, which includes an encoder, a polarization-aware module, a decoder, and a loss function. The encoder extracts multi-scale features, the polarization-aware module processes polarization information, and the decoder maps the features to a depth map. Then, using the preliminary depth map and the preprocessed polarization phase gradient spectrum, ray tracing simulation is performed to generate simulated images and depth maps for training data enhancement. Finally, the preliminary depth map and the enhanced depth map are fused to obtain an optimized depth map.

[0033] Step S3: Obtain optical system parameters; perform scene stratification on the optimized depth map to obtain a stratified scene; perform Fresnel diffraction simulation according to the optical system parameters and the stratified scene to obtain a Fresnel diffraction image set; perform focal plane prediction on the Fresnel diffraction image set to obtain a focal plane confidence map; In an embodiment of the present invention, scene layering and Fresnel diffraction simulation are performed mainly based on the optimized depth map, and the focal plane confidence is calculated. First, according to the optimized depth map, the scene is divided into multiple depth layers to obtain a layered scene. Then, the optical system parameters are obtained, and Fresnel diffraction simulation is performed according to the layered scene and the optical system parameters to generate a Fresnel diffraction image set. Next, the diffraction characteristics of the Fresnel diffraction image set are extracted, and the light field focusing evaluation of the layered scene is performed according to the polarization phase gradient spectrum to obtain a light field clarity map. Finally, the diffraction feature map and the light field clarity map are fused to calculate the focal plane confidence of each depth layer to obtain a focal plane confidence map.

[0034] Step S4: calculating the lens displacement according to the focal plane confidence map to obtain the lens displacement; correcting the lens displacement according to the lens displacement, and generating a lens control instruction to obtain a corrected lens displacement and a lens control instruction; In the embodiment of the present invention, the lens control instruction is calculated and generated mainly based on the focal plane confidence map. First, the pixel with the highest confidence in the focal plane confidence map is found to obtain the optimal focal plane depth. Then, the lens displacement is calculated according to the optimal focal plane depth and the camera parameters. Next, the lens displacement is corrected according to the lens displacement and the optical system parameters to obtain the corrected lens displacement. Finally, according to the corrected lens displacement, a lens control instruction is generated to drive the lens motor to move to the target position.

[0035] Step S5: moving the lens according to the lens control instruction, and determining the focus state to obtain the focus state; performing focus iteration optimization according to the focus state to obtain the optimized lens control instruction; In the embodiment of the present invention, the focus state judgment and iterative optimization are mainly performed. First, according to the lens control instruction, the lens motor is driven to move and the current image is captured. Then, the image clarity is evaluated for the current image to obtain an image clarity value. Next, the image clarity value is compared with a preset clarity threshold to determine the focus state. If the focus fails, the focus iterative optimization is performed according to the focus state, the image clarity value, the focal plane confidence map and the corrected lens displacement to generate a new lens control instruction. The above process is repeated until the focus is successful or the preset number of iterations is reached.

[0036] Preferably, step S1 comprises the following steps: Step S11: collecting polarized light fields through a light field camera to obtain a polarized multi-view image set; Step S12: performing viewpoint image Fourier transform on the polarization multi-view image set to obtain a Fourier transform spectrum; Step S13: performing phase gradient calculation on the Fourier transform spectrum to obtain a phase gradient map; Step S14: performing polarization phase gradient calculation on the phase gradient image to obtain a polarization phase gradient image; Step S15: generating a multi-dimensional light field feature vector for the polarization phase gradient image to obtain a polarization phase gradient spectrum.

[0037] In an embodiment of the present invention, a light field camera with a microlens array (MLA) is configured and integrated with a rotatable linear polarizer. The MLA is located in front of the image sensor and is used to sample light in space and obtain light information at different angles. The linear polarizer is placed in front of the MLA and is used to control the polarization direction of the light entering the camera. The rotation angle of the polarizer is set, for example, set to 0°, 45°, 90° and 135° in sequence. At each polarization angle, the camera is controlled to collect light field data. At each polarization angle, the light field camera collects multi-view images through the MLA. The MLA decomposes the incident light into multiple sub-beams, each of which corresponds to a viewpoint. The image sensor records the intensity and color information of each sub-beam to generate a multi-view image. Repeat the above process to collect multi-view images at all set polarization angles. The multi-view images collected at different polarization angles are combined to form a polarization multi-view image set. The polarization multi-view image set includes multi-view images under multiple polarization directions, and each view corresponds to a scene under a polarization direction.

[0038] A multi-view image of one polarization direction is extracted from the polarization multi-view image set. A viewpoint image is selected from the multi-view image, which represents the scene information observed from a specific viewing angle. A two-dimensional fast Fourier transform (FFT) is performed on the viewpoint image. FFT is applied to each pixel of the viewpoint image to calculate the discrete Fourier transform of each pixel. FFT converts the image from the spatial domain to the frequency domain to obtain a complex matrix, the real and imaginary parts of which represent the frequency components of the image respectively. The above FFT operation is repeated for all viewpoint images in the polarization multi-view image set. The Fourier transform results of all viewpoint images are combined into a Fourier transform spectrum. The Fourier transform spectrum consists of multiple complex matrices, each matrix corresponds to the Fourier transform result of a viewpoint image, and each matrix is ​​associated with a specific polarization direction.

[0039] Select a complex matrix from the Fourier transform spectrum, which corresponds to the Fourier transform result of a viewpoint image. Calculate the phase of the complex matrix. The phase represents the relative position of different frequency components in the image. For each element in the complex matrix, use the inverse tangent function (arctan) to calculate its phase. Calculate the gradient of the phase in the horizontal and vertical directions. Use the Sobel operator or Scharr operator to convolve the phase and calculate the difference of the phase in the horizontal and vertical directions. The Sobel operator or Scharr operator is a set of predefined convolution kernels used to detect the edges of the image. Combine the calculated horizontal and vertical phase gradients to form a phase gradient vector. Calculate the modulus of the phase gradient vector to obtain the magnitude of the phase gradient. Repeat the above operation for all complex matrices in the Fourier transform spectrum. Group the phase gradient sizes of all viewpoint images into a phase gradient map. The phase gradient map contains the phase gradient information of all viewpoint images, and each pixel value represents the magnitude of the phase gradient at the pixel position.

[0040] Select two phase gradient images from the phase gradient image, which correspond to the same viewpoint position but have different polarization directions. Calculate the difference between the two phase gradient images. You can calculate the difference between corresponding pixels, or calculate the absolute value of the difference between corresponding pixels. Use the calculated difference as the new phase gradient value. Repeat the above operation for all phase gradient images with different polarization directions but the same viewpoint position in the phase gradient image. Combine the calculated polarization phase gradient values ​​to form a polarization phase gradient image. The polarization phase gradient image contains the polarization phase gradient information of all viewpoint images, and each pixel value represents the polarization phase gradient value of the pixel position.

[0041] Select a polarization phase gradient map from the polarization phase gradient map. Reshape the pixel values ​​of the polarization phase gradient map and convert it into a one-dimensional vector. For example, the pixel values ​​of the polarization phase gradient map can be arranged into a vector in row-first or column-first order. Repeat this operation for all polarization phase gradient maps and convert each polarization phase gradient map into a vector. Combine all vectors together to form a multidimensional feature vector. This multidimensional feature vector represents the comprehensive characteristics of the light field and contains phase gradient information from different viewpoints and different polarization directions. Use this multidimensional feature vector as a polarization phase gradient spectrum. The polarization phase gradient spectrum is a multidimensional feature vector used for subsequent depth reconstruction and focal plane prediction.

[0042] Preferably, step S11 includes the following steps: Step S111: placing an angle-calibrated linear polarizer in front of the light field camera, performing synchronous control of the light field camera and the polarizer, and obtaining synchronous control parameters; Step S112: collecting light field data in multiple polarization directions according to the synchronization control parameters to obtain an original light field image set; Step S113: correcting the effect of the polarizer on the light field on the original light field image set to obtain a corrected light field image set; Step S114: storing and marking the corrected light field image set to obtain a polarization multi-view image set.

[0043] In an embodiment of the present invention, a light field camera is prepared, which is equipped with a microlens array (MLA) for capturing light field information and a rotatable linear polarizer. The linear polarizer is placed in front of the light field camera lens to ensure that it can cover the entire field of view. The angle of the linear polarizer is accurately calibrated using a precision angle calibration device, such as an optical platform and an angle meter. The purpose of the calibration is to ensure that the polarization direction of the linear polarizer is consistent with the set angle. A synchronous control system is designed, which can simultaneously control the image acquisition of the light field camera and the rotation of the linear polarizer. The system includes a microcontroller for sending a control signal. The microcontroller is connected to the light field camera and the polarizer rotation mechanism. The rotation angle of the polarizer is set, for example, 0°, 45°, 90° and 135°. The microcontroller sends a control signal to the polarizer rotation mechanism to rotate it to the set angle. At the same time, the microcontroller sends an instruction to the light field camera to collect images. The synchronous control parameters include: a sequence of polarizer rotation angles, an image acquisition trigger signal at each angle, and an exposure time and gain setting for the collected image. These parameters are stored in the memory of the microcontroller for subsequent data acquisition.

[0044] Start the synchronous control system and collect light field data in multiple polarization directions according to the preset synchronous control parameters. The microcontroller first controls the polarizer to rotate to the first set angle, for example, 0°. After the polarizer is rotated into place, the microcontroller sends an image acquisition trigger signal to the light field camera. The light field camera collects a multi-view image according to the set exposure time and gain settings. The collected multi-view image is stored in the internal memory of the light field camera, or transmitted to an external storage device via a data line. Repeat the above process, control the polarizer to rotate to the next set angle, for example, 45°, and collect multi-view images. Repeat this process in sequence until the multi-view image collection of all preset polarization directions is completed. All collected multi-view images are combined into an original light field image set. The original light field image set contains multi-view images under multiple polarization directions, and each multi-view image corresponds to a specific polarization angle.

[0045] Since the linear polarizer is not an ideal polarization device, its transmittance to light of different polarization directions will be different, and additional light loss will be introduced. The original light field image set is corrected to eliminate these effects. Prepare a uniform, polarization-independent light source, such as an integrating sphere light source. Place the light source in front of the light field camera. At each preset polarization angle, collect images of the integrating sphere light source. Collect multiple integrating sphere images to calculate the transmittance and light loss of the polarizer. For each polarization angle, calculate a correction coefficient. The correction coefficient can be calculated by comparing the average pixel values ​​of the images at different polarization angles. For each multi-view image in the original light field image set, apply the corresponding correction coefficient according to its corresponding polarization angle. For each pixel in the multi-view image, use the correction coefficient for correction. For example, the light intensity can be corrected by multiplying the pixel value by the correction coefficient. Combine the corrected multi-view images into a corrected light field image set. The corrected light field image set eliminates the influence of the polarizer on the light field and improves the accuracy of subsequent processing.

[0046] Data is stored and labeled for each multi-view image in the corrected light field image set. A suitable file format, such as TIFF, PNG, or RAW, is selected for storing the multi-view images. The file name of each multi-view image is named, and the file name contains the polarization angle information corresponding to the image. For example, "image_000_0.tif" can be used to represent an image with a polarization angle of 0°, and "image_000_45.tif" can be used to represent an image with a polarization angle of 45°. A data index file is established, which records the file name, polarization angle, and other related information of each multi-view image, such as acquisition time, exposure time, and gain setting. All multi-view images in the corrected light field image set are stored in a storage device, and the data index file is also stored in the storage device. The stored corrected light field image set and the data index file together constitute a polarization multi-view image set. The polarization multi-view image set contains the corrected multi-view images, as well as metadata and related information used to describe these images.

[0047] Preferably, step S2 comprises the following steps: Step S21: performing polarization phase gradient spectrum preprocessing on the polarization phase gradient spectrum to obtain a preprocessed polarization phase gradient spectrum; Step S22: constructing a polarization-aware depth reconstruction network according to the preprocessed polarization phase gradient spectrum to obtain a preliminary depth map; Step S23: performing ray tracing simulation according to the preliminary depth map and the preprocessed polarization phase gradient spectrum, and performing training data enhancement to obtain an enhanced depth map; Step S24: performing depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map.

[0048] As an example of the present invention, refer to Figure 2 As shown, in this example, step S2 includes: Step S21: performing polarization phase gradient spectrum preprocessing on the polarization phase gradient spectrum to obtain a preprocessed polarization phase gradient spectrum; In an embodiment of the present invention, starting from the polarization phase gradient spectrum (PPGS) obtained in step S1, a preprocessing operation is performed to improve its performance for deep reconstruction network. First, the dimension of PPGS is adjusted. Since PPGS has different dimensions and data types, it needs to be converted into the input format of the deep reconstruction network. If PPGS is a multidimensional feature vector, it is reshaped into a four-dimensional tensor with dimensions (batch_size, height, width, channels). Wherein batch_size represents the batch size, height and width represent the height and width of the feature map, respectively, and channels represent the number of channels of the feature map. Secondly, data normalization is performed. Each element in PPGS is normalized so that its value range is between 0 and 1. Normalization can use the Min-Max normalization method to calculate the minimum and maximum values ​​of each feature, and then scale all elements using the formula: (x - min) / (max - min). If the value range of PPGS is already close to 0 to 1, this step can be omitted. Third, data type conversion is performed. The data type of PPGS is converted to a data type supported by the deep reconstruction network, for example, float32. Data type conversion can improve computational efficiency and numerical stability. Fourth, perform data augmentation if necessary. For example, the PPGS can be randomly cropped, rotated, or noise added to increase data diversity and improve the generalization ability of the model. The PPGS after the above processing is used as a preprocessed polarization phase gradient spectrum (PPPGS) and is ready to be input into the deep reconstruction network.

[0049] Step S22: constructing a polarization-aware depth reconstruction network according to the preprocessed polarization phase gradient spectrum to obtain a preliminary depth map; In an embodiment of the present invention, a polarization-aware deep reconstruction network based on a deep convolutional neural network (CNN) is constructed. The network receives PPPGS as input and outputs a preliminary depth map. The network structure includes the following main components. First, an encoder is designed to extract multi-scale features from PPPGS. The encoder includes multiple convolutional layers, pooling layers, and activation functions. The convolutional layer is used to extract local features of PPPGS, the pooling layer is used to reduce the dimension of the feature map, and the ReLU activation function is used to introduce nonlinearity. The encoder uses residual connections to alleviate the gradient vanishing problem. Secondly, a polarization-aware module is constructed. This module is embedded in certain layers of the encoder and is used to process polarization information. The polarization-aware module can be a network branch including multiple convolutional layers and an attention mechanism. The module receives polarization-related features in PPPGS as input and uses a cross-channel attention mechanism to allow the network to learn the association between features in different polarization directions. For example, the correlation between different polarization channels can be calculated, and the features of different channels can be weighted using attention weights. The output of the polarization-aware module is fused with the features of the backbone encoder, for example, by splicing or adding. Third, a decoder is constructed to map the features extracted by the encoder to a depth map. The decoder contains multiple deconvolution layers, upsampling layers, and skip connections. Deconvolution layers are used to increase the spatial resolution of feature maps, upsampling layers are used to enlarge feature maps to the same size as the input image, and skip connections are used to pass low-level features in the encoder to the decoder to preserve detail information. The last layer of the decoder outputs a single-channel depth map, which represents the depth value of each pixel in the scene. The mean square error (MSE) is used as the loss function to measure the difference between the predicted depth map and the true depth map. The Adam optimizer is used to train the network and optimize the network parameters. The trained network is applied to PPPGS to obtain a preliminary depth map.

[0050] Step S23: performing ray tracing simulation according to the preliminary depth map and the preprocessed polarization phase gradient spectrum, and performing training data enhancement to obtain an enhanced depth map; In an embodiment of the present invention, ray tracing technology is used to enhance training data in combination with a preliminary depth map and PPPGS. First, a ray tracing engine is selected, such as Mitsuba or POV-Ray. A geometric model similar to the actual scene is prepared, including objects, materials, and lighting conditions in the scene. The depth information of the scene is constructed using the preliminary depth map. The preliminary depth map is used as input to the ray tracing engine to simulate the propagation of light in the scene. PPPGS is used to provide polarization information for ray tracing simulation. For example, the reflection and refraction of light at different polarization angles can be simulated based on the polarization information in PPPGS. The ray tracing engine is run to generate simulated images and corresponding depth maps. The simulated images contain images from different viewpoints, which can simulate changes in the camera's viewing angle. The simulated depth map contains the true depth information of each pixel in the scene. The training data is enhanced using the simulated images and depth maps. For example, different lighting conditions can be simulated, including changes in the intensity, color, and position of the light source. The movement and rotation of objects in the scene can be simulated. Changes in camera parameters, such as focal length and aperture, can be simulated. The enhanced data is mixed with the original training data for retraining the deep reconstruction network. When retraining the deep reconstruction network, the parameters of the network can be adjusted to better adapt it to the new training data. After training with the augmented data, the augmented depth map is used.

[0051] Step S24: performing depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map; In an embodiment of the present invention, the preliminary depth map is fused with the enhanced depth map to improve the quality of the depth map. First, the preliminary depth map and the enhanced depth map are registered. The purpose of registration is to align the two depth maps to ensure that the same pixel position corresponds to the same scene point. Registration can use an image registration algorithm, such as feature point-based registration or image block-based registration. After registration, the preliminary depth map and the enhanced depth map are fused. The fusion method can use weighted averaging. Different weights are assigned to the preliminary depth map and the enhanced depth map. The selection of weights can be adjusted according to the quality of the depth map. For example, if the quality of the enhanced depth map is higher, a larger weight can be assigned to it. The weight can be determined according to the confidence of the preliminary depth map and the enhanced depth map. For example, if the confidence of the preliminary depth map is higher, a larger weight can be assigned to it. The weighted depth maps are fused to obtain an optimized depth map. The optimized depth map contains the depth information of each pixel in the scene and has higher accuracy and robustness.

[0052] Preferably, step S22 includes the following steps: Step S221: performing initial convolution layer feature extraction on the preprocessed polarization phase gradient spectrum to obtain a preliminary feature map; performing downsampling layer feature extraction based on the preliminary feature map to obtain a downsampled feature map; Step S222: performing polarization channel separation on the downsampled feature map to obtain separated polarization features; Step S223: introducing channel attention to the separated polarization features to obtain polarization features with channel attention; Step S224: performing polarization feature fusion on the polarization features with channel attention to obtain a fused feature map; Step S225: performing feature extraction based on the residual block encoder on the fused feature map to obtain encoder output features; Step S226: Perform decoder-based feature mapping on the encoder output features to obtain decoder output features; Step S227: Output the decoder output features as a depth map to obtain a preliminary depth map.

[0053] In an embodiment of the present invention, the preprocessed polarization phase gradient spectrum (PPPGS) obtained in step S21 is received and used as the input of the deep reconstruction network. An initial convolution layer is designed, which is composed of multiple convolution kernels, and the size of the convolution kernel is set to 3x3, for example. PPPGS is input to the initial convolution layer and a convolution operation is performed. The convolution operation uses the ReLU activation function to introduce nonlinearity. The ReLU activation function sets values ​​less than 0 to 0, and values ​​greater than 0 remain unchanged. The output result of the initial convolution layer is a preliminary feature map (Preliminary Feature Map). The preliminary feature map contains preliminary features extracted from PPPGS. The number of channels of the preliminary feature map depends on the number of convolution kernels in the initial convolution layer. According to the preliminary feature map, multiple downsampling layers are designed to extract higher-level features and reduce the size of the feature map. The downsampling layer includes a convolution layer and a pooling layer. The convolution kernel size of the convolution layer is set to 3x3, for example, and the ReLU activation function is used. The pooling layer uses maximum pooling (maxpooling), and the pooling window size is set to 2x2, for example, with a step size of 2. The maximum pooling operation selects the largest value from the pooling window. After the downsampling layer, the size of the feature map is reduced, but the feature abstraction is higher. The feature map processed by multiple downsampling layers is used as the downsampled feature map. The number of channels of the downsampled feature map depends on the number of convolution kernels in the downsampling layer.

[0054] The preprocessed polarization phase gradient spectrum (PPPGS) obtained in step S21 is received and used as the input of the deep reconstruction network. An initial convolution layer is designed, which is composed of multiple convolution kernels, and the size of the convolution kernel is set to 3x3, for example. The PPPGS is input to the initial convolution layer for convolution operation. The convolution operation uses the ReLU activation function to introduce nonlinearity. The ReLU activation function sets the values ​​less than 0 to 0, and the values ​​greater than 0 remain unchanged. The output result of the initial convolution layer is a preliminary feature map (Preliminary Feature Map). The preliminary feature map contains preliminary features extracted from PPPGS. The number of channels of the preliminary feature map depends on the number of convolution kernels in the initial convolution layer. According to the preliminary feature map, multiple downsampling layers are designed to extract higher-level features and reduce the size of the feature map. The downsampling layer includes a convolution layer and a pooling layer. The convolution kernel size of the convolution layer is set to 3x3, for example, and the ReLU activation function is used. The pooling layer uses max pooling, and the pooling window size is set to 2x2, for example, with a step size of 2. The maximum pooling operation selects the largest value from the pooling window. After the downsampling layer, the size of the feature map is reduced, but the feature abstraction is higher. The feature map processed by multiple downsampling layers is used as the downsampled feature map. The number of channels of the downsampled feature map depends on the number of convolution kernels in the downsampling layer.

[0055] The channel attention mechanism is introduced to the separated polarization features, so that the network can adaptively focus on important polarization feature channels. For each separated polarization feature, a channel attention module is constructed. The channel attention module contains a global average pooling layer, a fully connected layer, and a sigmoid activation function. The global average pooling layer averages the pixel values ​​of each channel of the feature map to obtain a value representing the importance of the channel. The fully connected layer is used to reduce and increase the dimension of the output of the global average pooling, for example, reducing the number of channels to 1 / 16 of the original number and then increasing it back to the original number of channels. The sigmoid activation function maps the output of the fully connected layer to between 0 and 1 to obtain the channel attention weight. The channel attention weight is multiplied by the corresponding polarization feature to obtain the polarization features with channel attention. For example, the features of each channel are multiplied by the corresponding attention weight to enhance the features of important channels and suppress the features of unimportant channels.

[0056] Fuse the polarization features with channel attention to integrate the information of different polarization directions. Concatenate all polarization features with channel attention. The concatenation operation connects the features of different polarization directions in the channel dimension to form a fused feature map. For example, if four polarization directions are used and the number of channels of the feature map of each polarization direction is C, the number of channels of the concatenated feature map is 4C. Perform a convolution operation on the concatenated feature map. The convolution kernel size of the convolution operation is set to 1x1, for example, to further extract features from the fused features. The convolution operation uses the ReLU activation function to introduce nonlinearity. The output of the convolution operation is used as the fused feature map.

[0057] The residual block is used as the core building block of the encoder to perform deeper feature extraction on the fused feature map. The design of the residual block can alleviate the gradient vanishing problem, allowing the network to be trained deeper. The residual block contains multiple convolutional layers, batch normalization layers, and ReLU activation functions. The kernel size of the convolutional layer is set to 3x3, for example. The batch normalization layer normalizes the feature map of each batch, speeding up the training process and improving the generalization ability of the model. The ReLU activation function introduces nonlinearity. The residual block also contains a skip connection that connects the input features directly to the output of the residual block. The skip connection enables the network to learn the residual, which is the difference between the input and output. Multiple residual blocks are designed and connected in series to form an encoder. The output of each residual block of the encoder is used as the input of the next residual block. The output of the last layer of the encoder is used as the encoder output feature. The encoder output feature contains higher-level features extracted from the fused feature map after multiple layers of processing.

[0058] Construct a decoder to map the encoder output feature to the depth map. The decoder contains multiple deconvolutional layers, upsampling layers, and skip connections. The deconvolutional layer is used to increase the spatial resolution of the feature map. The kernel size of the deconvolutional layer is set to 3x3 with a stride of 2. The upsampling layer is used to enlarge the feature map to the same size as the input image. The upsampling layer can use methods such as bilinear interpolation or transposed convolution. The skip connection passes the low-level features in the encoder to the decoder to preserve the detail information. The encoder output feature is input to the decoder. The decoder gradually restores the feature map to the same size as the input image through multiple layers of deconvolution, upsampling, and skip connections. The output of the last layer of the decoder is the decoder output feature. The decoder output feature contains the features mapped from the encoder output feature for generating the depth map.

[0059] Generate a preliminary depth map using the decoder output feature. Perform a 1x1 convolution on the decoder output feature. The output channel number of this convolution operation is 1, which is used to map the feature map to a single-channel depth map. The convolution operation does not use an activation function. The output result of the convolution operation is used as the preliminary depth map. Each pixel value in the preliminary depth map represents the depth value of the corresponding position in the scene. The depth value can be the distance from the object in the scene to the camera, or other ways to represent depth.

[0060] Preferably, step S3 comprises the following steps: Step S31: performing scene stratification on the optimized depth map to obtain a layered scene; Step S32: acquiring optical system parameters; performing Fresnel diffraction simulation according to the layered scene to obtain a Fresnel diffraction image set; Step S33: extracting diffraction features from the Fresnel diffraction image set to obtain a diffraction feature map; Step S34: performing light field focusing evaluation on the layered scene according to the polarization phase gradient spectrum to obtain a light field clarity map; Step S35: performing focal plane confidence calculation on the diffraction characteristic map and the light field clarity map to obtain a focal plane confidence map.

[0061] In an embodiment of the present invention, the optimized depth map obtained in step S2 is received, and the three-dimensional scene is decomposed into multiple two-dimensional depth layers based on the depth information of the depth map. First, the layering strategy is determined. A uniform layering strategy can be adopted, that is, layering is performed according to a fixed depth interval. For example, the depth value range is from 0 to 1000 units, and layering is performed according to each 10 units as a depth layer. An adaptive layering strategy can also be adopted, that is, layering is performed according to the rate of change of the depth value. For areas with drastic depth changes, a smaller depth interval can be used, and for areas with gentle depth changes, a larger depth interval can be used. After determining the layering strategy, the optimized depth map is traversed. For each pixel in the depth map, its corresponding depth value is obtained. According to the depth value and the layering strategy, the depth layer to which the pixel belongs is determined. Pixels with the same depth layer are divided into the same depth layer. A data structure of a layered scene is created. The data structure contains multiple depth layers, and each depth layer corresponds to a two-dimensional image. Pixels belonging to the same depth layer are copied to the corresponding two-dimensional image. Pixels that are not in any depth layer are processed. For example, its depth value can be set to a default value or ignored. All depth layers are combined to obtain a layered scene. The layered scene contains multiple two-dimensional images, each image corresponds to a depth layer, and each image contains the scene information in the depth layer.

[0062] Get the optical system parameters, including the focal length (f), pixel pitch (Δx, Δy) and aperture diameter (D) of the camera. These parameters can be obtained from the camera specifications or calibration process. Secondly, simulate the Fresnel diffraction of light waves at different depth layers according to the layered scene. For each depth layer in the layered scene, perform the following operations. Construct the input light field. For each depth layer, calculate the complex amplitude of each pixel in the depth layer according to the pixel value of the depth map. The amplitude part of the complex amplitude represents the intensity of the light, and the phase part represents the phase of the light. For example, it can be assumed that the amplitude of the light is proportional to the pixel value, and the phase is related to the propagation distance of the light. For each depth layer, generate a two-dimensional complex amplitude distribution as the input light field. Fill the input light field. Since the Fresnel diffraction calculation needs to process the entire space, the input light field needs to be filled to avoid edge effects. Methods such as zero-padding or mirror-padding can be used. The size of the filled light field is larger than the original light field. Perform a light field Fourier transform on the filled light field. Use the fast Fourier transform (FFT) algorithm to transform the filled light field from the spatial domain to the frequency domain. Obtain a frequency domain light field, which represents the frequency components of the light wave. Calculate the propagation function based on the optical system parameters and the frequency domain light field. The propagation function describes the changes of the light wave during propagation. For Fresnel diffraction, the propagation function can be calculated based on the Fresnel diffraction integral formula. The propagation function is related to the focal length, pixel spacing and wavelength. Calculate the propagation function based on these parameters. Perform frequency domain multiplication on the propagation function and the frequency domain light field. Multiply the propagation function and the frequency domain light field element by element. The multiplication result represents the frequency domain information of the light wave after propagation. Perform an inverse Fourier transform on the propagated frequency domain light field. Use the inverse fast Fourier transform (IFFT) algorithm to transform the propagated frequency domain light field from the frequency domain to the spatial domain. Obtain a diffraction image. The diffraction image represents the light intensity distribution of the light wave after propagation to the focal plane. Perform Fresnel diffraction simulation for each depth layer according to the diffraction image and the layered scene. Repeat the above process to perform Fresnel diffraction simulation for each depth layer in the layered scene to generate a Fresnel diffraction image. Combine the Fresnel diffraction images of all depth layers into a Fresnel Diffraction Image Set. The Fresnel Diffraction Image Set contains diffraction images of different depth layers for subsequent feature extraction and focal plane prediction.

[0063] Receive the Fresnel Diffraction Image Set obtained in step S32, and extract diffraction features from it to evaluate the clarity of the image. For each diffraction image in the Fresnel diffraction image set, perform the following operations. Select a suitable diffraction feature. The diffraction feature is used to describe the characteristics of the diffraction image and reflect the clarity of the image. A variety of diffraction features can be selected, such as image contrast, edge sharpness, and frequency domain features. Calculate the contrast of the diffraction image. There are many methods for calculating the image contrast, for example, calculating the local contrast or the global contrast. The local contrast can use indicators such as standard deviation or root mean square contrast. The global contrast can use the difference between the maximum pixel value and the minimum pixel value. Calculate the edge sharpness of the diffraction image. Edge sharpness reflects the clarity of the edge in the image. Edge detection operators such as the Sobel operator and the Canny operator can be used to calculate the strength of the edge. Calculate the frequency domain features of the diffraction image. Perform Fourier transform on the diffraction image to obtain a frequency domain representation. Extract the energy of the high-frequency component as the frequency domain feature. The high-frequency component reflects the detail information of the image. The calculated diffraction features are combined to form a diffraction feature vector. The diffraction feature vector contains different types of diffraction features, which are used to fully describe the characteristics of the diffraction image. Repeat the above operation for all diffraction images in the Fresnel diffraction image set. The diffraction feature vectors of all diffraction images are combined to form a diffraction feature map. The diffraction feature map contains the diffraction features of each depth layer, which are used for subsequent focal plane confidence calculations.

[0064] Using the Polarization Phase Gradient Spectrum (PPGS) and the Layered Scene, a light field focusing evaluation is performed on each depth layer to determine whether the depth layer is in focus. For each depth layer in the layered scene, the following operations are performed: According to the polarization phase gradient spectrum, the light field information of the depth layer is reconstructed. Using PPGS, combined with the depth information of the depth layer, a virtual perspective image of the depth layer can be reconstructed. A light field refocusing algorithm can be used, such as integral refocusing or parallax-based refocusing. The light field refocusing algorithm can generate images at different perspectives based on the depth information. The clarity of the refocused image is evaluated. The clarity evaluation is used to evaluate the clarity of the refocused image. A variety of clarity evaluation indicators can be used, such as image gradient, variance, or frequency domain energy. To calculate the gradient of the image, the Sobel operator or the Laplacian operator can be used. The variance of the image is calculated. The larger the variance, the higher the contrast of the image. The frequency domain energy of the image is calculated to extract the energy of the high-frequency component. The calculated clarity index is used as the light field focusing evaluation result of the depth layer. Repeat the above operation for all depth layers in the layered scene. Combine the light field focus evaluation results of all depth layers into a light field sharpness map (LFSM). The light field sharpness map contains the sharpness evaluation results of each depth layer and is used for subsequent focal plane confidence calculation.

[0065] The Diffraction Feature Map (DFM) and the Light Field Sharpness Map (LFSM) are fused to calculate the focal plane confidence of each depth layer. For each depth layer, the following operations are performed: obtain the diffraction feature value of the depth layer in the DFM, and obtain the sharpness value of the depth layer in the LFSM. Fuse the DFM and LFSM. A weighted fusion method can be used to weighted average the DFM and LFSM. Different weights are assigned to the DFM and LFSM, and the weights can be adjusted according to the characteristics of the scene. For example, in low-contrast scenes, the weight of the DFM can be increased to compensate for the lack of light field information. A neural network can also be used to learn the relationship between the DFM and the LFSM and predict the focal plane confidence. Normalize the calculated confidence. Normalize the confidence value to between 0 and 1 for subsequent processing. Combine the focal plane confidences of all depth layers to form a focal plane confidence map (FPCM). The focus plane confidence map is a two-dimensional map, each value of which represents the focus plane confidence at different depths of field. The focus plane confidence map is used to determine the optimal focus plane position.

[0066] Preferably, step S32 includes the following steps: Step S321: constructing an input light field according to the layered scene to obtain an input light field; Step S322: performing light field filling on the input light field to obtain a filled light field; Step S323: performing a light field Fourier transform on the filled light field to obtain a frequency domain light field; Step S324: Calculate the propagation function according to the optical system parameters and the frequency domain light field to obtain the propagation function; Step S325: performing frequency domain multiplication on the propagation function and the frequency domain light field to obtain the frequency domain light field after propagation; Step S326: performing an inverse Fourier transform on the propagated frequency domain light field to obtain a diffraction image; Step S327: performing Fresnel diffraction simulation for each depth layer according to the diffraction image and the layered scene to obtain a Fresnel diffraction image set.

[0067] In an embodiment of the present invention, a layered scene obtained in step S31 is received, and an input light field of each depth layer is constructed according to the depth information in the layered scene. For each depth layer in the layered scene, the following operations are performed: according to the depth value corresponding to the depth layer, the coordinates and the corresponding complex amplitude value of each pixel in the depth layer are obtained. The complex amplitude represents the amplitude and phase information of the light wave. The amplitude is usually related to the grayscale value or color value of the image, and the phase is related to the propagation distance of the light wave. A two-dimensional complex matrix is ​​constructed, and the size of the matrix is ​​the same as the image size of the depth layer. The complex amplitude value of each pixel is filled into the corresponding pixel position in the complex matrix. For example, if the grayscale value of the pixel is g and the depth value is z, the complex amplitude of the pixel can be defined as: A * exp(j * k * z), where A is the amplitude, j is the imaginary unit, and k is the wave number. A is set to be proportional to g, and k is set to a constant. The constructed complex matrix is ​​used as the input light field of the depth layer. For all depth layers in the layered scene, the above operation is repeated to obtain the input light field of each depth layer. The input light field represents the complex amplitude distribution of the light wave at this depth layer.

[0068] The input light field obtained in step S321 is received, and the input light field is padded to avoid edge effects in subsequent Fourier transform and diffraction simulation. Due to the periodicity of Fourier transform and the influence of boundary conditions in Fresnel diffraction simulation, the input light field needs to be padded. The filling method, for example, uses zero-padding. For each input light field, a new two-dimensional complex matrix is ​​created, and the size of the matrix is ​​larger than the size of the input light field. For example, if the size of the input light field is M x N, the size of the padded light field can be set to 2M x 2N. The input light field is copied to the center of the padded light field. The remaining pixel values ​​of the padded light field are set to 0. Mirror-padding can also be used. For each input light field, the edge pixels of the input light field are mirror-copied and filled into the padded light field. Mirror-padding can reduce edge effects and maintain the continuity of the image. The size of the padded light field is the same as the size of the input light field. Use zero padding or mirror padding to fill the input light field to obtain a padded light field (PaddingLight Field). The size of the filled light field is larger than that of the input light field and is used for subsequent Fourier transform and diffraction simulation.

[0069] Receive the filled light field obtained in step S322, and use Fourier transform to convert it from the spatial domain to the frequency domain. For each filled light field, use a two-dimensional fast Fourier transform (FFT) algorithm. The FFT algorithm can efficiently calculate the Fourier transform. The FFT algorithm converts each pixel of the filled light field into a frequency component in the frequency domain. The result of the Fourier transform is a complex matrix, and the real and imaginary parts of the matrix represent different frequency components in the frequency domain respectively. For each pixel of the filled light field, calculate its Fourier transform. Store the result of the Fourier transform in a new two-dimensional complex matrix. The size of the complex matrix is ​​the same as the size of the filled light field. Use the complex matrix as the frequency domain light field (FrequencyDomain Light Field). The frequency domain light field represents the distribution of light waves in the frequency domain.

[0070] According to the optical system parameters and the frequency domain light field, the propagation function is calculated. The propagation function describes the changes of light waves during propagation. Obtain the optical system parameters, including the focal length (focal length, f), pixel pitch (pixelpitch, Δx, Δy) and wavelength (wavelength, λ) of the camera. These parameters can be obtained from the camera specifications or calibration process. Calculate the propagation function. According to the Fresnel diffraction theory, the propagation function can be expressed as: H(fx, fy) = exp(-j * π * λ * z * (fx^2 + fy^2)), where fx and fy are frequency domain coordinates, and z is the propagation distance, that is, the depth value of the depth layer. Calculate the frequency domain coordinates fx and fy, fx = m / (M * Δx), fy = n / (N * Δy), where m and n are the indices of the frequency domain coordinates, and M and N are the sizes of the light field after filling. Calculate the propagation function according to the optical system parameters and the frequency domain coordinates. For each pixel (fx, fy) in the frequency domain, calculate the corresponding propagation function value. The propagation function is a complex number that represents the phase change and amplitude attenuation of the light wave during propagation. The calculated propagation function values ​​are stored in a two-dimensional complex matrix, the size of which is the same as the size of the frequency domain light field. The complex matrix is ​​used as the propagation function. The propagation function describes the changes of the light wave during propagation and is used for subsequent frequency domain multiplication.

[0071] Receive the frequency domain light field obtained in step S323 and the propagation function obtained in step S324, and multiply the propagation function and the frequency domain light field element by element. For each pixel in the frequency domain light field, multiply the corresponding propagation function value. For example, if the complex value of a pixel in the frequency domain light field is A, and the complex value of the corresponding pixel in the propagation function is H, the result of the multiplication is A * H. Repeat the above operation for all pixels in the frequency domain light field. Store the multiplication result in a new two-dimensional complex matrix, the size of which is the same as the size of the frequency domain light field and the propagation function. Use the complex matrix as the propagated frequency domain light field (Propagated Frequency Domain Light Field). The propagated frequency domain light field represents the frequency domain distribution of the light wave after propagation.

[0072] The frequency domain light field after propagation obtained in step S325 is received, and the inverse Fourier transform is used to convert it from the frequency domain to the spatial domain. For each frequency domain light field after propagation, a two-dimensional fast inverse Fourier transform (IFFT) algorithm is used. The IFFT algorithm can efficiently calculate the inverse Fourier transform. The IFFT algorithm converts each frequency component of the frequency domain light field after propagation to a pixel in the spatial domain. For each pixel of the frequency domain light field after propagation, its inverse Fourier transform is calculated. The result of the inverse Fourier transform is stored in a new two-dimensional complex matrix. The size of the complex matrix is ​​the same as the size of the frequency domain light field after propagation. The diffraction image is calculated, and the square of the modulus of the result of the inverse Fourier transform is taken. The square of the modulus represents the intensity of the light wave. The square of the modulus is used as the diffraction image (Diffraction Image). The diffraction image represents the intensity distribution of the light wave after propagation to the focal plane.

[0073] According to the diffraction image obtained in step S326 and the layered scene, the Fresnel diffraction simulation of each depth layer is performed to obtain a Fresnel diffraction image set. For each depth layer in the layered scene, the above steps S321 to S326 are repeated to simulate the Fresnel diffraction after the light wave propagates in the depth layer. For each depth layer, a diffraction image is obtained. The diffraction images of all depth layers are combined into a Fresnel diffraction image set (Fresnel Diffraction ImageSet). The Fresnel diffraction image set contains multiple diffraction images, each of which corresponds to a depth layer.

[0074] Preferably, step S4 comprises the following steps: Step S41: determining the optimal focal plane position on the focal plane confidence map to obtain the optimal focal plane depth; Step S42: obtaining camera parameters; performing lens displacement calculation according to the optimal focal plane depth and the camera parameters to obtain the lens displacement; Step S43: performing lens displacement correction according to the lens displacement amount and the optical system parameters to obtain a corrected lens displacement; Step S44: generating a lens control instruction for the corrected lens displacement to obtain a lens control instruction.

[0075] In an embodiment of the present invention, the focal plane confidence map (FPCM) obtained in step S3 is received to determine the optimal focal plane position. FPCM is a two-dimensional image, each pixel of which represents the focus confidence at different depths of field. First, find the pixel with the highest confidence in FPCM. Traverse all pixels in FPCM to find the pixel with the highest value. Record the row index and column index of the pixel. According to the row index and column index of the pixel, and the depth range corresponding to FPCM, calculate the optimal focal plane depth. For example, if the depth range of FPCM is from 0 to 1000 units, and the row index of FPCM corresponds to the depth value, linear interpolation or a lookup table can be used to determine the optimal focal plane depth. The optimal focal plane depth corresponds to the depth value with the highest confidence in FPCM. The depth value is used as the optimal focal plane depth (OFPD). OFPD represents the clearest focal plane position in the scene.

[0076] Get the relevant parameters of the camera and calculate the distance the lens needs to move based on the optimal focal plane depth (OFPD). First, get the focal length (focal length, f) of the camera. The focal length can be obtained from the camera specifications or calibration process. Get the initial lens position of the camera. The initial lens position of the camera indicates the current distance from the lens to the image sensor. Get the sensor size of the camera. The sensor size is used to calculate the image distance. Calculate the distance the lens needs to move based on the OFPD and camera parameters. Use the thin lens imaging formula: 1 / f = 1 / u + 1 / v, where f is the focal length, u is the object distance (OFPD), and v is the image distance. Calculate the image distance v based on the OFPD and focal length. Lens displacement = v - initial image distance. The lens displacement indicates the distance and direction the lens needs to move. If the lens displacement is positive, it means that the lens needs to move away from the sensor. If the lens displacement is negative, it means that the lens needs to move closer to the sensor. Use the calculated lens displacement (LensDisplacement, LD) as the instruction for lens movement.

[0077] Correct the lens displacement according to the lens displacement (LD) and optical system parameters to improve the focusing accuracy. Obtain the optical system parameters, including the aperture value (f / #), zoom information, and other optical element parameters. The aperture value affects the depth of field, and the zoom information affects the focal length. Correct the lens displacement according to the aperture value. If a smaller depth of field (large aperture) is required and more precise focusing is required, the LD needs to be fine-tuned. For example, a small offset can be made to the LD according to the aperture value. The larger the aperture value, the larger the offset. Correct the lens displacement according to the focal length. If it is a zoom lens, the lens movement needs to be adjusted according to the current focal length. For example, the LD can be scaled as a function of the focal length. Correct the lens displacement according to other optical element parameters. Consider other optical system parameters, such as lens distortion, chromatic aberration, etc. Correct the LD according to these parameters. Use the corrected lens displacement (CLD) as the final lens movement instruction.

[0078] According to the Corrected Lens Displacement (CLD), a lens control command is generated to drive the lens motor to move to the target position. Convert CLD to a control signal that the lens motor can understand. The lens motor usually uses a stepper motor or a DC motor. Convert CLD to the number of steps of the stepper motor. The number of steps of the stepper motor is proportional to the distance the lens moves. The CLD needs to be converted to the corresponding number of steps according to the stepping accuracy of the lens motor. Convert CLD to a voltage signal. If a DC motor is used, CLD needs to be converted to a corresponding voltage signal to control the speed and direction of the motor. Send instructions through the lens control interface. The lens control interface usually uses serial communication protocols such as I2C and UART. Send the control signal to the lens motor through the lens control interface. The control signal includes: the direction of lens movement (forward or reverse), the number of steps or voltage value. Generate a lens control command (LCC). LCC contains information such as the direction of lens movement and the distance (number of steps or voltage value) of movement, which is used to control the lens motor to move to the target position.

[0079] Preferably, step S5 comprises the following steps: Step S51: moving the lens according to the lens control instruction and acquiring the current image to obtain the current image; Step S52: Evaluate the image clarity of the current image to obtain an image clarity value; Step S53: using a preset clarity threshold to determine the focus state of the image clarity value to obtain the focus state; Step S54: performing focus iterative optimization according to the focus state, the image clarity value, the focal plane confidence map and the corrected lens displacement to obtain an optimized lens control instruction.

[0080] In an embodiment of the present invention, a lens control command (Lens Control Command, LCC) generated in step S4 is received to control the lens motor to move to a target position. The lens motor is controlled according to the LCC. The LCC contains information about the direction and distance (number of steps or voltage value) of the lens movement. The LCC is sent to the lens motor driver through a lens control interface, such as I2C or UART. The lens motor driver drives the lens motor to move according to the LCC. Wait for the lens motor to move to the target position. After the lens motor has moved, the current image is captured. Use a light field camera to capture the current image. The exposure time and gain settings when capturing the image can refer to the previous settings, or be adjusted according to the scene brightness. The captured image is used as the current image (Current Image, CI). CI represents the scene image at the current lens position.

[0081] Receive the current image (CI) obtained in step S51, and perform image clarity evaluation on it to determine the clarity of the image. Select a suitable clarity evaluation index. A variety of clarity evaluation indicators can be used, such as image gradient, frequency domain energy, or clarity evaluation based on a pre-trained neural network. Calculate the image gradient. The image gradient can reflect the edge and detail information of the image. The Sobel operator, Prewitt operator or Laplacian operator can be used to calculate the horizontal and vertical gradients of the image. Calculate the modulus or sum of squares of the gradient image to obtain the gradient value of the image. Calculate the frequency domain energy. Perform Fourier transform on the image to obtain a frequency domain representation. Extract the energy of the high-frequency component as the clarity index of the image. The high-frequency component reflects the detail information of the image. Use a pre-trained neural network. Use a pre-trained neural network, input CI, and output a clarity value. The neural network needs to be trained using a large number of clear images and blurred images. Calculate the clarity value of CI. Calculate the clarity value of CI based on the selected clarity evaluation index. The calculated sharpness value is used as the image sharpness value (ISV). ISV represents the clarity of CI.

[0082] Receive the image sharpness value (ISV) obtained in step S52, and compare it with the preset sharpness threshold to determine the focus state. Set a sharpness threshold (ST). ST represents the minimum requirement for image sharpness. ST can be adjusted according to the actual application scenario and the performance of the camera. Compare ISV with ST. If ISV is greater than or equal to ST, it is considered that the image has been focused and the focus state is "focus successful". If ISV is less than ST, it is considered that the image has not been focused and the focus state is "focus failed". Generate a focus state (FS). FS represents the status information of whether the focus is successful.

[0083] Focus iterative optimization is performed based on the focus status (FS), image sharpness value (ISV), focal plane confidence map (FPCM) and corrected lens displacement (CLD) to further improve the focusing accuracy. If FS indicates that the focus fails, iterative focus optimization is required. The focal plane offset is calculated based on the difference between the ISV and the expected value. The expected value can be set to the maximum ISV ​​or a threshold close to the maximum ISV. The difference between the ISV and the expected value is calculated. For example, the absolute value of the difference can be calculated. The search direction is determined based on the distribution of confidence in the FPCM. The search direction should point to a direction with higher confidence. Based on the focal plane offset and the search direction, the distance and direction of the lens movement are calculated. Based on the focal plane offset, the CLD is fine-tuned. The fine-tuned CLD indicates the distance and direction the lens needs to move. A new lens control command is generated. Based on the fine-tuned CLD, a new lens control command (RLCC) is generated. RLCC contains the direction and distance information of the lens movement, which is used to control the movement of the lens motor. Repeat steps S51 to S53, use RLCC to control the lens movement, collect images again, and perform clarity evaluation. Until FS indicates that the focus is successful, or the preset number of iterations is reached. If FS indicates that the focus is successful, the focus iterative optimization is completed. If the preset number of iterations is reached, but FS still indicates that the focus fails, the iterative optimization is stopped and the current lens position is returned.

[0084] Preferably, the present invention further provides an optical lens automatic focusing system, which is used to execute the optical lens automatic focusing method as described above, and the optical lens automatic focusing system comprises: The light field feature extraction module is used to collect polarized light fields through a light field camera to obtain a polarized multi-view image set; extract light field features from the polarized multi-view image set to obtain a polarization phase gradient spectrum; The depth information reconstruction module is used to construct a polarization-aware depth reconstruction network according to the polarization phase gradient spectrum to obtain a preliminary depth map; perform ray tracing simulation processing according to the preliminary depth map to obtain an enhanced depth map; perform depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map; The focal plane prediction module is used to obtain optical system parameters; perform scene stratification on the optimized depth map to obtain a stratified scene; perform Fresnel diffraction simulation based on the optical system parameters and the stratified scene to obtain a Fresnel diffraction image set; perform focal plane prediction on the Fresnel diffraction image set to obtain a focal plane confidence map; The lens control instruction generation module is used to calculate the lens displacement according to the focal plane confidence map to obtain the lens displacement; perform lens displacement correction according to the lens displacement, and generate the lens control instruction to obtain the corrected lens displacement and the lens control instruction; The focus confirmation and optimization module is used to move the lens according to the lens control instruction, and to judge the focus state to obtain the focus state; and to perform focus iterative optimization according to the focus state to obtain the optimized lens control instruction.

[0085] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is therefore intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.

[0086] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.

Claims

1. An optical lens automatic focusing method, characterized in that: The following steps are involved: Step S1: collecting polarized light fields through a light field camera to obtain a polarized multi-view image set; extracting light field features from the polarized multi-view image set to obtain a polarized phase gradient spectrum; Step S2: constructing a polarization-aware depth reconstruction network according to the polarization phase gradient spectrum to obtain a preliminary depth map; performing ray tracing simulation processing according to the preliminary depth map to obtain an enhanced depth map; performing depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map; Step S3: Obtain optical system parameters; perform scene stratification on the optimized depth map to obtain a layered scene; Performing Fresnel diffraction simulation according to optical system parameters and layered scenes to obtain a Fresnel diffraction image set; performing focal plane prediction on the Fresnel diffraction image set to obtain a focal plane confidence map; Step S4: calculating the lens displacement according to the focal plane confidence map to obtain the lens displacement; correcting the lens displacement according to the lens displacement, and generating a lens control instruction to obtain a corrected lens displacement and a lens control instruction; Step S5: moving the lens according to the lens control instruction, and determining the focus state to obtain the focus state; Focus iterative optimization is performed according to the focus state to obtain optimized lens control instructions.

2. The optical lens automatic focusing method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: collecting polarized light fields through a light field camera to obtain a polarized multi-view image set; Step S12: performing viewpoint image Fourier transform on the polarization multi-view image set to obtain a Fourier transform spectrum; Step S13: performing phase gradient calculation on the Fourier transform spectrum to obtain a phase gradient map; Step S14: performing polarization phase gradient calculation on the phase gradient image to obtain a polarization phase gradient image; Step S15: generating a multi-dimensional light field feature vector for the polarization phase gradient image to obtain a polarization phase gradient spectrum.

3. The optical lens automatic focusing method according to claim 2, characterized in that: Step S11 includes the following steps: Step S111: placing an angle-calibrated linear polarizer in front of the light field camera, performing synchronous control of the light field camera and the polarizer, and obtaining synchronous control parameters; Step S112: collecting light field data in multiple polarization directions according to the synchronization control parameters to obtain an original light field image set; Step S113: correcting the effect of the polarizer on the light field on the original light field image set to obtain a corrected light field image set; Step S114: storing and marking the corrected light field image set to obtain a polarization multi-view image set.

4. The optical lens automatic focusing method according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: performing polarization phase gradient spectrum preprocessing on the polarization phase gradient spectrum to obtain a preprocessed polarization phase gradient spectrum; Step S22: constructing a polarization-aware depth reconstruction network according to the preprocessed polarization phase gradient spectrum to obtain a preliminary depth map; Step S23: performing ray tracing simulation according to the preliminary depth map and the preprocessed polarization phase gradient spectrum, and performing training data enhancement to obtain an enhanced depth map; Step S24: performing depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map.

5. The optical lens automatic focusing method according to claim 4, characterized in that: Step S22 includes the following steps: Step S221: performing initial convolution layer feature extraction on the preprocessed polarization phase gradient spectrum to obtain a preliminary feature map; performing downsampling layer feature extraction based on the preliminary feature map to obtain a downsampled feature map; Step S222: performing polarization channel separation on the downsampled feature map to obtain separated polarization features; Step S223: introducing channel attention to the separated polarization features to obtain polarization features with channel attention; Step S224: performing polarization feature fusion on the polarization features with channel attention to obtain a fused feature map; Step S225: performing feature extraction based on the residual block encoder on the fused feature map to obtain encoder output features; Step S226: Perform decoder-based feature mapping on the encoder output features to obtain decoder output features; Step S227: Output the decoder output features as a depth map to obtain a preliminary depth map.

6. The optical lens automatic focusing method according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: performing scene stratification on the optimized depth map to obtain a layered scene; Step S32: acquiring optical system parameters; performing Fresnel diffraction simulation according to the layered scene to obtain a Fresnel diffraction image set; Step S33: extracting diffraction features from the Fresnel diffraction image set to obtain a diffraction feature map; Step S34: performing light field focusing evaluation on the layered scene according to the polarization phase gradient spectrum to obtain a light field clarity map; Step S35: performing focal plane confidence calculation on the diffraction characteristic map and the light field clarity map to obtain a focal plane confidence map.

7. The optical lens automatic focusing method according to claim 6, characterized in that: Step S32 includes the following steps: Step S321: constructing an input light field according to the layered scene to obtain an input light field; Step S322: performing light field filling on the input light field to obtain a filled light field; Step S323: performing a light field Fourier transform on the filled light field to obtain a frequency domain light field; Step S324: Calculate the propagation function according to the optical system parameters and the frequency domain light field to obtain the propagation function; Step S325: performing frequency domain multiplication on the propagation function and the frequency domain light field to obtain the frequency domain light field after propagation; Step S326: performing an inverse Fourier transform on the propagated frequency domain light field to obtain a diffraction image; Step S327: performing Fresnel diffraction simulation for each depth layer according to the diffraction image and the layered scene to obtain a Fresnel diffraction image set.

8. The optical lens automatic focusing method according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: determining the optimal focal plane position on the focal plane confidence map to obtain the optimal focal plane depth; Step S42: obtaining camera parameters; performing lens displacement calculation according to the optimal focal plane depth and the camera parameters to obtain the lens displacement; Step S43: performing lens displacement correction according to the lens displacement amount and the optical system parameters to obtain a corrected lens displacement; Step S44: generating a lens control instruction for the corrected lens displacement to obtain a lens control instruction.

9. The optical lens automatic focusing method according to claim 1, characterized in that: Step S5 includes the following steps: Step S51: moving the lens according to the lens control instruction and acquiring the current image to obtain the current image; Step S52: Evaluate the image clarity of the current image to obtain an image clarity value; Step S53: using a preset clarity threshold to determine the focus state of the image clarity value to obtain the focus state; Step S54: performing focus iterative optimization according to the focus state, the image clarity value, the focal plane confidence map and the corrected lens displacement to obtain an optimized lens control instruction.

10. An optical lens automatic focusing system, characterized in that: Used to perform the optical lens automatic focusing method as claimed in claim 1, the optical lens automatic focusing system comprises: The light field feature extraction module is used to collect polarized light fields through a light field camera to obtain a polarized multi-view image set; extract light field features from the polarized multi-view image set to obtain a polarization phase gradient spectrum; The depth information reconstruction module is used to construct a polarization-aware depth reconstruction network according to the polarization phase gradient spectrum to obtain a preliminary depth map; perform ray tracing simulation processing according to the preliminary depth map to obtain an enhanced depth map; perform depth map optimization on the preliminary depth map and the enhanced depth map to obtain an optimized depth map; The focal plane prediction module is used to obtain optical system parameters; perform scene stratification on the optimized depth map to obtain a stratified scene; perform Fresnel diffraction simulation based on the optical system parameters and the stratified scene to obtain a Fresnel diffraction image set; perform focal plane prediction on the Fresnel diffraction image set to obtain a focal plane confidence map; The lens control instruction generation module is used to calculate the lens displacement according to the focal plane confidence map to obtain the lens displacement; perform lens displacement correction according to the lens displacement, and generate the lens control instruction to obtain the corrected lens displacement and the lens control instruction; The focus confirmation and optimization module is used to move the lens according to the lens control instruction, and to judge the focus state to obtain the focus state; and to perform focus iterative optimization according to the focus state to obtain the optimized lens control instruction.

Citation Information

Patent Citations

  • Light field image processing method for depth acquisition

    CN111670576A

  • Multi-depth target focusing method based on computational ghost imaging

    CN112165570A

  • Automatic focusing method for focal plane of imaging ellipsometer

    CN112969026A

  • Electric control focusing full-field optical coherence tomography system and method thereof

    CN114111623A

  • Monocular snapshot type depth polarization four-dimensional imaging method and system

    CN114595636A