Optical lens autofocus system and method
The polarized light field is acquired through the light field camera and combined with the polarization phase gradient spectrum and Fresnel diffraction simulation, the automatic focus problem in complex light and low contrast scenes is solved, and high-precision and fast lens control are achieved to adapt to the shooting of complex environments and fast moving objects.
Patent Information
- Application Number
- CN202510482490.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing automatic focus technology is limited in complex lighting environments, low-contrast scenes, and when shooting quickly moving objects, focusing accuracy and speed are difficult to accurately adjust the focus.
Polarized light field is collected by the light field camera, depth reconstruction and Fresnel diffraction simulation are performed using the polarization phase gradient spectrum, combined with ray tracing and depth map optimization, lens displacement is calculated and control instructions are generated to achieve accurate movement and focus state judgment of the lens.
Significantly improve the accuracy and robustness of autofocus in complex lighting and low contrast scenarios, ensure image clarity, and adapt to the shooting needs of fast moving objects.
Smart Images

Figure CN120010086B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic focusing of optical lenses, and particularly to an automatic focusing system and method for optical lenses. Background Art
[0002] The goal of autofocus (AF) technology is to automatically adjust the focus of a camera lens to make the captured image clearer. The mainstream autofocus technologies are mainly divided into two categories:
[0003] Phase Detection Autofocus (PDAF): The principle is to set dedicated phase detection pixels on the image sensor (or use a separate phase detection sensor). These pixels are divided into several pairs, and each pair of pixels receives light from different parts of the lens. By comparing the phase difference between the images received by these paired pixels, it can be determined whether the focus is in front of or behind the target, as well as the direction and distance that the lens needs to move.
[0004] Contrast Detection Autofocus (CDAF): The principle is to determine the focus by analyzing the contrast of the image captured by the image sensor. When the image contrast is the highest, the image is considered to be the clearest, that is, at the focus position. The camera continuously moves the lens and compares the image contrast to find the lens position with the highest contrast.
[0005] However, in complex lighting environments, low-contrast scenes, and when shooting fast-moving objects, the focusing accuracy and speed of existing autofocus technologies are limited. In complex lighting environments, the image sensor is overexposed, the contrast is reduced, it is difficult for CDAF to find the contrast peak, and the phase detection pixels of PDAF are also saturated. In low-contrast scenes, there is a lack of texture and details, it is difficult for CDAF to find the contrast change, and it is also difficult for the phase detection pixels of PDAF to distinguish. Fast-moving objects make it difficult to predict the object's movement trajectory and impossible to predict the focus position. Summary of the Invention
[0006] Based on this, it is necessary to provide an automatic focusing system and method for optical lenses to solve at least one of the above technical problems.
[0007] To achieve the above object, an automatic focusing method for an optical lens includes the following steps:
[0008] Step S1: Collect a polarized light field through a light field camera to obtain a set of polarized multi-view images; extract light field features from the set of polarized multi-view images to obtain a polarized phase gradient spectrum;
[0009] Step S2: Construct a polarization-aware depth reconstruction network based on the polarization phase gradient spectrum to obtain a preliminary depth map; perform ray tracing simulation processing on the preliminary depth map to obtain an enhanced depth map; optimize the preliminary depth map and the enhanced depth map to obtain an optimized depth map.
[0010] Step S3: Obtain the optical system parameters; perform scene stratification on the optimized depth map to obtain a stratified scene; perform Fresnel diffraction simulation based on the optical system parameters and the stratified scene to obtain a set of Fresnel diffraction images; perform focal plane prediction on the set of Fresnel diffraction images to obtain a focal plane confidence map.
[0011] Step S4: Calculate the lens displacement according to the focal plane confidence map to obtain the lens displacement amount; correct the lens displacement according to the lens displacement amount and generate a lens control command to obtain the corrected lens displacement and the lens control command.
[0012] Step S5: Move the lens according to the lens control command and judge the focusing state to obtain the focusing state; perform iterative optimization of the focusing according to the focusing state to obtain an optimized lens control command.
[0013] The present invention combines a light field camera and a polarizer to achieve a comprehensive capture of the scene light field information, providing richer and more reliable information for subsequent depth reconstruction and focal plane prediction. Especially in complex lighting and low-contrast scenes, it can significantly improve the accuracy and robustness of autofocus. By constructing a polarization-aware depth reconstruction network and combining ray tracing simulation and data augmentation techniques, accurate reconstruction of depth information is achieved, and the generalization ability of the network is improved, laying a foundation for subsequent focal plane prediction. Through scene stratification and Fresnel diffraction simulation of the optimized depth map, the imaging effects on different depth layers can be simulated, and combined with the light field information, accurate prediction of the focal plane position is achieved, improving the accuracy of autofocus. By calculating and correcting the lens displacement and generating a lens control command, precise control of the lens motor is achieved, so that the lens is moved to the predicted optimal focal plane position to achieve fast autofocus. Through focusing state judgment and iterative optimization, gradual adjustment of the lens position is achieved, further improving the accuracy and efficiency of focusing, ensuring the clarity of the image, and thus completing the entire autofocus process. Therefore, the present invention provides an optical lens autofocus method, comprehensively utilizing polarization information, wave optics principles, light field information, and deep learning techniques to construct a new autofocus framework. This framework not only overcomes the limitations of traditional methods in complex lighting, low-contrast, and fast-moving object scenes, but also improves the focusing accuracy, speed, and robustness.
[0014] Preferably, step S1 includes the following steps:
[0015] Step S11: Perform polarized light field acquisition through a light field camera to obtain a polarized multi-view image set;
[0016] Step S12: Perform Fourier transform on the viewpoint images of the polarized multi-view image set to obtain a Fourier transform spectrum;
[0017] Step S13: Calculate the phase gradient of the Fourier transform spectrum to obtain a phase gradient map;
[0018] Step S14: Calculate the polarized phase gradient of the phase gradient map to obtain a polarized phase gradient map;
[0019] Step S15: Generate a multi-dimensional light field feature vector from the polarized phase gradient map to obtain a polarized phase gradient spectrum.
[0020] The present invention realizes a more comprehensive and detailed sampling of the scene light field by combining a light field camera and a linear polarizer. By collecting multi-view images with multiple polarization directions, the polarization characteristics of the light in the scene can be obtained, which provides richer information for subsequent depth reconstruction and focal plane prediction. Compared with only using a traditional camera, this polarization light field acquisition can better capture the depth and structure information of the scene. Especially in complex lighting environments and low-contrast scenes, it can more effectively extract scene features, thereby improving the accuracy and robustness of subsequent autofocus. By transforming the viewpoint image from the spatial domain to the frequency domain, the frequency components in the image can be better analyzed. The Fourier transform spectrum provides the energy distribution of the image at different frequencies, which is crucial for subsequent phase gradient calculation and feature extraction. Frequency domain analysis can highlight the texture and detail information in the image. Especially when dealing with low-contrast scenes, the frequency domain information can more effectively help extract the features of the image, thereby improving the accuracy of subsequent autofocus. By calculating the phase gradient of the Fourier transform spectrum, the edge and structure information of the image can be extracted. The phase gradient reflects the phase change between pixels in the image and can be used to detect the edge, texture, and depth information of the image. The phase gradient map can highlight the detail information in the image, reduce the influence of light changes, and improve the accuracy of subsequent depth reconstruction and focal plane prediction. By fusing the phase gradient information of different polarization directions, the polarization characteristics of the scene can be enhanced, thereby improving the extraction accuracy of depth and focus information. The polarization phase gradient map combines polarization information and phase gradient information and can better reflect the polarization state of the light in the scene, providing more reliable features for subsequent depth reconstruction and focal plane prediction. In complex lighting environments and scenes with polarization characteristics, the polarization phase gradient can more effectively extract the features of the scene, thereby improving the performance of autofocus. By converting the polarization phase gradient map into a multi-dimensional feature vector, the complex information of the light field can be compressed and encoded, facilitating subsequent depth reconstruction and focal plane prediction. The polarization phase gradient spectrum contains the polarization-related phase gradient information of the light field in the Fourier domain and can comprehensively describe the characteristics of the light field, providing rich input features for the subsequent depth reconstruction network. The form of the feature vector is convenient for the processing of deep learning networks, can improve the calculation efficiency and the generalization ability of the model, and thus improve the performance of the entire autofocus system.
[0021] Preferably, step S11 includes the following steps:
[0022] Step S111: Place the angle-calibrated linear polarizer in front of the light field camera, perform synchronous control of the light field camera and the polarizer, and obtain the synchronous control parameters;
[0023] Step S112: Collect light field data with multiple polarization directions according to the synchronous control parameters to obtain the original light field image set;
[0024] Step S113: Correct the influence of the polarizer on the original light field image set to obtain a corrected light field image set;
[0025] Step S114: Store and label the corrected light field image set to obtain a polarized multi-view image set.
[0026] In the present invention, by precisely calibrating the angle of the linear polarizer, it is ensured that the acquired light field data has accurate polarization direction information. Synchronously controlling the light field camera and the polarizer guarantees the temporal synchronization of the data acquired under different polarization directions, which is crucial for subsequent fusion of multi-polarization direction information. The setting of the synchronous control parameters enables the collaborative work of the light field camera and the polarizer to be carried out efficiently and accurately, laying a solid foundation for subsequent multi-polarized light field data acquisition, thereby improving the accuracy and robustness of autofocus. By acquiring multi-view images under different preset polarization directions, the light field information of the scene under different polarization states is obtained. The original light field image set contains rich polarization information, providing a data basis for subsequent polarization characteristic analysis and feature extraction. Acquiring multi-view images can obtain the stereoscopic information of the scene, further enhancing the subsequent depth reconstruction ability. The acquisition of such multi-polarization direction data enables subsequent autofocus methods to better adapt to complex lighting conditions and scenes with polarization characteristics, improving the accuracy and reliability of focusing. By correcting the influence of the polarizer, the non-uniformity brought by the polarizer to the light field data is eliminated, ensuring the accuracy and reliability of the data. The correction can effectively reduce the errors caused by the transmittance difference and light loss of the polarizer, thereby improving the accuracy of subsequent calculations and analyses. The corrected light field image set enables subsequent processing to more accurately reflect the true polarization characteristics of the scene, further improving the performance of autofocus. Through standardized data storage and labeling, subsequent data management and processing are facilitated. Data storage can ensure that the acquired light field data can be completely saved and used in subsequent steps. Data labeling, for example, including polarization angle information in the file name, enables convenient identification and invocation of data in different polarization directions. The polarized multi-view image set provides well-organized data for subsequent feature extraction, depth reconstruction, and focal plane prediction, improving the efficiency and maintainability of the entire autofocus system.
[0027] Preferably, step S2 includes the following steps:
[0028] Step S21: Perform preprocessing on the polarization phase gradient spectrum to obtain a preprocessed polarization phase gradient spectrum;
[0029] Step S22: Construct a polarization-aware depth reconstruction network based on the preprocessed polarization phase gradient spectrum to obtain a preliminary depth map;
[0030] Step S23: Perform ray tracing simulation based on the preliminary depth map and the preprocessed polarization phase gradient spectrum, and perform training data augmentation to obtain an enhanced depth map;
[0031] Step S24: Optimize the preliminary depth map and the enhanced depth map to obtain an optimized depth map.
[0032] In the present invention, by preprocessing the polarization phase gradient spectrum (PPGS), it is ensured that it can be effectively utilized by the depth reconstruction network, and the training efficiency and stability of the network are improved. The preprocessing includes dimension adjustment, normalization, and data type conversion. These operations can unify the format of the input data, reduce the impact of numerical differences on training, and improve the calculation efficiency. Data augmentation, such as random cropping, rotation, and adding noise, can increase the diversity of the data and improve the generalization ability of the network, thereby enhancing the robustness of the autofocus system in different scenarios. Constructing a polarization-aware depth reconstruction network enables the effective utilization of polarization information and improves the accuracy of depth reconstruction. The network structure includes an encoder, a polarization-aware module, a decoder, and a loss function. The encoder extracts multi-scale features of the PPPGS, the polarization-aware module specifically processes polarization information, and the decoder maps the features to a depth map. The design of the polarization-aware module enables the network to learn the correlation between features in different polarization directions, thereby better capturing the depth information of the scene. Using the mean square error (MSE) as the loss function can effectively train the network to accurately predict the depth map, thereby improving the accuracy of autofocus. Through ray tracing simulation and training data augmentation, the training dataset is enriched, and the generalization ability and robustness of the depth reconstruction network are enhanced. Ray tracing simulation can generate images and depth maps similar to real scenes, which are used to simulate different lighting conditions, perspective changes, and object movements. Using the data generated by ray tracing simulation to augment the training data enables the network to perform depth reconstruction in various complex scenarios, improving the adaptability of the autofocus system in different environments. By fusing the preliminary depth map and the enhanced depth map, the accuracy and reliability of the depth map are further improved. The fusion operation can combine the prior information of the preliminary depth map and the detailed information of the enhanced depth map to generate a depth map with higher quality. The weighted average fusion method can dynamically adjust the weights according to the quality of the depth map, making the fusion result more accurate. The optimized depth map provides more reliable depth information for subsequent focal plane prediction, thereby improving the overall performance of autofocus.
[0033] Preferably, step S22 includes the following steps:
[0034] Step S221: Extract initial convolutional layer features from the preprocessed polarization phase gradient spectrum to obtain a preliminary feature map; extract downsampling layer features based on the preliminary feature map to obtain a downsampled feature map;
[0035] Step S222: Perform polarization channel separation on the downsampled feature map to obtain the separated polarization features;
[0036] Step S223: Introduce channel attention to the separated polarization features to obtain the polarization features with channel attention;
[0037] Step S224: Perform polarization feature fusion on the polarization features with channel attention to obtain the fused feature map;
[0038] Step S225: Perform feature extraction based on the residual block encoder on the fused feature map to obtain the encoder output features;
[0039] Step S226: Perform feature mapping based on the decoder on the encoder output features to obtain the decoder output features;
[0040] Step S227: Perform depth map output on the decoder output features to obtain the preliminary depth map.
[0041] Through the initial convolutional layer and the downsampling layer, the present invention realizes effective feature extraction and multi-scale representation of PPPGS. The initial convolutional layer can extract preliminary features from the original data, laying a foundation for subsequent processing. The downsampling layer reduces the size of the feature map through convolutional and pooling operations, and extracts higher-level abstract features, enabling the network to capture the global information of the image, reducing the computational amount, and improving the robustness of the network. The multi-scale features enable the depth reconstruction network to better understand the depth information of the scene. Through channel separation, independent processing of information in different polarization directions is realized, enabling the network to better utilize polarization information. The separated polarization features enable the network to perform independent feature learning and processing for each polarization channel, improving the sensitivity of the network to polarization information. This separation processing method can better capture the possible differences and correlations between different polarization directions, thereby improving the accuracy of depth reconstruction. By introducing a channel attention mechanism, adaptive weighting of important polarization feature channels is realized, improving the performance of the network. The channel attention mechanism enables the network to automatically learn the importance of features in different polarization directions, enhance important feature channels, and suppress unimportant feature channels, thereby enhancing the expression ability of features and improving the accuracy of depth reconstruction. By fusing the polarization features with channel attention, integration of information in different polarization channels is realized, thereby enhancing the expression ability of the network. The fusion operation can integrate features in different polarization directions to form a more comprehensive feature representation, thereby improving the network's ability to understand scene information. The fused feature map contains information in different polarization directions, providing a richer feature representation for subsequent depth reconstruction. Using an encoder based on residual blocks can effectively extract deeper-level features and alleviate the problem of gradient disappearance, thereby improving the performance of the model. The design of the residual block enables the network to be trained deeper and extract more complex features. The encoder can transform the fused feature map into a higher-level feature representation, providing more effective features for subsequent depth map generation. Through the decoder, the features extracted by the encoder are mapped to a depth map. The decoder gradually restores the spatial resolution of the feature map through deconvolution, upsampling, and skip connections. The output features of the decoder contain the features used to generate the depth map, providing the necessary information for subsequent depth map output. By mapping the output features of the decoder to a single-channel depth map, a preliminary depth map is obtained. The preliminary depth map contains the depth information of each pixel in the scene, providing a basis for subsequent focal plane prediction.
[0042] Preferably, step S3 includes the following steps:
[0043] Step S31: Perform scene stratification on the optimized depth map to obtain a stratified scene;
[0044] Step S32: Obtain the optical system parameters; perform Fresnel diffraction simulation according to the layered scene to obtain a set of Fresnel diffraction images;
[0045] Step S33: Extract the diffraction features from the set of Fresnel diffraction images to obtain a diffraction feature map;
[0046] Step S34: Evaluate the light field focusing of the layered scene according to the polarization phase gradient spectrum to obtain a light field clarity map;
[0047] Step S35: Calculate the focal plane confidence of the diffraction feature map and the light field clarity map to obtain a focal plane confidence map.
[0048] In the present invention, the optimized depth map is subjected to scene layering, and the three-dimensional scene is decomposed into multiple two-dimensional depth layers, which facilitates subsequent Fresnel diffraction simulation and focal plane prediction. Scene layering can simplify the complexity of subsequent processing, enabling independent analysis of each depth layer. The construction of the layered scene enables more effective evaluation of the focusing situation on different depth layers, improving the accuracy of focal plane prediction. Through Fresnel diffraction simulation, the propagation process of light waves on different depth layers is simulated, thereby enabling prediction of the imaging effects of different depth layers. The Fresnel diffraction simulation takes into account the wave characteristics of light waves and can more accurately simulate the diffraction phenomenon of light waves during propagation. The set of Fresnel diffraction images, which contains the diffraction images of different depth layers, provides key data for subsequent feature extraction and focal plane prediction, thereby improving the accuracy of autofocus. By extracting the diffraction features of the Fresnel diffraction images, the clarity of the images on different depth layers is quantified, providing a basis for subsequent calculation of the focal plane confidence. Diffraction features, such as image contrast, edge sharpness, and frequency domain features, can reflect the clarity of the image. Extracting the diffraction feature map can effectively evaluate the imaging effects of different depth layers, providing important information for focal plane prediction, thereby improving the accuracy of autofocus. Combining the polarization phase gradient spectrum and the layered scene realizes the evaluation of the light field focusing of each depth layer, providing more comprehensive information for the calculation of the focal plane confidence. The light field focusing evaluation combines light field information and depth information and can more accurately evaluate the focusing degree of different depth layers. The light field clarity map, which contains the clarity evaluation results of each depth layer, provides important supplementary information for subsequent calculation of the focal plane confidence, thereby improving the robustness of autofocus. By fusing the diffraction feature map and the light field clarity map, the focal plane confidence of each depth layer is calculated, enabling accurate prediction of the optimal focal plane position. The focal plane confidence map can comprehensively consider the results of Fresnel diffraction simulation and light field focusing evaluation, thereby more comprehensively and accurately evaluating the focusing degree of different depth layers. The focal plane confidence map, which contains the focusing confidence information of different depth layers, provides key information for subsequent generation of lens control instructions, thereby improving the accuracy of autofocus.
[0049] Preferably, step S32 includes the following steps:
[0050] Step S321: Construct an input light field according to the layered scenario to obtain the input light field;
[0051] Step S322: Perform light field filling on the input light field to obtain the filled light field;
[0052] Step S323: Perform a light field Fourier transform on the filled light field to obtain the frequency-domain light field;
[0053] Step S324: Calculate the propagation function according to the optical system parameters and the frequency-domain light field to obtain the propagation function;
[0054] Step S325: Perform frequency-domain multiplication on the propagation function and the frequency-domain light field to obtain the propagated frequency-domain light field;
[0055] Step S326: Perform an inverse light field Fourier transform on the propagated frequency-domain light field to obtain the diffraction image;
[0056] Step S327: Perform Fresnel diffraction simulation for each depth layer according to the diffraction image and the layered scenario to obtain a set of Fresnel diffraction images.
[0057] The present invention provides a basis for subsequent Fresnel diffraction simulation by constructing an input optical field and converting the depth information of a layered scene into the complex amplitude distribution of light waves. The input optical field can accurately describe the initial state of light waves on different depth layers, ensuring the accuracy of the diffraction simulation. The construction of the input optical field enables the simulation of the propagation process of light waves based on depth information, improving the accuracy of autofocus. By filling the input optical field, edge effects in subsequent Fourier transform and diffraction simulation are avoided, ensuring the accuracy of the simulation. The filling operation can eliminate the influence of the periodic boundary conditions of the Fourier transform, making the simulation results more reliable. The filled optical field provides a more stable input for subsequent calculations, thereby improving the accuracy of autofocus. Through Fourier transform, the filled optical field is converted from the spatial domain to the frequency domain, facilitating subsequent calculation of the propagation function and diffraction simulation. The frequency-domain optical field can describe the frequency components of light waves, simplifying the subsequent calculation process. Fourier transform can be calculated efficiently, improving the efficiency of the simulation. By calculating the propagation function, the changes of light waves during propagation are described, providing key information for subsequent diffraction simulation. The propagation function can accurately describe the propagation characteristics of light waves on different depth layers, ensuring the accuracy of the diffraction simulation. The calculation of the propagation function enables the simulation of the propagation process of light waves based on optical system parameters and depth information, improving the accuracy of autofocus. By multiplying the propagation function and the frequency-domain optical field in the frequency domain, the changes of light waves during propagation are simulated, and the propagated frequency-domain optical field is obtained. Frequency-domain multiplication can efficiently simulate the propagation process of light waves, and the simulation results are more accurate. The propagated frequency-domain optical field contains the frequency-domain information of light waves after propagation, providing an input for subsequent inverse Fourier transform. Through inverse Fourier transform, the propagated frequency-domain optical field is converted from the frequency domain to the spatial domain, obtaining a diffraction image and simulating the light intensity distribution of light waves after propagation to the focal plane. The diffraction image can intuitively display the imaging effects of different depth layers, providing key data for subsequent feature extraction and focal plane prediction. Inverse Fourier transform can be calculated efficiently, improving the efficiency of the simulation. By performing Fresnel diffraction simulation on each depth layer in the layered scene, a set of Fresnel diffraction images is obtained, enabling the simulation of the imaging effects of different depth layers and providing key data for subsequent focal plane prediction. The set of Fresnel diffraction images contains diffraction images of different depth layers, comprehensively reflecting the light intensity distribution of different depth layers. The Fresnel diffraction simulation for each depth layer enables the clarity evaluation for different depth layers, improving the accuracy of focal plane prediction.
[0058] Preferably, step S4 includes the following steps:
[0059] Step S41: Determine the optimal focal plane position for the focal plane confidence map to obtain the optimal focal plane depth;
[0060] Step S42: Obtain camera parameters; calculate the lens displacement according to the optimal focal plane depth and the camera parameters to obtain the lens displacement amount;
[0061] Step S43: Correct the lens displacement according to the lens displacement amount and the optical system parameters to obtain the corrected lens displacement;
[0062] Step S44: Generate a lens control command for the corrected lens displacement to obtain the lens control command.
[0063] The present invention realizes the prediction of the position of the clearest focal plane in the scene by determining the optimal focal plane depth, providing a target for subsequent lens control. The processing of the focal plane confidence map can quickly and accurately find the depth value with the highest confidence, that is, the optimal focal plane depth. The determination of the optimal focal plane depth enables the lens to be adjusted to the clearest position, improving the accuracy of autofocus. By calculating the lens displacement amount, the distance that the lens needs to move is determined, providing a basis for subsequent lens control. The calculation of the lens displacement amount takes into account parameters such as the focal length of the camera and the sensor size, ensuring the accuracy of the calculation. The calculation of the lens displacement amount enables the lens to be moved to the position corresponding to the optimal focal plane depth, thereby realizing autofocus. Through lens displacement correction, the accuracy of the lens displacement is improved, thereby improving the focusing accuracy. The lens displacement correction takes into account optical system parameters such as the aperture value and zoom information, making the lens displacement amount more accurate. The corrected lens displacement can more accurately control the movement of the lens, thereby realizing more precise autofocus. By generating a lens control command, the control of the lens motor is realized, thereby controlling the lens to move to the target position. The lens control command converts the corrected lens displacement into a signal that the lens motor can understand, facilitating the driving of the lens motor. The generation of the lens control command enables the automatic control of the lens movement, thereby realizing autofocus.
[0064] Preferably, step S5 includes the following steps:
[0065] Step S51: Move the lens according to the lens control command and perform current image acquisition to obtain the current image;
[0066] Step S52: Evaluate the clarity of the current image to obtain the image clarity value;
[0067] Step S53: Use a preset clarity threshold to judge the focusing state of the image clarity value to obtain the focusing state;
[0068] Step S54: Perform focusing iterative optimization according to the focusing state, the image clarity value, the focal plane confidence map, and the corrected lens displacement to obtain the optimized lens control command.
[0069] The present invention realizes precise control of the lens position by moving the lens according to the lens control instruction, and acquires the current image, providing a data basis for subsequent focus state judgment and iterative optimization. The lens movement can move the lens to the target position and acquire the current image for evaluating the clarity of the image, thereby judging the focus state. Acquiring the current image can obtain the scene information at the current lens position, providing a data basis for subsequent processing. By evaluating the clarity of the current image, the clarity degree of the image is quantified, providing a basis for subsequent focus state judgment. The image clarity evaluation can effectively evaluate the clarity degree of the image and obtain the image clarity value for judging the focus state. The calculation of the image clarity value enables the quantification of the clarity degree of the image, providing a reference for subsequent iterative optimization. By comparing the image clarity value with the clarity threshold, the focus state is judged, providing a control condition for subsequent iterative optimization. The focus state judgment can determine whether the current image is already in focus and provide a control condition for subsequent iterative optimization. The judgment of the focus state enables the determination of whether further iterative optimization is required, thereby realizing autofocus. Through focus iterative optimization, the step-by-step adjustment of the lens position is realized, improving the accuracy and efficiency of focusing. The focus iterative optimization can perform step-by-step adjustment of the lens position according to the focus state, image clarity value, focal plane confidence map, and corrected lens displacement. The iterative optimization can continuously adjust the lens position until the image clarity reaches a preset threshold or a preset number of iterations, thereby improving the accuracy and efficiency of autofocus.
[0070] Preferably, the present invention also provides an optical lens autofocus system for executing the optical lens autofocus method as described above. The optical lens autofocus system includes:
[0071] A light field feature extraction module for collecting a polarized light field through a light field camera to obtain a polarized multi-view image set; and extracting light field features from the polarized multi-view image set to obtain a polarized phase gradient spectrum.
[0072] A depth information reconstruction module for constructing a polarization-aware depth reconstruction network according to the polarized phase gradient spectrum to obtain a preliminary depth map; performing ray tracing simulation processing on the preliminary depth map to obtain an enhanced depth map; and optimizing the preliminary depth map and the enhanced depth map to obtain an optimized depth map.
[0073] A focal plane prediction module for obtaining optical system parameters; performing scene layering on the optimized depth map to obtain a layered scene; performing Fresnel diffraction simulation according to the optical system parameters and the layered scene to obtain a set of Fresnel diffraction images; and predicting the focal plane from the set of Fresnel diffraction images to obtain a focal plane confidence map.
[0074] A lens control instruction generation module, which is used to calculate the lens displacement according to the focal plane confidence map to obtain the lens displacement amount; correct the lens displacement according to the lens displacement amount, and generate a lens control instruction to obtain the corrected lens displacement and the lens control instruction;
[0075] A focusing confirmation and optimization module, which is used to move the lens according to the lens control instruction, and judge the focusing state to obtain the focusing state; perform focusing iterative optimization according to the focusing state to obtain an optimized lens control instruction. Description of the Drawings
[0076] Figure 1 It is a schematic flow chart of the steps of an optical lens autofocus method;
[0077] Figure 2 It is a schematic detailed implementation step flow chart of step S2 in the present invention.
[0078] The realization of the object, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments
[0079] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention.
[0080] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0081] It should be understood that although the terms "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.
[0082] To achieve the above object, please refer to Figures 1 to 2, an automatic focusing method for an optical lens, comprising the following steps:
[0083] Step S1: Collect a polarized light field through a light field camera to obtain a set of polarized multi-view images; extract light field features from the set of polarized multi-view images to obtain a polarized phase gradient spectrum;
[0084] Step S2: Construct a polarization-aware depth reconstruction network based on the polarized phase gradient spectrum to obtain a preliminary depth map; perform ray tracing simulation processing on the preliminary depth map to obtain an enhanced depth map; optimize the preliminary depth map and the enhanced depth map to obtain an optimized depth map;
[0085] Step S3: Obtain optical system parameters; layer the scene of the optimized depth map to obtain a layered scene; perform Fresnel diffraction simulation based on the optical system parameters and the layered scene to obtain a set of Fresnel diffraction images; perform focal plane prediction on the set of Fresnel diffraction images to obtain a focal plane confidence map;
[0086] Step S4: Calculate the lens displacement according to the focal plane confidence map to obtain a lens displacement amount; correct the lens displacement according to the lens displacement amount and generate a lens control instruction to obtain a corrected lens displacement and a lens control instruction;
[0087] Step S5: Move the lens according to the lens control instruction and judge the focusing state to obtain a focusing state; perform focusing iterative optimization according to the focusing state to obtain an optimized lens control instruction.
[0088] In the embodiment of the present invention, refer to Figure 1 As shown, it is a schematic flow chart of the steps of the automatic focusing method for the optical lens of the present invention. In this example, the automatic focusing method for the optical lens comprises the following steps:
[0089] Step S1: Collect a polarized light field through a light field camera to obtain a set of polarized multi-view images; extract light field features from the set of polarized multi-view images to obtain a polarized phase gradient spectrum;
[0090] In the embodiment of the present invention, the core lies in collecting and extracting light field information and fusing polarization characteristics. First, a combination of a light field camera and a rotatable linear polarizer is used to collect multi-view images at different polarization directions to form a set of polarized multi-view images. Then, perform Fourier transform on each viewpoint image to obtain a Fourier transform spectrum. Next, calculate the phase gradient of the Fourier transform spectrum to obtain a phase gradient map, and further calculate the polarization-related phase gradient to obtain a polarized phase gradient map. Finally, generate a multi-dimensional light field feature vector from the polarized phase gradient map to obtain a polarized phase gradient spectrum, which contains the polarization-related phase gradient information of the light field in the Fourier domain and serves as the input for subsequent depth reconstruction.
[0091] Step S2: Construct a polarization perception depth reconstruction network based on the polarization phase gradient spectrum to obtain a preliminary depth map; perform ray tracing simulation processing on the preliminary depth map to obtain an enhanced depth map; optimize the preliminary depth map and the enhanced depth map to obtain an optimized depth map.
[0092] In the embodiment of the present invention, a depth reconstruction network is mainly constructed, and ray tracing is used to enhance the training data. First, preprocess the polarization phase gradient spectrum, including dimension adjustment, normalization, and data type conversion. Then, construct a polarization perception depth reconstruction network, which includes an encoder, a polarization perception module, a decoder, and a loss function. The encoder extracts multi-scale features, the polarization perception module processes polarization information, and the decoder maps the features to a depth map. Then, use the preliminary depth map and the preprocessed polarization phase gradient spectrum to perform ray tracing simulation to generate simulated images and depth maps for enhancing the training data. Finally, fuse the preliminary depth map and the enhanced depth map to obtain an optimized depth map.
[0093] Step S3: Obtain the optical system parameters; perform scene stratification on the optimized depth map to obtain a stratified scene; perform Fresnel diffraction simulation according to the optical system parameters and the stratified scene to obtain a set of Fresnel diffraction images; perform focal plane prediction on the set of Fresnel diffraction images to obtain a focal plane confidence map.
[0094] In the embodiment of the present invention, based on the optimized depth map, scene stratification, Fresnel diffraction simulation are mainly performed, and the focal plane confidence is calculated. First, according to the optimized depth map, the scene is segmented into multiple depth layers to obtain a stratified scene. Then, obtain the optical system parameters, and perform Fresnel diffraction simulation according to the stratified scene and the optical system parameters to generate a set of Fresnel diffraction images. Next, extract the diffraction features of the set of Fresnel diffraction images, and perform light field focusing evaluation on the stratified scene according to the polarization phase gradient spectrum to obtain a light field clarity map. Finally, fuse the diffraction feature map and the light field clarity map, calculate the focal plane confidence of each depth layer, and obtain a focal plane confidence map.
[0095] Step S4: Calculate the lens displacement according to the focal plane confidence map to obtain the lens displacement amount; perform lens displacement correction according to the lens displacement amount, and generate a lens control instruction to obtain the corrected lens displacement and the lens control instruction.
[0096] In the embodiments of the present invention, a lens control instruction is mainly calculated and generated based on a focal plane confidence map. First, the pixel with the highest confidence in the focal plane confidence map is found to obtain the optimal focal plane depth. Then, according to the optimal focal plane depth and camera parameters, the lens displacement amount is calculated. Next, according to the lens displacement amount and optical system parameters, the lens displacement amount is corrected to obtain the corrected lens displacement. Finally, according to the corrected lens displacement, a lens control instruction is generated to drive the lens motor to move to the target position.
[0097] Step S5: Move the lens according to the lens control instruction, and judge the focusing state to obtain the focusing state; perform focusing iterative optimization according to the focusing state to obtain an optimized lens control instruction;
[0098] In the embodiments of the present invention, focusing state judgment and iterative optimization are mainly performed. First, according to the lens control instruction, the lens motor is driven to move, and the current image is collected. Then, the current image is evaluated for image sharpness to obtain an image sharpness value. Next, the image sharpness value is compared with a preset sharpness threshold to judge the focusing state. If the focusing fails, focusing iterative optimization is performed according to the focusing state, the image sharpness value, the focal plane confidence map, and the corrected lens displacement to generate a new lens control instruction. Repeat the above process until the focusing is successful or the preset number of iterations is reached.
[0099] Preferably, step S1 includes the following steps:
[0100] Step S11: Collect a polarized light field through a light field camera to obtain a set of polarized multi-view images;
[0101] Step S12: Perform a Fourier transform on the viewpoint images of the set of polarized multi-view images to obtain a Fourier transform spectrum;
[0102] Step S13: Calculate the phase gradient of the Fourier transform spectrum to obtain a phase gradient map;
[0103] Step S14: Calculate the polarized phase gradient of the phase gradient map to obtain a polarized phase gradient map;
[0104] Step S15: Generate a multi-dimensional light field feature vector for the polarized phase gradient map to obtain a polarized phase gradient spectrum.
[0105] In an embodiment of the present invention, a light field camera equipped with a microlens array (MLA) is configured and integrated with a rotatable linear polarizer. The MLA is located in front of the image sensor and is used to sample light rays spatially to obtain light ray information at different angles. The linear polarizer is placed in front of the MLA and is used to control the polarization direction of the light rays entering the camera. Set the rotation angle of the polarizer, for example, set it to 0°, 45°, 90°, and 135° in sequence. At each polarization angle, control the camera to collect light field data. At each polarization angle, the light field camera collects multi-view images through the MLA. The MLA decomposes the incident light rays into multiple sub-beams, and each sub-beam corresponds to a view point. The image sensor records the intensity and color information of each sub-beam to generate multi-view images. Repeat the above process to collect multi-view images at all set polarization angles. Combine the multi-view images collected at different polarization angles to form a polarized multi-view image set. The polarized multi-view image set contains multi-view images at multiple polarization directions, and each view corresponds to a scene at a specific polarization direction.
[0106] Extract the multi-view images of one polarization direction from the polarized multi-view image set. Select a view point image from the multi-view images, and this view point image represents the scene information observed from a specific perspective. Perform a two-dimensional fast Fourier transform (FFT) on the view point image. Apply the FFT to each pixel of the view point image to calculate the discrete Fourier transform of each pixel. The FFT transforms the image from the spatial domain to the frequency domain to obtain a complex matrix, and the real part and the imaginary part of this matrix respectively represent the frequency components of the image. Repeat the above FFT operation for all view point images in the polarized multi-view image set. Combine the Fourier transform results of all view point images into a Fourier transform spectrum. The Fourier transform spectrum is composed of multiple complex matrices, each matrix corresponds to the Fourier transform result of a view point image, and each matrix is associated with a specific polarization direction.
[0107] Select a complex matrix from the Fourier transform spectrum, where the matrix corresponds to the Fourier transform result of a viewpoint image. Calculate the phase of the complex matrix. The phase represents the relative positions of different frequency components in the image. For each element in the complex matrix, calculate its phase using the arctangent function (arctan). Calculate the gradients of the phase in the horizontal and vertical directions. Convolve the phase using the Sobel operator or the Scharr operator to calculate the differences of the phase in the horizontal and vertical directions. The Sobel operator or the Scharr operator is a set of predefined convolution kernels used to detect edges in an image. Combine the calculated phase gradients in the horizontal and vertical directions to form a phase gradient vector. Calculate the magnitude of the phase gradient vector to obtain the size of the phase gradient. Repeat the above operations for all complex matrices in the Fourier transform spectrum. Combine the phase gradient sizes of all viewpoint images to form a phase gradient map. The phase gradient map contains the phase gradient information of all viewpoint images, and each pixel value represents the size of the phase gradient at that pixel position.
[0108] Select two phase gradient maps from the phase gradient map. These two phase gradient maps correspond to the same viewpoint position but have different polarization directions. Calculate the difference between these two phase gradient maps. The difference can be calculated as the difference of the corresponding pixels or the absolute value of the difference of the corresponding pixels. Use the calculated difference as the new phase gradient value. Repeat the above operations for all phase gradient maps with different polarization directions but the same viewpoint position in the phase gradient map. Combine the calculated polarization phase gradient values to form a polarization phase gradient map. The polarization phase gradient map contains the polarization phase gradient information of all viewpoint images, and each pixel value represents the polarization phase gradient value at that pixel position.
[0109] Select a polarization phase gradient map from the polarization phase gradient map. Reshape the pixel values of the polarization phase gradient map to convert it into a one-dimensional vector. For example, the pixel values of the polarization phase gradient map can be arranged into a vector in row-major or column-major order. Repeat this operation for all polarization phase gradient maps to convert each polarization phase gradient map into a vector. Combine all the vectors together to form a multi-dimensional feature vector. This multi-dimensional feature vector represents the comprehensive features of the light field and contains the phase gradient information of different viewpoints and different polarization directions. Use this multi-dimensional feature vector as the polarization phase gradient spectrum. The polarization phase gradient spectrum is a multi-dimensional feature vector used for subsequent depth reconstruction and focal plane prediction.
[0110] Preferably, step S11 includes the following steps:
[0111] Step S111: Place an angle-calibrated linear polarizer in front of the light field camera, perform synchronous control of the light field camera and the polarizer, and obtain synchronous control parameters;
[0112] Step S112: Collect multi-polarization-direction light field data according to the synchronization control parameters to obtain an original light field image set;
[0113] Step S113: Correct the influence of the polarizer on the light field for the original light field image set to obtain a corrected light field image set;
[0114] Step S114: Store and label the corrected light field image set to obtain a polarization multi-view image set.
[0115] In an embodiment of the present invention, prepare a light field camera equipped with a microlens array (MLA) for capturing light field information and a rotatable linear polarizer. Place the linear polarizer in front of the light field camera lens to ensure that it can cover the entire field of view. Use precise angle calibration equipment, such as an optical platform and an angle measuring instrument, to accurately calibrate the angle of the linear polarizer. The purpose of calibration is to ensure that the polarization direction of the linear polarizer is consistent with the set angle. Design a synchronization control system that can simultaneously control the image acquisition of the light field camera and the rotation of the linear polarizer. The system includes a microcontroller for sending control signals. Connect the microcontroller to the light field camera and the polarizer rotation mechanism. Set the rotation angles of the polarizer, such as 0°, 45°, 90°, and 135°. The microcontroller sends control signals to the polarizer rotation mechanism to rotate it to the set angle. At the same time, the microcontroller sends an instruction to the light field camera to acquire images. The synchronization control parameters include: the polarizer rotation angle sequence, the image acquisition trigger signal at each angle, and the exposure time and gain settings for the acquired images. Store these parameters in the memory of the microcontroller for subsequent data acquisition.
[0116] Start the synchronization control system and collect multi-polarization-direction light field data according to the preset synchronization control parameters. The microcontroller first controls the polarizer to rotate to the set first angle, such as 0°. After the polarizer rotates in place, the microcontroller sends an image acquisition trigger signal to the light field camera. The light field camera acquires a multi-view image according to the set exposure time and gain settings. Store the acquired multi-view image in the internal memory of the light field camera or transmit it to an external storage device through a data cable. Repeat the above process, control the polarizer to rotate to the next set angle, such as 45°, and acquire a multi-view image. Repeat this process in sequence until the acquisition of multi-view images in all preset polarization directions is completed. Combine all the acquired multi-view images into an original light field image set. The original light field image set contains multi-view images in multiple polarization directions, and each multi-view image corresponds to a specific polarization angle.
[0117] Since the linear polarizer is not an ideal polarization device, there are differences in the transmittance of light with different polarization directions, and additional optical losses will be introduced. The original light field image set is corrected to eliminate these effects. Prepare a uniform, polarization-independent light source, for example, an integrating sphere light source. Place this light source in front of the light field camera. At each preset polarization angle, capture the image of the integrating sphere light source. Capture multiple integrating sphere images for calculating the transmittance and optical loss of the polarizer. For each polarization angle, calculate a correction coefficient. The correction coefficient can be calculated by comparing the average pixel values of the images at different polarization angles. For each multi-view image in the original light field image set, apply the corresponding correction coefficient according to its corresponding polarization angle. For each pixel in the multi-view image, use the correction coefficient for correction. For example, the light intensity can be corrected by multiplying the pixel value by the correction coefficient. Combine the corrected multi-view images into a corrected light field image set. The corrected light field image set eliminates the influence of the polarizer on the light field and improves the accuracy of subsequent processing.
[0118] For each multi-view image in the corrected light field image set, perform data storage and labeling. Select a suitable file format, for example, TIFF, PNG, or RAW, for storing the multi-view images. Name the file name of each multi-view image, and the file name contains the polarization angle information corresponding to the image. For example, "image_000_0.tif" can be used to represent the image with a polarization angle of 0°, and "image_000_45.tif" can be used to represent the image with a polarization angle of 45°. Establish a data index file, which records the file name, polarization angle, and other relevant information of each multi-view image, such as acquisition time, exposure time, and gain setting. Store all the multi-view images in the corrected light field image set in a storage device, and also store the data index file in the storage device. The stored corrected light field image set and the data index file together form a polarization multi-view image set. The polarization multi-view image set contains the corrected multi-view images, as well as the metadata and relevant information for describing these images.
[0119] Preferably, step S2 includes the following steps:
[0120] Step S21: Perform preprocessing on the polarization phase gradient spectrum to obtain a preprocessed polarization phase gradient spectrum;
[0121] Step S22: Construct a polarization perception depth reconstruction network based on the preprocessed polarization phase gradient spectrum to obtain a preliminary depth map;
[0122] Step S23: Perform ray tracing simulation based on the preliminary depth map and the preprocessed polarization phase gradient spectrum, and perform training data augmentation to obtain an enhanced depth map;
[0123] Step S24: Optimize the preliminary depth map and the enhanced depth map to obtain an optimized depth map.
[0124] As an example of the present invention, refer to Figure 2 As shown, in this example, step S2 includes:
[0125] Step S21: Perform preprocessing on the polarization phase gradient spectrum to obtain a preprocessed polarization phase gradient spectrum;
[0126] In the embodiment of the present invention, starting from the polarization phase gradient spectrum (PPGS) obtained in step S1, a preprocessing operation is performed to improve its performance for the depth reconstruction network. First, adjust the dimensions of the PPGS. Since the PPGS has different dimensions and data types, it needs to be converted into the input format of the depth reconstruction network. If the PPGS is a multi-dimensional feature vector, reshape it into a four-dimensional tensor with dimensions (batch_size, height, width, channels). Here, batch_size represents the batch size, height and width represent the height and width of the feature map respectively, and channels represent the number of channels of the feature map. Second, perform data normalization. Normalize each element in the PPGS so that its numerical range is between 0 and 1. Normalization can adopt the Min - Max normalization method, calculate the minimum and maximum values of each feature, and then use the formula: (x - min) / (max - min) to scale all elements. If the numerical range of the PPGS is already close to 0 to 1, this step can be omitted. Third, perform data type conversion. Convert the data type of the PPGS to the data type supported by the depth reconstruction network, for example, float32. Data type conversion can improve calculation efficiency and numerical stability. Fourth, if necessary, perform data augmentation. For example, operations such as randomly cropping, rotating, or adding noise to the PPGS can be performed to increase data diversity and improve the generalization ability of the model. The PPGS after the above processing is used as the preprocessed polarization phase gradient spectrum (PPPGS) and is ready to be input into the depth reconstruction network.
[0127] Step S22: Construct a polarization-aware depth reconstruction network based on the preprocessed polarization phase gradient spectrum to obtain a preliminary depth map;
[0128] In the embodiments of the present invention, a polarization-aware depth reconstruction network based on a deep convolutional neural network (CNN) is constructed. This network receives PPPGS as input and outputs a preliminary depth map. The network structure includes the following main components. First, an encoder is designed to extract multi-scale features from PPPGS. The encoder includes multiple convolutional layers, pooling layers, and activation functions. The convolutional layers are used to extract local features of PPPGS, the pooling layers are used to reduce the dimension of the feature map, and the ReLU activation function is used to introduce non-linearity. The encoder adopts residual connections to alleviate the problem of gradient vanishing. Second, a polarization-aware module is constructed. This module is embedded in some layers of the encoder and is used to process polarization information. The polarization-aware module can be a network branch containing multiple convolutional layers and an attention mechanism. This module receives polarization-related features in PPPGS as input and uses a cross-channel attention mechanism to enable the network to learn the correlation between features in different polarization directions. For example, the correlation between different polarization channels can be calculated, and attention weights are used to weight the features of different channels. The output of the polarization-aware module is fused with the features of the backbone encoder, for example, by concatenation or addition. Third, a decoder is constructed to map the features extracted by the encoder to a depth map. The decoder includes multiple transposed convolutional layers, upsampling layers, and skip connections. The transposed convolutional layers are used to increase the spatial resolution of the feature map, the upsampling layers are used to enlarge the feature map to the same size as the input image, and the skip connections are used to transfer low-level features in the encoder to the decoder to retain detailed information. The last layer of the decoder outputs a single-channel depth map representing the depth value of each pixel in the scene. The mean square error (MSE) is used as the loss function to measure the difference between the predicted depth map and the true depth map. The Adam optimizer is used to train the network and optimize the network parameters. The trained network is applied to PPPGS to obtain a preliminary depth map.
[0129] Step S23: Perform ray tracing simulation based on the preliminary depth map and the preprocessed polarization phase gradient spectrum, and perform training data augmentation to obtain an enhanced depth map;
[0130] In the embodiments of the present invention, ray tracing technology is used, combined with a preliminary depth map and PPPGS, to enhance training data. First, a ray tracing engine is selected, for example, Mitsuba or POV-Ray. A geometric model similar to the actual scene is prepared, including the objects, materials, and lighting conditions in the scene. Using the preliminary depth map, the depth information of the scene is constructed. The preliminary depth map is used as the input of the ray tracing engine to simulate the propagation of light in the scene. PPPGS is used to provide polarization information for the ray tracing simulation. For example, according to the polarization information in PPPGS, the reflection and refraction of light at different polarization angles can be simulated. The ray tracing engine is run to generate simulated images and corresponding depth maps. The simulated images contain images from different viewpoints and can simulate the perspective changes of the camera. The simulated depth map contains the true depth information of each pixel in the scene. The simulated images and depth maps are used to enhance the training data. For example, different lighting conditions can be simulated, including changing the intensity, color, and position of the light source. The movement and rotation of objects in the scene can be simulated. The changes in camera parameters, such as focal length and aperture, can be simulated. The enhanced data is mixed with the original training data for retraining the depth reconstruction network. When retraining the depth reconstruction network, the parameters of the network can be adjusted to better adapt to the new training data. After the training using the enhanced data is completed, the enhanced depth map is used.
[0131] Step S24: Optimize the preliminary depth map and the enhanced depth map to obtain an optimized depth map;
[0132] In the embodiments of the present invention, the preliminary depth map and the enhanced depth map are fused to improve the quality of the depth map. First, the preliminary depth map and the enhanced depth map are registered. The purpose of registration is to align the two depth maps to ensure that the same pixel positions correspond to the same scene points. Registration can use image registration algorithms, such as feature point-based registration or patch-based registration. After registration, the preliminary depth map and the enhanced depth map are fused. The fusion method can adopt weighted averaging. Different weights are assigned to the preliminary depth map and the enhanced depth map. The selection of weights can be adjusted according to the quality of the depth map. For example, if the quality of the enhanced depth map is higher, a larger weight can be assigned to it. The weights can be determined according to the confidence levels of the preliminary depth map and the enhanced depth map. For example, if the confidence level of the preliminary depth map is higher, a larger weight can be assigned to it. The weighted depth maps are fused to obtain an optimized depth map. The optimized depth map contains the depth information of each pixel in the scene and has higher accuracy and robustness.
[0133] Preferably, step S22 includes the following steps:
[0134] Step S221: Extract initial convolutional layer features from the preprocessed polarization phase gradient spectrum to obtain a preliminary feature map; extract downsampling layer features based on the preliminary feature map to obtain a downsampled feature map;
[0135] Step S222: Separate the polarization channels of the downsampled feature map to obtain separated polarization features;
[0136] Step S223: Introduce channel attention to the separated polarization features to obtain polarization features with channel attention;
[0137] Step S224: Fuse the polarization features with channel attention to obtain a fused feature map;
[0138] Step S225: Extract features based on the residual block encoder from the fused feature map to obtain encoder output features;
[0139] Step S226: Perform feature mapping based on the decoder on the encoder output features to obtain decoder output features;
[0140] Step S227: Output a depth map from the decoder output features to obtain a preliminary depth map.
[0141] In the embodiment of the present invention, the preprocessed polarization phase gradient spectrum (PPPGS) obtained in step S21 is received and used as the input of the depth reconstruction network. An initial convolutional layer is designed, which consists of multiple convolutional kernels, and the size of the convolutional kernel is set to 3x3, for example. The PPPGS is input into the initial convolutional layer for convolution operation. The convolution operation uses the ReLU activation function to introduce non-linearity. The ReLU activation function sets the values less than 0 to 0 and keeps the values greater than 0 unchanged. The output result of the initial convolutional layer is a preliminary feature map. The preliminary feature map contains the preliminary features extracted from the PPPGS. The number of channels of the preliminary feature map depends on the number of convolutional kernels in the initial convolutional layer. Based on the preliminary feature map, multiple downsampling layers are designed to extract higher-level features and reduce the size of the feature map. The downsampling layer includes a convolutional layer and a pooling layer. The size of the convolutional kernel of the convolutional layer is set to 3x3, for example, and the ReLU activation function is used. The pooling layer uses maxpooling, and the size of the pooling window is set to 2x2, for example, with a stride of 2. The maxpooling operation selects the largest value from the pooling window. After passing through the downsampling layer, the size of the feature map is reduced, but the abstraction level of the features is higher. The feature map processed by multiple downsampling layers is used as the downsampled feature map. The number of channels of the downsampled feature map depends on the number of convolutional kernels in the downsampling layer.
[0142] Receive the preprocessed polarization phase gradient spectrum (PPPGS) obtained in step S21 and use it as the input to the depth reconstruction network. Design an initial convolutional layer composed of multiple convolutional kernels, and the size of the convolutional kernels is set to 3x3, for example. Input the PPPGS into the initial convolutional layer for convolutional operation. The convolutional operation uses the ReLU activation function to introduce non-linearity. The ReLU activation function sets the values less than 0 to 0 and keeps the values greater than 0 unchanged. The output result of the initial convolutional layer is a preliminary feature map. The preliminary feature map contains the preliminary features extracted from the PPPGS. The number of channels of the preliminary feature map depends on the number of convolutional kernels in the initial convolutional layer. According to the preliminary feature map, design multiple downsampling layers to extract higher-level features and reduce the size of the feature map. The downsampling layer contains a convolutional layer and a pooling layer. The size of the convolutional kernels in the convolutional layer is set to 3x3, for example, and the ReLU activation function is used. The pooling layer uses max pooling, and the size of the pooling window is set to 2x2, for example, with a stride of 2. The max pooling operation selects the maximum value from the pooling window. After passing through the downsampling layer, the size of the feature map decreases, but the degree of abstraction of the features is higher. Use the feature map processed by multiple downsampling layers as the downsampled feature map. The number of channels of the downsampled feature map depends on the number of convolutional kernels in the downsampling layer.
[0143] Introduce a channel attention mechanism for the separated polarization features so that the network can adaptively focus on important polarization feature channels. For each separated polarization feature, construct a channel attention module. The channel attention module contains a global average pooling layer, a fully connected layer, and a Sigmoid activation function. The global average pooling layer averages the pixel values of each channel of the feature map to obtain a value representing the importance of the channel. The fully connected layer is used to reduce and increase the dimension of the output of the global average pooling. For example, reduce the number of channels to 1 / 16 of the original, and then increase it back to the original number of channels. The Sigmoid activation function maps the output of the fully connected layer to between 0 and 1 to obtain the channel attention weight. Multiply the channel attention weight by the corresponding polarization feature to obtain the polarization features with channel attention. For example, multiply the features of each channel by the corresponding attention weight to enhance the features of important channels and suppress the features of unimportant channels.
[0144] Fuse the polarization features with channel attention to integrate information in different polarization directions. Concatenate all the polarization features with channel attention. The concatenation operation connects the features in different polarization directions in the channel dimension to form a fused feature map. For example, if four polarization directions are used and the number of channels of the feature map in each polarization direction is C, the number of channels of the concatenated feature map is 4C. Perform a convolution operation on the concatenated feature map. The size of the convolution kernel of the convolution operation is set to 1x1, for example, to perform further feature extraction on the fused features. The convolution operation uses the ReLU activation function to introduce non-linearity. Take the output result of the convolution operation as the fused feature map.
[0145] Adopt the residual block as the core building block of the encoder to perform deeper feature extraction on the fused feature map. The design of the residual block can alleviate the problem of gradient disappearance, enabling the network to be trained deeper. The residual block contains multiple convolutional layers, batch normalization layers, and ReLU activation functions. The size of the convolution kernel of the convolutional layer is set to 3x3, for example. The batch normalization layer normalizes the feature map of each batch to accelerate the training process and improve the generalization ability of the model. The ReLU activation function introduces non-linearity. The residual block also contains a skip connection that directly connects the input feature to the output of the residual block. The skip connection enables the network to learn the residual, that is, the difference between the input and the output. Design multiple residual blocks and connect them in series to form an encoder. The output of each residual block of the encoder is used as the input of the next residual block. The output of the last layer of the encoder is used as the encoder output feature. The encoder output feature contains higher-level features extracted after the fused feature map has been processed through multiple layers.
[0146] Construct a decoder for mapping the encoder output feature to a depth map. The decoder includes multiple deconvolutional layers, upsampling layers, and skip connections. The deconvolutional layer is used to increase the spatial resolution of the feature map. The kernel size of the deconvolutional layer is set to 3x3, for example, and the stride is 2. The upsampling layer is used to enlarge the feature map to the same size as the input image. The upsampling layer can use methods such as bilinear interpolation or transposed convolution. The skip connection passes the low-level features in the encoder to the decoder to retain detailed information. Input the encoder output feature into the decoder. Through multiple deconvolutions, upsamplings, and skip connections, the decoder gradually restores the feature map to the same size as the input image. The output of the last layer of the decoder is used as the decoder output feature. The decoder output feature contains the features mapped from the encoder output feature for generating the depth map.
[0147] Use the decoder output feature to generate a preliminary depth map. Perform a 1x1 convolution operation on the decoder output feature. The number of output channels of this convolution operation is 1, which is used to map the feature map to a single-channel depth map. No activation function is used in the convolution operation. Take the output result of the convolution operation as the preliminary depth map. Each pixel value in the preliminary depth map represents the depth value at the corresponding position in the scene. The depth value can be the distance from an object in the scene to the camera, or other ways of representing depth.
[0148] Preferably, step S3 includes the following steps:
[0149] Step S31: Perform scene stratification on the optimized depth map to obtain a stratified scene;
[0150] Step S32: Obtain the optical system parameters; perform Fresnel diffraction simulation according to the stratified scene to obtain a set of Fresnel diffraction images;
[0151] Step S33: Extract diffraction features from the set of Fresnel diffraction images to obtain a diffraction feature map;
[0152] Step S34: Evaluate the light field focusing of the stratified scene according to the polarization phase gradient spectrum to obtain a light field clarity map;
[0153] Step S35: Calculate the focal plane confidence for the diffraction feature map and the light field clarity map to obtain a focal plane confidence map.
[0154] In an embodiment of the present invention, the optimized depth map obtained in step S2 is received. Based on the depth information of the depth map, the three-dimensional scene is decomposed into multiple two-dimensional depth layers. First, a layering strategy is determined. A uniform layering strategy can be adopted, that is, layering is performed at fixed depth intervals. For example, the depth value range is from 0 to 1000 units, and layering is performed with every 10 units as a depth layer. An adaptive layering strategy can also be adopted, that is, layering is performed according to the change rate of the depth value. For regions with a large depth change, a smaller depth interval can be used, and for regions with a gentle depth change, a larger depth interval can be used. After determining the layering strategy, the optimized depth map is traversed. For each pixel in the depth map, its corresponding depth value is obtained. According to the depth value and the layering strategy, the depth layer to which the pixel belongs is determined. Pixels with the same depth layer are divided into the same depth layer. A data structure of a layered scene is created. This data structure contains multiple depth layers, and each depth layer corresponds to a two-dimensional image. The pixels belonging to the same depth layer are copied into the corresponding two-dimensional image. Pixels not in any depth layer are processed. For example, their depth values can be set to a default value, or they can be ignored. All depth layers are combined to obtain a layered scene. The layered scene contains multiple two-dimensional images, each image corresponds to a depth layer, and each image contains the scene information in that depth layer.
[0155] Obtain the optical system parameters, including the focal length (focal length, f), pixel pitch (pixel pitch, Δx, Δy), and aperture diameter (aperture diameter, D) of the camera. These parameters can be obtained from the camera's specification or the calibration process. Secondly, according to the Layered Scene, simulate the Fresnel diffraction of light waves at different depth layers. For each depth layer in the layered scene, perform the following operations. Construct the input light field. For each depth layer, calculate the complex amplitude of each pixel in that depth layer according to the pixel values of the depth map. The amplitude part of the complex amplitude represents the intensity of the light ray, and the phase part represents the phase of the light ray. For example, it can be assumed that the amplitude of the light ray is proportional to the pixel value, and the phase is related to the propagation distance of the light ray. For each depth layer, generate a two-dimensional complex amplitude distribution as the input light field. Perform light field padding on the input light field. Since the Fresnel diffraction calculation needs to process the entire space, it is necessary to pad the input light field to avoid edge effects. Methods such as zero-padding or mirror-padding can be used. The size of the padded light field is larger than the original light field. Perform a light field Fourier transform on the padded light field. Use the Fast Fourier Transform (FFT) algorithm to transform the padded light field from the spatial domain to the frequency domain. Obtain the frequency domain light field, which represents the frequency components of the light wave. Calculate the propagation function according to the optical system parameters and the frequency domain light field. The propagation function describes the changes of the light wave during propagation. For Fresnel diffraction, the propagation function can be calculated according to the Fresnel diffraction integral formula. The propagation function is related to the focal length, pixel pitch, and wavelength. Calculate the propagation function according to these parameters. Perform a frequency domain multiplication on the propagation function and the frequency domain light field. Multiply the propagation function and the frequency domain light field element by element. The multiplication result represents the frequency domain information of the light wave after propagation. Perform an inverse light field Fourier transform on the propagated frequency domain light field. Use the Inverse Fast Fourier Transform (IFFT) algorithm to transform the propagated frequency domain light field from the frequency domain to the spatial domain. Obtain the diffraction image. The diffraction image represents the light intensity distribution of the light wave after propagating to the focal plane. Perform the Fresnel diffraction simulation for each depth layer according to the diffraction image and the layered scene. Repeat the above process to perform the Fresnel diffraction simulation for each depth layer in the layered scene and generate a Fresnel Diffraction Image Set. The Fresnel Diffraction Image Set contains the diffraction images of different depth layers for subsequent feature extraction and focal plane prediction.
[0156] Receive the Fresnel Diffraction Image Set obtained in step S32, and extract diffraction features from it for evaluating the clarity of the image. For each diffraction image in the Fresnel Diffraction Image Set, perform the following operations. Select appropriate diffraction features. Diffraction features are used to describe the characteristics of the diffraction image and reflect the clarity of the image. Multiple diffraction features can be selected, for example, image contrast, edge sharpness, and frequency domain features. Calculate the contrast of the diffraction image. There are multiple methods to calculate the image contrast, for example, calculating local contrast or global contrast. Local contrast can use metrics such as standard deviation or root mean square contrast. Global contrast can use the difference between the maximum pixel value and the minimum pixel value. Calculate the edge sharpness of the diffraction image. Edge sharpness reflects the clarity of the edges in the image. Edge detection operators such as Sobel operator and Canny operator can be used to calculate the intensity of the edges. Calculate the frequency domain features of the diffraction image. Perform Fourier transform on the diffraction image to obtain the frequency domain representation. Extract the energy of the high-frequency components as the frequency domain feature. High-frequency components reflect the detailed information of the image. Combine the calculated diffraction features to form a diffraction feature vector. The diffraction feature vector contains different types of diffraction features and is used to comprehensively describe the characteristics of the diffraction image. Repeat the above operations for all diffraction images in the Fresnel Diffraction Image Set. Combine the diffraction feature vectors of all diffraction images to form a Diffraction Feature Map. The Diffraction Feature Map contains the diffraction features of each depth layer, and these features are used for subsequent calculation of the focal plane confidence.
[0157] Using the Polarization Phase Gradient Spectrum (PPGS) and the Layered Scene, the light field focusing of each depth layer is evaluated to determine whether the depth layer is in focus. For each depth layer in the layered scene, the following operations are performed: According to the polarization phase gradient spectrum, the light field information of the depth layer is reconstructed. Using PPGS and combining with the depth information of the depth layer, the virtual view image of the depth layer can be reconstructed. Light field refocusing algorithms can be used, for example, integral refocusing or disparity-based refocusing. The light field refocusing algorithm can generate images at different viewpoints according to the depth information. The sharpness of the refocused image is evaluated. The sharpness evaluation is used to evaluate the sharpness of the refocused image. Multiple sharpness evaluation metrics can be used, such as image gradient, variance, or frequency domain energy. To calculate the gradient of the image, the Sobel operator or the Laplacian operator can be used. To calculate the variance of the image, the larger the variance, the higher the contrast of the image. To calculate the frequency domain energy of the image, the energy of the high-frequency components is extracted. The calculated sharpness metric is used as the light field focusing evaluation result of the depth layer. The above operations are repeated for all depth layers in the layered scene. The light field focusing evaluation results of all depth layers are combined into a Light Field Sharpness Map (LFSM). The light field sharpness map contains the sharpness evaluation results of each depth layer and is used for subsequent focal plane confidence calculation.
[0158] The Diffraction Feature Map (DFM) and the Light Field Sharpness Map (LFSM) are fused to calculate the focal plane confidence of each depth layer. For each depth layer, the following operations are performed: Obtain the diffraction feature value of the depth layer in the DFM and the sharpness value of the depth layer in the LFSM. The DFM and LFSM are fused. A weighted fusion method can be used to perform a weighted average of the DFM and LFSM. Different weights are assigned to the DFM and LFSM, and the weights can be adjusted according to the scene characteristics. For example, in a low-contrast scene, the weight of the DFM can be increased to compensate for the lack of light field information. A neural network can also be used to learn the relationship between the DFM and LFSM and predict the focal plane confidence. The calculated confidence is normalized. The confidence value is normalized to the range of 0 to 1 for subsequent processing. The focal plane confidences of all depth layers are combined to form a Focal Plane Confidence Map (FPCM). The focal plane confidence map is a two-dimensional map, and each value represents the focal plane confidence at different scene depths. The focal plane confidence map is used to determine the optimal focal plane position.
[0159] Preferably, step S32 includes the following steps:
[0160] Step S321: Construct an input light field according to the layered scene to obtain the input light field;
[0161] Step S322: Perform light field filling on the input light field to obtain the filled light field;
[0162] Step S323: Perform a light field Fourier transform on the filled light field to obtain the frequency-domain light field;
[0163] Step S324: Calculate the propagation function according to the optical system parameters and the frequency-domain light field to obtain the propagation function;
[0164] Step S325: Perform a frequency-domain multiplication on the propagation function and the frequency-domain light field to obtain the propagated frequency-domain light field;
[0165] Step S326: Perform an inverse light field Fourier transform on the propagated frequency-domain light field to obtain the diffraction image;
[0166] Step S327: Perform Fresnel diffraction simulation for each depth layer according to the diffraction image and the layered scene to obtain a set of Fresnel diffraction images.
[0167] In the embodiment of the present invention, the layered scene (Layered Scene) obtained in step S31 is received, and the input light field of each depth layer is constructed according to the depth information in the layered scene. For each depth layer in the layered scene, the following operations are performed: According to the depth value corresponding to the depth layer, the coordinates and the corresponding complex amplitude values of each pixel in the depth layer are obtained. The complex amplitude represents the amplitude and phase information of the light wave. The amplitude is usually related to the gray value or color value of the image, and the phase is related to the propagation distance of the light wave. A two-dimensional complex matrix is constructed, and the size of the matrix is the same as the image size of the depth layer. The complex amplitude value of each pixel is filled into the corresponding pixel position in the complex matrix. For example, if the gray value of the pixel is g and the depth value is z, the complex amplitude of the pixel can be defined as: A * exp(j * k * z), where A is the amplitude, j is the imaginary unit, and k is the wave number. A is set to be proportional to g, and k is set to be a constant. The constructed complex matrix is used as the input light field of the depth layer. For all depth layers in the layered scene, the above operations are repeated to obtain the input light field of each depth layer. The input light field characterizes the complex amplitude distribution of the light wave on the depth layer.
[0168] Receive the input optical field obtained in step S321, and fill the input optical field to avoid edge effects in subsequent Fourier transform and diffraction simulation. Due to the periodicity of the Fourier transform and the influence of boundary conditions in Fresnel diffraction simulation, it is necessary to fill the input optical field. The filling method can be, for example, zero-padding. For each input optical field, create a new two-dimensional complex matrix whose size is larger than that of the input optical field. For example, if the size of the input optical field is M x N, the size of the filled optical field can be set to 2M x 2N. Copy the input optical field to the center position of the filled optical field. Set the remaining pixel values of the filled optical field to 0. Mirror-padding can also be used. For each input optical field, mirror-copy the edge pixels of the input optical field and fill them into the filled optical field. Mirror-padding can reduce edge effects and maintain the continuity of the image. The size of the filled optical field is the same as that of the input optical field. Use zero-padding or mirror-padding to fill the input optical field to obtain the filled optical field (PaddingLight Field). The size of the filled optical field is larger than that of the input optical field and is used for subsequent Fourier transform and diffraction simulation.
[0169] Receive the filled optical field obtained in step S322, and use Fourier transform to convert it from the spatial domain to the frequency domain. For each filled optical field, use the two-dimensional fast Fourier transform (FFT) algorithm. The FFT algorithm can efficiently calculate the Fourier transform. The FFT algorithm converts each pixel of the filled optical field into a frequency component in the frequency domain. The result of the Fourier transform is a complex matrix, and the real part and the imaginary part of this matrix respectively represent different frequency components in the frequency domain. Calculate the Fourier transform for each pixel of the filled optical field. Store the result of the Fourier transform in a new two-dimensional complex matrix. The size of this complex matrix is the same as that of the filled optical field. Take this complex matrix as the frequency domain optical field (FrequencyDomain Light Field). The frequency domain optical field represents the distribution of light waves in the frequency domain.
[0170] Calculate the propagation function based on the optical system parameters and the frequency-domain light field. The propagation function describes the changes of light waves during propagation. Obtain the optical system parameters, including the focal length (focal length, f), pixel pitch (pixel pitch, Δx, Δy), and wavelength (wavelength, λ) of the camera. These parameters can be obtained from the camera's specifications or the calibration process. Calculate the propagation function. According to the Fresnel diffraction theory, the propagation function can be expressed as: H(fx, fy) = exp(-j * π * λ * z * (fx^2 + fy^2)), where fx and fy are the frequency-domain coordinates, and z is the propagation distance, which is the depth value of the depth layer. Calculate the frequency-domain coordinates fx and fy, fx = m / (M * Δx), fy = n / (N * Δy), where m and n are the indices of the frequency-domain coordinates, and M and N are the sizes of the padded light field. Calculate the propagation function based on the optical system parameters and the frequency-domain coordinates. For each pixel (fx, fy) in the frequency domain, calculate its corresponding propagation function value. The propagation function is a complex number, representing the phase change and amplitude attenuation of light waves during propagation. Store the calculated propagation function values in a two-dimensional complex matrix, whose size is the same as that of the frequency-domain light field. Take this complex matrix as the propagation function (Propagation Function). The propagation function describes the changes of light waves during propagation and is used for subsequent element-wise multiplication in the frequency domain.
[0171] Receive the frequency-domain light field obtained in step S323 and the propagation function obtained in step S324, and perform element-wise multiplication of the propagation function and the frequency-domain light field. For each pixel in the frequency-domain light field, multiply its corresponding propagation function value. For example, if the complex value of a certain pixel in the frequency-domain light field is A, and the complex value of the corresponding pixel in the propagation function is H, then the result after multiplication is A * H. Repeat the above operation for all pixels in the frequency-domain light field. Store the result after multiplication in a new two-dimensional complex matrix, whose size is the same as that of the frequency-domain light field and the propagation function. Take this complex matrix as the propagated frequency-domain light field (Propagated Frequency Domain Light Field). The propagated frequency-domain light field represents the frequency-domain distribution of light waves after propagation.
[0172] Receive the propagated frequency-domain optical field obtained in step S325, and use the inverse Fourier transform to convert it from the frequency domain to the spatial domain. For each propagated frequency-domain optical field, use the two-dimensional fast inverse Fourier transform (IFFT) algorithm. The IFFT algorithm can efficiently calculate the inverse Fourier transform. The IFFT algorithm converts each frequency component of the propagated frequency-domain optical field to a pixel in the spatial domain. Calculate the inverse Fourier transform for each pixel of the propagated frequency-domain optical field. Store the result of the inverse Fourier transform in a new two-dimensional complex matrix. The size of this complex matrix is the same as that of the propagated frequency-domain optical field. Calculate the diffraction image, take the square of the modulus of the result of the inverse Fourier transform. The square of the modulus represents the light intensity of the light wave. Use the square of the modulus result as the diffraction image (Diffraction Image). The diffraction image represents the light intensity distribution of the light wave after propagating to the focal plane.
[0173] According to the diffraction image obtained in step S326 and the layered scene, perform Fresnel diffraction simulation for each depth layer to obtain a set of Fresnel diffraction images. For each depth layer in the layered scene, repeat the above steps S321 to S326 to simulate the Fresnel diffraction after the light wave propagates through this depth layer. For each depth layer, obtain a diffraction image. Combine the diffraction images of all depth layers into a set of Fresnel diffraction images (Fresnel Diffraction ImageSet). The set of Fresnel diffraction images contains multiple diffraction images, and each diffraction image corresponds to a depth layer.
[0174] Preferably, step S4 includes the following steps:
[0175] Step S41: Determine the optimal focal plane position for the focal plane confidence map to obtain the optimal focal plane depth;
[0176] Step S42: Obtain the camera parameters; calculate the lens displacement according to the optimal focal plane depth and the camera parameters to obtain the lens displacement amount;
[0177] Step S43: Correct the lens displacement according to the lens displacement amount and the optical system parameters to obtain the corrected lens displacement;
[0178] Step S44: Generate a lens control instruction for the corrected lens displacement to obtain the lens control instruction.
[0179] In the embodiment of the present invention, the Focal Plane Confidence Map (FPCM) obtained in step S3 is received to determine the optimal focal plane position. The FPCM is a two-dimensional image, and each pixel of it represents the focusing confidence at different depths of field. First, find the pixel with the highest confidence in the FPCM. Traverse all the pixels in the FPCM to find the pixel with the highest value. Record the row index and column index of this pixel. According to the row index and column index of this pixel, and the depth range corresponding to the FPCM, calculate the optimal focal plane depth. For example, if the depth range of the FPCM is from 0 to 1000 units, and the row index of the FPCM corresponds to the depth value, linear interpolation or a look-up table can be used to determine the optimal focal plane depth. The optimal focal plane depth corresponds to the depth value with the highest confidence in the FPCM. Take this depth value as the Optimal Focal Plane Depth (OFPD). The OFPD represents the clearest focal plane position in the scene.
[0180] Obtain the relevant parameters of the camera, and calculate the distance that the lens needs to move according to the optimal focal plane depth (OFPD). First, obtain the focal length (f) of the camera. The focal length can be obtained through the camera's specification or calibration process. Obtain the initial lens position of the camera. The initial lens position of the camera represents the current distance from the lens to the image sensor. Obtain the sensor size of the camera. The sensor size is used to calculate the image distance. According to the OFPD and the camera parameters, calculate the distance that the lens needs to move. Use the thin lens imaging formula: 1 / f = 1 / u + 1 / v, where f is the focal length, u is the object distance (OFPD), and v is the image distance. Calculate the image distance v according to the OFPD and the focal length. Lens displacement = v - initial image distance. The lens displacement amount represents the distance and direction that the lens needs to move. If the lens displacement amount is positive, it means the lens needs to move away from the sensor. If the lens displacement amount is negative, it means the lens needs to move closer to the sensor. Take the calculated lens displacement amount (LensDisplacement, LD) as the instruction for the lens to move.
[0181] According to the lens displacement (LD) and optical system parameters, the lens displacement is corrected to improve the focusing accuracy. Obtain the optical system parameters, including the aperture value (f / #), zoom information, and other optical element parameters. The aperture value affects the depth of field, and the zoom information affects the focal length. Correct the lens displacement according to the aperture value. If a smaller depth of field (large aperture) is required, more precise focusing is needed, and LD needs to be fine-tuned. For example, a small offset can be made to LD according to the aperture value. The larger the aperture value, the larger the offset. Correct the lens displacement according to the focal length. If it is a zoom lens, the lens movement amount needs to be adjusted according to the current focal length. For example, LD can be scaled using a function of the focal length. Correct the lens displacement according to other optical element parameters. Consider other optical system parameters, such as lens distortion, chromatic aberration, etc. Correct LD according to these parameters. Take the corrected lens displacement (CLD) as the final lens movement command.
[0182] According to the corrected lens displacement (CLD), generate a lens control command to drive the lens motor to move to the target position. Convert CLD into a control signal that the lens motor can understand. The lens motor usually uses a stepper motor or a DC motor. Convert CLD into the number of steps of the stepper motor. The number of steps of the stepper motor is proportional to the distance the lens moves. It is necessary to convert CLD into the corresponding number of steps according to the stepping accuracy of the lens motor. Convert CLD into a voltage signal. If a DC motor is used, CLD needs to be converted into the corresponding voltage signal to control the speed and direction of the motor. Send the command through the lens control interface. The lens control interface usually uses serial communication protocols such as I2C, UART, etc. Send the control signal to the lens motor through the lens control interface. The control signal includes: the direction of lens movement (forward or reverse), the number of steps or voltage value of movement. Generate a lens control command (LCC). LCC contains information such as the direction of lens movement, the distance of movement (number of steps or voltage value), etc., and is used to control the lens motor to move to the target position.
[0183] Preferably, step S5 includes the following steps:
[0184] Step S51: Move the lens according to the lens control command and perform current image acquisition to obtain the current image;
[0185] Step S52: Evaluate the clarity of the current image to obtain the image clarity value;
[0186] Step S53: Use a preset clarity threshold to determine the focusing state based on the image sharpness value, and obtain the focusing state;
[0187] Step S54: Perform focusing iterative optimization based on the focusing state, the image sharpness value, the focal plane confidence map, and the corrected lens displacement to obtain an optimized lens control instruction.
[0188] In the embodiment of the present invention, receive the lens control instruction (Lens Control Command, LCC) generated in step S4, and control the lens motor to move to the target position. Control the lens motor according to the LCC. The LCC contains information about the direction and distance (number of steps or voltage value) of the lens movement. Send the LCC to the lens motor driver through a lens control interface, such as I2C or UART. The lens motor driver drives the lens motor to move according to the LCC. Wait for the lens motor to move to the target position. After the lens motor movement is completed, perform current image acquisition. Use a light field camera to acquire the current image. The exposure time and gain settings during image acquisition can refer to the previous settings or be adjusted according to the scene brightness. Take the acquired image as the current image (Current Image, CI). The CI represents the scene image at the current lens position.
[0189] Receive the current image (Current Image, CI) obtained in step S51, and perform image sharpness evaluation on it to determine the clarity of the image. Select a suitable sharpness evaluation index. Multiple sharpness evaluation indexes can be used, for example: image gradient, frequency domain energy, or sharpness evaluation based on a pre-trained neural network. Calculate the image gradient. The image gradient can reflect the edge and detail information of the image. The Sobel operator, Prewitt operator, or Laplacian operator can be used to calculate the horizontal and vertical gradients of the image. Calculate the modulus or sum of squares of the gradient image to obtain the image gradient value. Calculate the frequency domain energy. Perform a Fourier transform on the image to obtain the frequency domain representation. Extract the energy of the high-frequency components as the sharpness index of the image. The high-frequency components reflect the detail information of the image. Use a pre-trained neural network. Use a pre-trained neural network, input the CI, and output a sharpness value. This neural network needs to be trained using a large number of clear images and blurred images. Calculate the sharpness value of the CI. According to the selected sharpness evaluation index, calculate the sharpness value of the CI. Take the calculated sharpness value as the image sharpness value (Image Sharpness Value, ISV). The ISV represents the clarity of the CI.
[0190] Receive the image sharpness value (ISV) obtained in step S52, compare it with a preset sharpness threshold, and determine the focusing state. Set a sharpness threshold (ST). ST represents the minimum requirement for image sharpness. ST can be adjusted according to the actual application scenario and the performance of the camera. Compare ISV with ST. If ISV is greater than or equal to ST, it is considered that the image is in focus and the focusing state is "focus success". If ISV is less than ST, it is considered that the image is not in focus and the focusing state is "focus failure". Generate a focus status (FS). FS represents the status information indicating whether the focusing is successful.
[0191] According to the focus status (FS), image sharpness value (ISV), focal plane confidence map (FPCM), and corrected lens displacement (CLD), perform focus iteration optimization to further improve the focusing accuracy. If FS indicates focus failure, focus iteration optimization is required. Calculate the focal plane offset according to the difference between ISV and the expected value. The expected value can be set to the maximum ISV or a threshold close to the maximum ISV. Calculate the difference between ISV and the expected value. For example, the absolute value of the difference can be calculated. Determine the search direction according to the confidence distribution in FPCM. The search direction should point to the direction with higher confidence. Calculate the distance and direction of lens movement according to the focal plane offset and the search direction. Fine-tune CLD according to the focal plane offset. The fine-tuned CLD represents the distance and direction that the lens needs to move. Generate a new lens control command. Generate a new lens control command (RLCC) according to the fine-tuned CLD. RLCC contains the direction and distance information of lens movement for controlling the movement of the lens motor. Repeat steps S51 to S53, use RLCC to control the lens movement, acquire the image again, and perform sharpness evaluation. Until FS indicates focus success or reaches the preset number of iterations. If FS indicates focus success, the focus iteration optimization is completed. If the preset number of iterations is reached but FS still indicates focus failure, stop the iteration optimization and return the current lens position.
[0192] Preferably, the present invention further provides an optical lens automatic focusing system for performing the optical lens automatic focusing method as described above. The optical lens automatic focusing system includes:
[0193] A light field feature extraction module, which is used to collect a polarized light field through a light field camera to obtain a set of polarized multi-view images; and extract light field features from the set of polarized multi-view images to obtain a polarized phase gradient spectrum;
[0194] A depth information reconstruction module, which is used to construct a polarization-aware depth reconstruction network based on the polarized phase gradient spectrum to obtain a preliminary depth map; perform ray tracing simulation processing on the preliminary depth map to obtain an enhanced depth map; and optimize the preliminary depth map and the enhanced depth map to obtain an optimized depth map;
[0195] A focal plane prediction module, which is used to obtain optical system parameters; perform scene stratification on the optimized depth map to obtain a stratified scene; perform Fresnel diffraction simulation based on the optical system parameters and the stratified scene to obtain a set of Fresnel diffraction images; and perform focal plane prediction on the set of Fresnel diffraction images to obtain a focal plane confidence map;
[0196] A lens control instruction generation module, which is used to calculate the lens displacement according to the focal plane confidence map to obtain a lens displacement amount; correct the lens displacement according to the lens displacement amount, and generate a lens control instruction to obtain a corrected lens displacement and a lens control instruction;
[0197] A focusing confirmation and optimization module, which is used to move the lens according to the lens control instruction and judge the focusing state to obtain a focusing state; and perform focusing iterative optimization according to the focusing state to obtain an optimized lens control instruction.
[0198] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention.
[0199] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. An automatic focusing method for an optical lens, characterized in that, It includes the following steps: Step S1: Perform polarized light field acquisition through a light field camera to obtain a set of polarized multi-view images; extract light field features from the set of polarized multi-view images to obtain a polarized phase gradient spectrum; Step S2: Construct a polarization-aware depth reconstruction network based on the polarized phase gradient spectrum to obtain a preliminary depth map; perform ray tracing simulation processing on the preliminary depth map to obtain an enhanced depth map; optimize the preliminary depth map and the enhanced depth map to obtain an optimized depth map; Step S3: Obtain optical system parameters; perform scene layering on the optimized depth map to obtain a layered scene; Perform Fresnel diffraction simulation based on the optical system parameters and the layered scene to obtain a set of Fresnel diffraction images; perform focal plane prediction on the set of Fresnel diffraction images to obtain a focal plane confidence map; wherein, Step S3 includes the following steps: Step S31: Perform scene layering on the optimized depth map to obtain a layered scene; Step S32: Obtain optical system parameters, perform Fresnel diffraction simulation based on the layered scene and the optical system parameters to obtain a set of Fresnel diffraction images; wherein, Step S32 includes the following steps: Step S321: Construct an input light field based on the layered scene to obtain an input light field; Step S322: Perform light field filling on the input light field to obtain a filled light field; Step S323: Perform light field Fourier transform on the filled light field to obtain a frequency-domain light field; Step S324: Calculate a propagation function based on the optical system parameters and the frequency-domain light field to obtain a propagation function; Step S325: Perform frequency-domain multiplication on the propagation function and the frequency-domain light field to obtain a propagated frequency-domain light field; Step S326: Perform inverse light field Fourier transform on the propagated frequency-domain light field to obtain a diffraction image; Step S327: Perform Fresnel diffraction simulation for each depth layer based on the diffraction image and the layered scene to obtain a set of Fresnel diffraction images; Step S33: Extract diffraction features from the set of Fresnel diffraction images to obtain a diffraction feature map; Step S34: Perform light field focusing evaluation on the layered scene based on the polarized phase gradient spectrum to obtain a light field clarity map; Step S35: Calculate the focal plane confidence for the diffraction feature map and the light field clarity map to obtain a focal plane confidence map; Step S4: Calculate the lens displacement amount based on the focal plane confidence map; perform lens displacement correction based on the lens displacement amount, and generate a lens control instruction to obtain the corrected lens displacement and the lens control instruction; Step S5: Move the lens according to the lens control instruction, and judge the focusing state to obtain the focusing state; perform focusing iterative optimization based on the focusing state to obtain an optimized lens control instruction.
2. The automatic focusing method of an optical lens according to claim 1, wherein Step S1 includes the following steps: Step S11: Perform polarized light field acquisition through a light field camera to obtain a set of polarized multi-view images; Step S12: Perform Fourier transform on the view point images of the set of polarized multi-view images to obtain a Fourier transform spectrum; Step S13: Calculate the phase gradient of the Fourier transform spectrum to obtain a phase gradient map; Step S14: Calculate the polarized phase gradient of the phase gradient map to obtain a polarized phase gradient map; Step S15: Generate multi-dimensional light field eigenvectors for the polarization phase gradient map to obtain a polarization phase gradient spectrum.
3. The automatic focusing method of an optical lens according to claim 2, characterized in that, Step S11 includes the following steps: Step S111: Place an angle-calibrated linear polarizer in front of the light field camera, perform synchronous control of the light field camera and the polarizer to obtain synchronous control parameters; Step S112: Collect multi-polarization-direction light field data according to the synchronous control parameters to obtain an original light field image set; Step S113: Correct the influence of the polarizer on the light field for the original light field image set to obtain a corrected light field image set; Step S114: Store and label the data of the corrected light field image set to obtain a polarization multi-view image set.
4. The optical lens autofocus method according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Preprocess the polarization phase gradient spectrum to obtain a preprocessed polarization phase gradient spectrum; Step S22: Construct a polarization perception depth reconstruction network according to the preprocessed polarization phase gradient spectrum to obtain a preliminary depth map; Step S23: Perform ray tracing simulation according to the preliminary depth map and the preprocessed polarization phase gradient spectrum, and perform training data enhancement to obtain an enhanced depth map; Step S24: Optimize the preliminary depth map and the enhanced depth map to obtain an optimized depth map.
5. The automatic focusing method of an optical lens according to claim 4, wherein, Step S22 includes the following steps: Step S221: Extract initial convolutional layer features from the preprocessed polarization phase gradient spectrum to obtain a preliminary feature map; extract downsampled layer features from the preliminary feature map to obtain a downsampled feature map; Step S222: Separate the polarization channels of the downsampled feature map to obtain separated polarization features; Step S223: Introduce channel attention to the separated polarization features to obtain polarization features with channel attention; Step S224: Fuse the polarization features with channel attention to obtain a fused feature map; Step S225: Extract features based on a residual block encoder from the fused feature map to obtain encoder output features; Step S226: Perform feature mapping based on a decoder on the encoder output features to obtain decoder output features; Step S227: Output a depth map from the decoder output features to obtain a preliminary depth map.
6. The automatic focusing method of an optical lens according to claim 1, wherein Step S4 includes the following steps: Step S41: Determine the best focal plane position for the focal plane confidence map to obtain the best focal plane depth; Step S42: Obtain camera parameters; calculate the lens displacement according to the best focal plane depth and the camera parameters to obtain a lens displacement amount; Step S43: Correct the lens displacement according to the lens displacement amount and the optical system parameters to obtain a corrected lens displacement; Step S44: Generate a lens control command for the corrected lens displacement to obtain a lens control command.
7. The optical lens autofocus method according to claim 1, characterized in that, Step S5 includes the following steps: Step S51: Move the lens according to the lens control command, and collect the current image to obtain the current image; Step S52: Evaluate the clarity of the current image to obtain an image clarity value; Step S53: Use a preset clarity threshold to judge the focusing state of the image clarity value to obtain a focusing state; Step S54: Based on the focusing state, the image sharpness value, the focal plane confidence map, and the corrected lens displacement, perform focusing iterative optimization to obtain an optimized lens control instruction.
8. An automatic focusing system for an optical lens, characterized in that, An optical lens autofocus system for executing the optical lens autofocus method according to claim 1, the system comprising: A light field feature extraction module, configured to collect a polarized light field through a light field camera to obtain a set of polarized multi-view images; perform light field feature extraction on the set of polarized multi-view images to obtain a polarized phase gradient spectrum; A depth information reconstruction module, configured to construct a polarization-aware depth reconstruction network based on the polarized phase gradient spectrum to obtain a preliminary depth map; perform ray tracing simulation processing on the preliminary depth map to obtain an enhanced depth map; optimize the preliminary depth map and the enhanced depth map to obtain an optimized depth map; A focal plane prediction module, configured to obtain optical system parameters; perform scene stratification on the optimized depth map to obtain a stratified scene; perform Fresnel diffraction simulation based on the optical system parameters and the stratified scene to obtain a set of Fresnel diffraction images; perform focal plane prediction on the set of Fresnel diffraction images to obtain a focal plane confidence map; A lens control instruction generation module, configured to calculate a lens displacement based on the focal plane confidence map to obtain a lens displacement amount; correct the lens displacement according to the lens displacement amount, and generate a lens control instruction to obtain a corrected lens displacement and a lens control instruction; A focusing confirmation and optimization module, configured to move the lens according to the lens control instruction, and determine the focusing state to obtain a focusing state; perform focusing iterative optimization based on the focusing state to obtain an optimized lens control instruction.
Citation Information
Patent Citations
Automatic focusing method for focal plane of imaging ellipsometer
CN112969026A
Autofocus device for microscopy
US20100033811A1