Liquid lens focal stack imaging method, system, and electronic device
By employing a liquid lens-based focal stacking imaging method, and utilizing the image at its lowest driving voltage for focal stacking alignment and feature matching algorithms, the problem of low imaging accuracy in focal stacking imaging technology is solved. This enables the generation of aberration-free constant magnification focal stacking image sequences and full-focus images.
Patent Information
- Application Number
- CN202411424894.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-10-12
AI Technical Summary
Existing focus stacking imaging technology suffers from low imaging accuracy, mainly due to changes in magnification and field of view caused by variations in the distance between the lens and the CMOS image sensor, resulting in misalignment and pixel loss between image frames in the focus stacking image sequence.
A liquid lens focal stack imaging method is adopted. By controlling the driving voltage of the liquid lens, an image is formed on the CMOS array. The focal stack is aligned based on the image when the driving voltage is the lowest, and an aberration-free constant magnification focal stack image sequence is obtained. The image is magnified to the same imaging magnification using a feature matching algorithm. A full-focus image is generated by combining focus state evaluation and full-focus fusion.
It effectively reduces stitching misalignment and pixel loss in fully focused images generated from focus stacked images, improves imaging accuracy, and enables the generation of aberration-free constant magnification focus stacked image sequences.
Smart Images

Figure CN119105119B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of optical imaging detection, in particular to a liquid lens focal stack imaging method, system and electronic equipment. BACKGROUND
[0002] In recent years, with the significant development of deep learning technology, focal stack imaging technology has been widely used in many fields such as depth expansion and depth estimation. However, how to quickly and real-time collect focal stack image sequences with obvious defocus and focus features is still a challenge. In previous research work, optical zoom systems such as mechanical focusing cameras, light field cameras or telecentric cameras are often used to collect focal stack image sequences. When using a mechanical focusing camera to collect focal stack images, the imaging magnification and field angle of the captured focal stack image sequence will change due to the change of the distance between the lens and the CMOS image sensor, which is called lens breathing effect. This change in imaging magnification causes misalignment between pixel points in different image frames in the stack, and if the focal stack image sequence is not aligned, the all-in-focus image generated from the focal stack image will have obvious stitching misalignment and pixel loss, resulting in low image accuracy. SUMMARY
[0003] Therefore, it is necessary to provide a liquid lens focal stack imaging method, system and electronic equipment to solve the problem of low imaging accuracy in the existing focal stack imaging technology.
[0004] To solve the above problems, on the one hand, the present application provides a liquid lens focal stack imaging method, which is applied to a focal stack imaging device, the device comprising: a power supply module, and a coaxial fixed focus lens, liquid lens and CMOS array; the method comprising:
[0005] controlling the power supply module to provide a driving voltage for the liquid lens to form an image on the CMOS array;
[0006] acquiring the image collected by the CMOS array, and performing focal stack alignment on all collected images based on the image when the driving voltage is lowest to obtain an image sequence of constant magnification focal stack without aberration;
[0007] evaluating the focus state of the image sequence, and performing all-in-focus fusion on the image sequence based on the focus state evaluation result to obtain an all-in-focus image;
[0008] determining the object distance of the collected image, generating a depth map based on the object distance and the all-in-focus image, and generating a target image based on the all-in-focus image and the depth map.
[0009] In a possible implementation, the control power supply module provides a driving voltage for the liquid lens, comprising:
[0010] Based on the focal depth of the focus stack imaging device, the target depth of field and the image distance, the driving voltage required by the liquid lens is determined;
[0011] Based on the driving voltage required by the liquid lens, the power supply module provides a driving voltage to both ends of the liquid lens to form an image of the target depth of field on the CMOS array.
[0012] In a possible implementation, all images collected are aligned based on the image when the driving voltage is lowest to obtain an image sequence of anastigmatic constant magnification focus stack, comprising:
[0013] The image collected by the CMOS array when the driving voltage is lowest is taken as a reference image, and the images of the focus stack of the CMOS array are magnified to images with the same imaging magnification as the reference image based on a feature matching algorithm to obtain an anastigmatic constant magnification focus stack image sequence;
[0014] Wherein, the focus stack comprises images focused at different spatial depths collected by the CMOS array.
[0015] In a possible implementation, the images of the focus stack of the CMOS array are magnified to images with the same imaging magnification as the reference image based on a feature matching algorithm to obtain an anastigmatic constant magnification focus stack image sequence, comprising:
[0016] The feature points of two adjacent images in the focus stack of the CMOS array are matched to determine the transformation relationship between the two adjacent images;
[0017] Based on the transformation relationship between all two images between each image in the focus stack of the CMOS array and the reference image, each image is magnified to an image with the same imaging magnification as the reference image to obtain an anastigmatic constant magnification focus stack image sequence.
[0018] In a possible implementation, the focus state of the image sequence is evaluated, comprising:
[0019] The sharpness of the pixel points of each image in the image sequence is evaluated, and the focus state evaluation result is obtained based on the evaluation result of the sharpness of the pixel points.
[0020] In a possible implementation, the sharpness of each pixel point in each frame of the image sequence is evaluated, and a focus state evaluation result is obtained based on the evaluation result of the sharpness of the pixel point, including:
[0021] Each frame of the image sequence is converted into a gray-scale image.
[0022] The focus degree and the sharpness in a window of a set size of the gray-scale image are calculated.
[0023] Based on the focus degree and the sharpness in the window of the set size of the gray-scale image, an evaluation result of the sharpness of the pixel point is obtained, and a focus state evaluation result is obtained based on the evaluation result of the sharpness of the pixel point.
[0024] In a possible implementation, the image sequence is full-focus fused based on the focus state evaluation result, and a full-focus image is obtained, including:
[0025] Based on the focus state evaluation result, the most sharp pixel point of each frame of the image sequence is spliced into a full-focus image.
[0026] In a possible implementation, a depth map is generated based on the object distance and the full-focus image, including:
[0027] Two-dimensional features of each frame of the image are extracted, and a focal point body is formed based on the two-dimensional features; the focal point body is differentiated based on the object distance, and a differential focal point body is obtained;
[0028] Features of the differential focal point body are extracted, and a probability of a pixel point in an optimal focus state is predicted based on the extracted features;
[0029] Based on the predicted probability of the pixel point in the optimal focus state, a probability regression is performed, and a depth map reflecting spatial depth information of the image is generated in combination with a focal length of each frame of the image.
[0030] On the other hand, the application further provides an electronic device, including a memory and a processor, wherein,
[0031] The memory is configured to store a program.
[0032] The processor is coupled to the memory and is configured to execute the program stored in the memory, so as to implement the steps of the liquid lens focal point stack imaging method according to any one of the above.
[0033] In another aspect, the present application also provides a liquid lens focal stack imaging system, comprising: a power supply module, and a coaxial fixed-focus lens, a liquid lens and a CMOS array, and the electronic device described above, the electronic device is electrically connected with the power supply module, and is in communication connection with the CMOS array.
[0034] The beneficial effects of the above implementation manner are: the liquid lens focal stack imaging method, system and electronic device provided by the present application, by acquiring the image collected by the CMOS array, and based on the image when the driving voltage is lowest, all the collected images are aligned for focal stack, and the image sequence of constant magnification focal stack without aberration is obtained; when the driving voltage of the liquid lens is lowest, the imaging system focuses on the object farthest away, and has the largest imaging magnification, therefore, the image collected when the driving voltage is lowest is selected as the reference image, all the collected images are enlarged to have the same imaging magnification as the reference image, and then the image sequence of constant magnification focal stack without aberration is obtained, thereby reducing the obvious splicing misplacement and pixel loss of the all-focus image generated by the focal stack image, and solving the technical problem of low imaging accuracy existing in the existing focal stack imaging technology. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0036] Figure 1 The flow chart of one embodiment of the liquid lens focal stack imaging method provided by the present application;
[0037] Figure 2 The liquid lens focal stack acquisition method based on the present application;
[0038] Figure 3 The ZEMAX optical path simulation model under the action of the driving voltage of 80V and 55V provided by the present application;
[0039] Figure 4 The schematic diagram of the all-focus image synthesis provided by the present application;
[0040] Figure 5 The depth of field diagram of the optical imaging system provided by the present application;
[0041] Figure 6 The depth of field-image distance relationship curve diagram of the focal stack acquisition camera provided by the present application;
[0042] Figure 7 A schematic diagram of a camera driving voltage-focal length relationship curve provided by the focus stack acquisition of the present application;
[0043] Figure 8 A schematic diagram of a camera driving voltage-object distance relationship curve provided by the focus stack acquisition of the present application;
[0044] Figure 9 A schematic diagram of a depth estimation module based on the focal length electric addressing characteristics of the liquid lens provided by the present application;
[0045] Figure 10 A schematic diagram of an embodiment structure of an electronic device provided by the present application. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person skilled in the art without creative labor fall within the protection scope of the present application.
[0047] In the description of the embodiments of the present application, the meaning of "a plurality of" is two or more, unless otherwise specified.
[0048] In the embodiments of the present application, the terms "comprising" and "having" and any variations thereof are intended to cover the inclusions that are not exclusive, for example, a process, method, device, product or equipment comprising a series of steps or modules does not have to be limited to the clearly listed steps or modules, but can include other steps or modules that are not clearly listed or inherent to these processes, methods, products or equipment.
[0049] The naming or numbering of the steps appearing in the embodiments of the present application does not mean that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The flow steps that have been named or numbered can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.
[0050] Reference to "an embodiment" in this document means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears at various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily independent or alternative embodiments to other embodiments. It is explicitly and implicitly understood by a person skilled in the art that the embodiments described herein can be combined with other embodiments.
[0051] The application provides a focal point stack imaging method, system and electronic equipment of a liquid lens.
[0052] The application provides a focal point stack imaging method of a liquid lens, which is applied to a focal point stack imaging device. Figure 1 As shown in the figure, the method comprises:
[0053] S101, controlling the power supply module to provide a driving voltage for the liquid lens to form an image on the CMOS array;
[0054] S102, acquiring the image collected by the CMOS array, and performing focal point stack alignment on all the collected images based on the image when the driving voltage is lowest to obtain an image sequence of constant magnification focal point stack without aberration;
[0055] S103, performing focus state evaluation on the image sequence, and performing all-focus fusion on the image sequence based on the focus state evaluation result to obtain an all-focus image;
[0056] S104, determining the object distance of the collected image, generating a depth map based on the object distance and the all-focus image, and generating a target image based on the all-focus image and the depth map.
[0057] It can be understood that the CMOS array is a CMOS image sensor array. When focal point stack collection is performed using a traditional optical zoom system, the distance between the CMOS image sensor and the glass lens needs to be adjusted through mechanical displacement. This mechanical movement usually needs to use micro-electro-mechanical structures such as servo motors, piezoelectric motors and voice coil motors. The lens movement realized by the micro-electro-mechanical structure usually has large sequential traversal and inertia, so it is difficult to achieve jump interval sampling and dense sampling of the object space scene. Moreover, the mechanical movement of the lens is difficult to ensure complete parallelism with the optical axis, which may cause lateral drift misalignment and rotation misalignment between focal point stack images. Moreover, the mechanical movement of the glass lens will cause a relatively obvious change in the back focal length of the entire imaging system, thereby causing significant differences in the imaging magnification of different images in the focal point stack.
[0058] In some embodiments, the control of the power supply module to provide a driving voltage for the liquid lens comprises:
[0059] determining the required driving voltage of the liquid lens based on the focal depth, the target depth of field and the image distance of the focal point stack imaging device;
[0060] Based on the driving voltage required by the liquid lens, the power supply module provides a driving voltage to both ends of the liquid lens to form an image of a target depth of field on the CMOS array.
[0061] It can be understood that for a traditional optical imaging system, the light emitted by an object point is refracted by a lens to form a circular diffused spot on a CMOS image sensor, and only when the entire diffused spot is completely within a photosensitive pixel, the image of the object point is clear. Although part of the object point is not completely focused on the CMOS image sensor, the diffused spot formed by the lens is still completely within a pixel, and thus the image is still clear. These object points can all be clearly imaged, and this distance is the depth of field.
[0062] The target depth of field can be determined based on the focal depth and the image distance of the focal stack imaging device and the driving voltage required by the liquid lens, and conversely, the depth of field of the image formed on the CMOS array can be determined based on the driving voltage.
[0063] In some embodiments, all images collected based on the image when the driving voltage is lowest are aligned for focal stack to obtain an image sequence of anastigmatic constant magnification focal stack, including:
[0064] The image collected by the CMOS array when the driving voltage is lowest is taken as a reference image, and the images of the focal stack of the CMOS array are enlarged to have the same imaging magnification as the reference image based on a feature matching algorithm to obtain an anastigmatic constant magnification focal stack image sequence.
[0065] The focal stack includes images collected by the CMOS array focused at different spatial depths.
[0066] It can be understood that the present application proposes a focal stack alignment algorithm for the problem of changing imaging magnification. When the driving voltage of the liquid lens is lowest, the imaging system focuses on the farthest object, and has the largest imaging magnification. Therefore, the image collected when the driving voltage is lowest is selected as a reference image, and the remaining images in the focal stack are aligned with the reference image and enlarged to have the same imaging magnification as the reference image through a feature point matching algorithm, and then an anastigmatic constant magnification focal stack image sequence is obtained.
[0067] In some embodiments, the images of the focal stack of the CMOS array are enlarged to have the same imaging magnification as the reference image based on a feature matching algorithm to obtain an anastigmatic constant magnification focal stack image sequence, including:
[0068] Feature points of two adjacent images in the focal stack of the CMOS array are matched to determine the transformation relationship between the two adjacent images.
[0069] Based on the transformation relationship between each frame image in the focal stack of the CMOS array and all two frames of images between the reference image, each frame image is magnified to an image with the same imaging magnification as the reference image to obtain a constant magnification focal stack image sequence without aberration.
[0070] It is understandable that the present invention uses the SIFT operator for feature extraction and the Flann and Knn algorithms for feature matching between images. Since the focus and defocus phenomena of different areas in the focus stack image sequence collected by the liquid lens are very different, the two non-adjacent frames in the stack appear to be very different, and problems such as glare and overexposure that may occur during the image acquisition process will further reduce the success rate of the feature point matching algorithm. Therefore, when performing feature point matching, feature point matching is first performed between two adjacent frames in the focus stack, and the transformation relationship between the two adjacent frames is determined based on the RANSAC alignment algorithm. Two consecutive adjacent frames of the focus stack image have relatively close focus and defocus phenomena, which ensures the success rate of the feature point matching algorithm. Finally, the transformation relationship between each frame image and the adjacent frames of the reference image is connected in series to obtain its transformation relationship with the reference image.
[0071] In some embodiments, performing focus status evaluation on the image sequence includes:
[0072] The pixel clarity of each frame of the image in the image sequence is evaluated, and a focus state evaluation result is obtained based on the evaluation result of the pixel clarity.
[0073] In some embodiments, evaluating the pixel clarity of each frame in the image sequence, and obtaining a focus state evaluation result based on the pixel clarity evaluation result, includes:
[0074] Converting each frame of the image in the image sequence into a grayscale image;
[0075] Calculating the focus and clarity of the grayscale image within a window of a set size;
[0076] Based on the focus and clarity of the grayscale image in a window of a set size, an evaluation result of the pixel clarity is obtained, and based on the evaluation result of the pixel clarity, a focus state evaluation result is obtained.
[0077] It can be understood that the focus stack images collected by the liquid lens are converted into gray scale images. For each pixel, the focusing degree and the sharpness can be evaluated by calculating the Laplace variance of the pixels in the 3*3 window in which the pixel is located. Since the defocus area in the image is approximately a low-pass filter, which causes serious attenuation of high-frequency information in the image, the Laplacian operator is used to calculate the second derivative when solving the regional intensity variation in the 3*3 window.
[0078] In some embodiments, the image sequence is full-focus fused based on the focus state evaluation result to obtain a full-focus image, including:
[0079] Based on the focus state evaluation result, the most clear pixel points of each frame of image in the image sequence are spliced into a full-focus image.
[0080] It can be understood that the 15 frames of focus stack images collected in the imaging test experiment are aligned to obtain the focus stack images with constant magnification using the focus stack alignment algorithm proposed above. Since no alignment processing is performed, there are many misaligned splicing areas in the synthesized full-focus image, such as the metal rod on the left side of the image, which has a relatively obvious misaligned splicing. This shows that the algorithm proposed above can effectively solve the imaging magnification change caused by the focal length change of the liquid lens and generate focus stack images with constant magnification.
[0081] In some embodiments, based on the object distance and the full-focus image, a depth map is generated, including:
[0082] Two-dimensional features of each frame of image are extracted, and a focus volume is formed based on the two-dimensional features. The focus volume is differentiated based on the object distance to obtain a differential focus volume;
[0083] Features of the differential focus volume are extracted, and the probability of the pixel points in the best focus state is predicted based on the extracted features;
[0084] Based on the predicted probability of the pixel points in the best focus state, probability regression is performed and combined with the focal length of each frame of image to generate a depth map reflecting the depth information of the image space.
[0085] In some embodiments, the liquid lens-based focus stack rapid acquisition principle provided by the application is as follows Figure 2As shown in (a) of FIG. 1, the role of the fixed-focus lens is to increase the converging ability of the imaging system to the light beam, thereby reducing the mechanical total length of the whole imaging system, which is a functional lens for compressing the optical path. Its position in the whole imaging system remains unchanged, and the imaging system only relies on the liquid lens to achieve focal length adjustment, so that the whole imaging system does not have mechanical moving parts except for the liquid-liquid interface in the liquid lens. The imaging targets A, B and C are located at different object distances in the object space. At this time, the three objects A, B and C are imaged at A', B' and C' respectively after being refracted by the fixed-focus lens and the liquid lens. At this time, the image A' is located in front of the CMOS image sensor and is in a defocused state; the image B' coincides with the CMOS image sensor and is in a focused state; and the image C' is located behind the CMOS image sensor and is in a defocused state. When the driving voltage increases, the converging ability of the imaging system to the light beam increases, and the image C' moves to the lens direction to coincide with the CMOS image sensor. When the driving voltage V decreases, the converging ability of the imaging system to the light beam decreases, and the image A' moves to the CMOS image sensor direction to coincide with it.
[0086] The prepared liquid lens has a millisecond-level zoom speed, so that the imaging system can achieve a large amplitude of focal length adjustment in a short time. As shown in (b) of FIG. 1, Figure 2 As shown in (b) of FIG. 1, by adjusting the driving voltage of the liquid lens by a large amplitude, the focusing plane of the imaging system can be quickly moved between the objects A, B and C which are far apart, thereby realizing the jump interval sampling of the object space scene. This jump interval sampling mode can collect a focus stack image sequence with a large object distance range in a short time, and then synthesize a full-focus image with a large depth of field range. According to the imaging performance test results, the prepared liquid lens is still sensitive to a small amplitude driving voltage change of 0.1 V, so that dense sampling of the object space scene can be realized by continuously and densely adjusting the driving voltage. As shown in (c) of FIG. 1, Figure 2 As shown in (c) of FIG. 1, by continuously adjusting the driving voltage by a small amplitude, dense sampling of the object B can be realized. This dense sampling mode is helpful to synthesize a full-focus image with more rich texture details and higher resolution, and has a very broad application prospect in the field of super-resolution imaging.
[0087] Imaging magnification change caused by lens focal length change:
[0088] When using a traditional mechanical zoom system for focus stack acquisition, mechanical lens displacement and changes in the distance between the lens and the CMOS image sensor can lead to significant differences in imaging magnification and field of view between acquired image sequences. This phenomenon is known as "lens breathing." While the spatial relationship between the optical elements in an imaging system using a liquid lens as the core zoom element is fixed, changes in the drive voltage can still cause slight displacement and deformation of the liquid-liquid interface within the liquid lens. This deformation and displacement can still cause slight variations in the imaging magnification of the focus stack image sequence.
[0089] Based on the focal length data of the liquid lens obtained from experimental measurements, an optical path simulation model was established using ZEMAX simulation software. For an imaging system consisting of two fixed-focus glass lenses and a liquid lens, when an 80 V driving voltage is applied, the optical path simulation model parameter settings are shown in Table 1.
[0090] Table 1: Liquid lens ZEMAX simulation model parameters
[0091]
[0092] When the driving voltage changes, only the thickness and curvature of plane 3, plane 4, plane 5 and plane 6 in Table 1 need to be changed. The optical path simulation models when the driving voltage is set to 80V and 55V are respectively Figure 3 As shown in (a) and (b) in the figure, the distance from the object plane to the first fixed-focus glass lens is set to 1000mm, and the object height is 150mm. According to the simulation results of the ZEMAX optical path simulation model, when the drive voltage is 55V, the image height of the object on the CMOS image sensor is 4.542mm. When the drive voltage is 80V, the image height of the object on the CMOS image sensor is 3.688mm. The simulation results show that as the drive voltage changes, the image height of the object on the CMOS photosensitive chip changes, that is, the imaging magnification of different images in the focal stack is inconsistent, and there are aberrations. In addition, the smaller the drive voltage, the larger the imaging magnification.
[0093] There is a slight change in the imaging magnification between the focus stack image sequences captured by the liquid lens in the imaging performance test. When the driving voltage is 59.0 V, the pixel size of the "stop" pattern of the school bus body in the captured focus stack image is 66x66 pixels. When the driving voltage is 60.0 V, the pixel size of the "stop" image of the school bus body in the captured focus stack image is 63x63. The change in pixel size indicates that the imaging magnification of the focus stack image sequence captured by the liquid lens is inconsistent, and the imaging magnification when the driving voltage is 59.0 V is slightly larger than the imaging magnification when the driving voltage is 60.0 V, which is basically consistent with the analysis result of the optical simulation model established based on ZEMAX. This change in imaging magnification will cause the full-focus image synthesized from the focus stack image sequence to have a stitching misalignment and pixel loss, and the focus stack needs to be aligned by hardware or software methods to make the entire image sequence have a constant magnification.
[0094] Focus stack alignment algorithm:
[0095] Some existing solutions, such as using six liquid lenses to jointly focus in the imaging system to achieve constant magnification imaging, but the increase in the number of liquid lenses will cause the imaging quality of the imaging system to decrease, and simultaneously adjusting the focal length of multiple liquid lenses is more complex and time-consuming. Some people have proposed a liquid lens with a tunable liquid-liquid interface, which maintains the imaging magnification and back focal length of the imaging system constant by adjusting the position of the liquid-liquid interface, but its structure is relatively complex and the focusing speed is slow, making it difficult to use in real-time application scenarios. Inspired by the serial optical flow algorithm proposed by Suwajanakorn, the present invention proposes a focus stack alignment algorithm for the imaging magnification change problem. According to the simulation and experimental results described above, when the driving voltage of the liquid lens is the lowest, the imaging system focuses on the farthest object and has the largest imaging magnification. Therefore, the image captured when the driving voltage is the lowest is selected as the reference image, and the remaining images in the focus stack are aligned with the reference image and enlarged to have the same imaging magnification as the reference image by the feature point matching algorithm shown in Table 2, thereby obtaining a constant magnification focus stack image sequence without aberration.
[0096] Table 2: Focus stack alignment algorithm flow
[0097]
[0098] The algorithm flow is shown in Table 2, using SIFT operator for feature extraction, and using Flann and Knn algorithm for feature matching between images. Due to the significant difference in focusing and defocusing phenomenon of different regions in the image sequence collected by the liquid lens focus stack, the difference between the non-adjacent two frames in the stack is very large, and the glare and overexposure problems that may occur during image acquisition will further reduce the success rate of the feature point matching algorithm. Using SIFT operator for feature point matching of two images collected at driving voltages of 65V and 51V cannot establish sufficient pixel mapping relationship. Therefore, when performing feature point matching, first perform feature point matching between adjacent two frames in the focus stack, and determine the transformation relationship between the two adjacent images based on the RANSAC registration algorithm. The two consecutive adjacent frames of the focus stack have relatively close focusing and defocusing phenomenon, which ensures the success rate of the feature point matching algorithm. Finally, the transformation relationship between the adjacent frames of each frame and the reference image is concatenated to obtain the transformation relationship between the reference image.
[0099] Depth of field extension and all-in-focus image synthesis based on liquid lens:
[0100] Depth of field (DoF) refers to the distance in front of and behind the focusing plane that can be clearly imaged, and is an important performance parameter of optical imaging systems. The depth of field of traditional imaging systems is usually determined by parameters such as lens focal length, object distance, and CMOS image sensor photosensitive cell size, and is relatively limited. By fusing focus stack images to generate all-in-focus images that can clearly image everywhere in the object space, the imaging system equipped with a liquid lens can have a larger depth of field than traditional imaging systems. Figure 4 As shown in FIG. 1, generating an all-in-focus image with a large depth of field through a focus stack mainly includes two steps: focus evaluation (FM) and image stitching. Focus evaluation refers to evaluating the clarity of each pixel in each image collected by the liquid lens, and establishing an index relationship (x, y, I, E) between the pixel and the image number to which the pixel belongs, where (x, y) is the pixel with horizontal and vertical coordinates x and y in image I, and E is the focus degree evaluation result of the pixel. After establishing the index relationship (x, y, I, E), the all-in-focus image is stitched by selecting the most clear pixels in each frame, as shown in FIG. 1(b). Figure 4
[0101] 1. Focus evaluation function
[0102] The focus stack images collected by the liquid lens are converted into gray scale images. For each pixel, the focusing degree and the sharpness can be evaluated by calculating the Laplace variance of the pixels in the 3x3 window. Since the defocus region in the image is approximately a low-pass filter, which causes a serious attenuation of the high-frequency information in the image, the Laplacian operator is used to calculate the second derivative when solving the regional intensity variation in the 3x3 window. The regional intensity variation, i.e. the sharpness, is denoted as E(x, y).
[0103] (1)
[0104] (2)
[0105] The aligned focus stack images are respectively named as where k = 1, …, N. Where N is the number of image frames in the focus stack. First, the images in the focus stack are converted into gray scale images. Based on equations (1) and (2), the regional intensity variation E(x, y) of all pixel points (x, y) in the kth image is calculated in turn and a four-tuple (x, y, I, E) is established to record the sharpness-image number index relationship. The intensity variation results E(x, y) of each pixel point (x, y) are normalized and converted into gray scale values. Where the driving voltage is 53V, 55V and 59V. According to the gray scale image, it can be seen that under the action of 53V, 55V and 59V, the imaging system focuses on the tank car at the farthest distance, the fire truck at the middle distance and the school bus at the nearest distance, which is consistent with the actual situation.
[0106] 2. All-in-focus image fusion method
[0107] The stitching problem of all-in-focus image can be converted into a multi-label MRF optimization problem. The picture composed of pixel points can be regarded as a 4-connected grid, and each pixel point constitutes a node set v, and the connection between adjacent pixels constitutes an edge set ε. The energy function is defined as:
[0108] (3)
[0109] wherein is a weighted constant used to balance v and ε. is the defocus degree of the pixel point calculated by equation (2). characterizes the change of image number The iterative computation can be performed by an a-expansion algorithm to obtain the minimized energy function. The a-expansion algorithm attempts to switch its label to the label of the neighboring node in the iteration process, and accepts the label replacement if the energy function decreases, otherwise continues the iteration. After the iteration, the energy function is minimized, and the all-in-focus image with good stitching effect can be obtained according to the label of each pixel node.
[0110] Using the focus stack alignment algorithm proposed above, 15 frames of focus stack images collected in the imaging test experiment are aligned to obtain focus stack images with constant magnification. The all-in-focus image generated by the above image algorithm, the tank, the fire truck and the school bus are clearly imaged, and the image depth of field is obviously improved. There is a part of the blurred area on the roof of the school bus, which is caused by the too high brightness setting of the LED light source used during shooting, resulting in overexposure of the image in this area, and the original focus stack image collected loses the details of this part of the image, so there is a part of the blur. Since there is no alignment processing, there are many misaligned areas in the synthesized all-in-focus image, such as the metal rod on the left side of the image, which has a more obvious misaligned stitching, which shows that the algorithm proposed above can effectively solve the change of imaging magnification caused by the change of focal length of the liquid lens and generate focus stack images with constant magnification.
[0111] 3. Depth of field extension based on liquid lens
[0112] To further quantitatively analyze the depth of field extension capability of the liquid lens to the imaging system, in this section, the prepared liquid lens is combined with two fixed-focus glass lenses and a CMOS image sensor to make a focus stack acquisition camera. The fixed-focus glass lenses used in the focus stack acquisition camera are two flat-convex glass lenses made of H-K9L material, with a curvature radius of 25.75 mm and a diameter of 12.7 mm. The fixed-focus glass lens mainly plays two roles, one is to increase the convergence ability of the imaging system to the light beam, compress the optical path, and reduce the total length of the imaging system; the other is to couple with the liquid lens to expand the image field, so that the image field covers the CMOS image sensor as much as possible. The distance between the two fixed-focus glass lenses is set to 7 mm to facilitate the insertion of the liquid lens, and the distance between the fixed-focus glass lens and the CMOS image sensor is 22.2 mm. The selected CMOS image sensor has a photosensitive cell size of 2.2x2.2μm, a photosensitive cell array size of 2592x1944, a signal-to-noise ratio of 39dB, and a dynamic range of 85dB. The connecting piece is prepared by 3D printing to connect the components together.
[0113] The optical path simulation model of the focus stack camera was established based on ZEMAX, and the prepared liquid lens can be equivalent to a double-cemented lens with an optical aperture of 5 mm. The two component materials of the equivalent double-cemented lens are 0.1% SDS deionized water solution and dimethyl silicone oil, respectively. The refractive index fitting of the two liquid materials was carried out based on the Schott method, and was added to the ZEMAX glass material library, so that the focal length and depth of field of the focus stack camera under different curvatures can be simulated and analyzed.
[0114] For a traditional optical imaging system, the light emitted by the object point is refracted by the lens to form a circular diffused spot on the CMOS image sensor. Only when the entire diffused spot is completely within a photosensitive pixel, the image formed by the object point is clear. As shown in Figure 5 , object points A and C are not completely focused on the CMOS image sensor as object point B, but the diffused spots formed by the refraction of the lens are still completely within a pixel, so the images are still clear. The object points between object points A and C can be clearly imaged, and this distance is the depth of field. Assuming that the current image distance is focused on object point B with an object distance of u, the image distance is v. The size of a single pixel of the CMOS image sensor is δ, and D is the size of the main lens aperture. The front focal depth and the back focal depth of the imaging system can be represented as:
[0115] (4)
[0116] (5)
[0117] The depth of field of the imaging system can be obtained by combining equations (4) and (5) with the thin lens imaging formula:
[0118] (6)
[0119] When the liquid lens is not inserted, the focal length of the fixed-focus lens composed of two pieces of H-K9L fixed-focus glass with a curvature radius of 25.75 mm and a spacing of 7 mm is 27.4435 mm. When the pixel size of the CMOS image sensor is 2.2 μm, the relationship curve of the depth of field with the image distance can be calculated by equation (6) as shown in Figure 6 . The image distance is the distance between the equivalent optical centers of the two fixed-focus lenses and the CMOS image sensor, i.e. 28.3 mm. At this image distance, the depth of field of the focus stack camera is 23.2 mm, and the depth of field range is relatively narrow.
[0120] Three imaging targets, a tank, an airplane and a truck, were placed at 920 mm, 740 mm and 540 mm in front of the focus stack acquisition camera, respectively. Twelve focus stack images focusing on different object distances were acquired by adjusting the driving voltage. Among them, when the driving voltage was 61.2 V, 62.4 V and 63.0 V, the imaging system was focused on the tank, the airplane and the truck, respectively. The aberration-free image sequence with constant magnification was generated based on the focus stack alignment algorithm proposed above. In the all-in-focus image, the objects in the range of 540~920 mm object distance are clearly imaged, which shows that the system depth of field is successfully extended to 380 mm based on the liquid lens, which is at least one order of magnitude higher than the original depth of field range of 23.2 mm.
[0121] And because the prepared liquid lens has a working voltage range of 0~110V, it can continue to acquire focus stack images in a larger object distance range by expanding the driving voltage adjustment range and synthesize all-in-focus images with a depth of field range greater than 380mm. In summary, the proposed focus stack acquisition method and all-in-focus image synthesis method based on liquid lens can effectively realize large depth of field imaging beyond traditional optical imaging systems.
[0122] Fast depth estimation method based on liquid lens:
[0123] 1. Depth estimation principle based on liquid lens
[0124] The depth information of an image, i.e. the object distance in the image, is a key clue to restore the three-dimensional scene in the object space. In recent years, monocular depth estimation methods based on deep learning have developed rapidly, which can estimate depth directly from a single image, but this method relies heavily on image semantic information and requires a large amount of data set for training. Focus stack is essentially a series of images focusing on different spatial depths. In addition to the semantic information in each image in the stack, the difference in focus and defocus state between different images in the stack can also be used for depth estimation. The method of depth estimation based on focus stack can be divided into depth estimation method based on focus clarity (Depth from focus, DFF) and depth estimation method based on defocus blur (Depth from defocus, DFD). DFF technology assumes that each imaging target in the focus stack has an image with the best focus. For densely sampled focus stack images, this assumption is true. The main problems in using DFF technology for fast depth estimation of object space scene are as follows: (1) fast acquisition of focus stack image sequence; (2) locating the pixel points with the best focus state in each image; (3) determining the focus distance of each image in the focus stack.
[0125] (1) Fast acquisition of focus stack image sequence
[0126] The DFF technique has higher requirements for the dynamic focusing capability of the image sensor used to collect the focus stack. First, the image sensor should have a large focal length range to collect image sequences with a large object distance range. Second, the image sensor should have a high zoom speed to reduce the time required for image collection. Third, the focal length of the image sensor should be able to be continuously adjusted by a small amount to collect a dense focus stack. According to the optical performance of the prepared liquid lens, the effective focal length of the liquid lens can be continuously adjusted in the range of (-∞, -128.6mm) (48.6mm, +∞), which has a large focal length range. Second, the liquid lens has a millisecond-level zoom speed, which can achieve image collection of about 10-30 frames per second. According to the imaging performance test results of the liquid lens, the lens is still sensitive to small changes in the driving voltage, and can achieve dense sampling of the spatial scene on the object side. Therefore, the prepared liquid lens is sufficient to meet the requirements of the DFF technique for the image sensor. In addition to the above three points, the liquid lens has the advantages of no mechanical moving parts, etc., which avoids the aberration caused by mechanical disturbance of the lens.
[0127] (2) Locating the pixel points with the best focusing state in each frame of image
[0128] Depth estimation by DFF technique also faces the problem of evaluating the focusing degree of pixel points, especially when estimating the depth of a real spatial scene, there will be a large number of textureless regions. The traditional Laplacian operator has difficulty in accurately evaluating the focusing degree of pixel points in textureless regions. Second, in order to ensure that each coordinate pixel has a frame of its most clear image, a dense focus stack is often used as input, which usually requires dozens of frames of images as input. The increase in the number of input images increases the cost of image sequence collection on the one hand, and also reduces the running speed of the model on the other hand. The network model based on differential focus volume (Differential Focus Volume, DFV) solves the above problems well. It uses the powerful feature extraction capability of 2D convolutional neural network (Convolutional Neural Network, CNN) to evaluate the focusing degree of pixel points, and uses the focus distance of focus stack image and each frame of image to construct DFV. After successfully constructing DFV, it uses 3D CNN to predict the probability of each frame of image reaching the best focusing state, and combines it with the focus distance (object distance) when collecting each frame of image to complete depth estimation. The model currently has high depth estimation accuracy in multiple public datasets.
[0129] (3) Determination of the focus distance of each frame of image in the focus stack
[0130] How to obtain the focusing distance of each frame of focus stack image is the key to absolute depth estimation of real-world scenes using DFF technology. Unlike performing depth estimation tasks on datasets, when using network models such as DFV to perform absolute depth estimation on real space scenes, the focusing distance of each frame of image in the focus stack, i.e. the object distance corresponding to the sharpest pixel point in the image, needs to be obtained. In previous work, the focusing distance of the picture was mainly obtained by calibrating traditional optical zoom cameras through manual calibration, etc. This method has large deviation in determining the focusing distance and poor generalization. The liquid lens based on dielectric wetting effect has obvious advantages in fast calculation of focusing distance.
[0131] 2. Focal length electric addressing characteristics of liquid lens
[0132] There is a one-to-one electric addressing relationship between the focal length of the prepared liquid lens and the driving voltage. This electric addressing relationship makes the driving voltage of the liquid lens available for fast calculation of the focal length of the liquid lens and the focal length of the imaging system, thereby realizing fast calculation of spatial depth information such as depth of field and object distance. Taking the focus stack acquisition camera constructed above as an example, the system focal length-driving voltage relationship of the camera is calculated based on ZEMAX, and the obtained driving voltage-focal length relationship curve is as shown in Figure 7 When the image of the object formed after being refracted by the lens is located at the CMOS image sensor, the image of the object is the sharpest, and the image distance v is the distance from the equivalent optical center of the lens to the plane of the CMOS image sensor. According to the lens imaging formula, there is a relationship between the focal length f of the lens, the object distance u and the image distance v. Substituting the simulation data in Figure 7 into formula (7) can calculate the focusing distance of the imaging system under the action of different driving voltages, i.e. the object distance of the sharpest pixel point in the image, as shown in Figure 8 . Further, a one-to-one indexing relationship is established between the driving voltage and the focusing distance. This indexing relationship is calculated based on the objective physical law of the Lipmann-Yang equation, and using it as a clue to estimate the depth information of the object space scene can ensure the estimation accuracy and generalization of the depth estimation method.
[0133] In summary, the unique focal length electric addressing characteristics of the dielectric wetting liquid lens make it very suitable for combining with the DFV network model with higher depth estimation accuracy at the present stage to quantitatively estimate the absolute depth of the object space scene.
[0134] 3. Depth estimation module architecture and implementation
[0135] According to the DFV-DFF network architecture, the network is built, and the overall framework of the depth estimation module is as shown in Figure 9The network model is shown in FIG. 1. A driving voltage-object distance conversion module based on the focal length electrically addressable characteristics of the liquid lens is introduced on the basis of the network model. When training, the model input is the focus stack images in the public datasets FoD500 and DDFF-12. When performing depth estimation experiments on real space scenes, the model input is the focus stack image sequence collected by the liquid lens. The above focus stack alignment and aberration correction algorithm is used to convert it into an aberration-free focus stack image sequence with constant magnification. Each frame of image in the stack is labeled with the driving voltage at the time of collection. According to the Lipmann-Yang equation and the principle of lens imaging, the driving voltage at the time of collection of each frame of image is converted into the object distance (focusing distance of the imaging system). The model first uses a 2DCNN model to extract the two-dimensional features of each frame of image in the stack, i.e., the focus volume (FV). Next, by calculating the first derivative of the two-dimensional feature stack, a DFV can be constructed. For the DFV constructed, a 3DCNN is used to extract features of the DFV to predict the probability of whether a pixel point is in the best focusing state. Finally, the depth map reflecting the depth information of the photo space is formed through probability regression and combining the focusing distances of each frame of image.
[0136] The above model is implemented using PyTorch 1.10.2, and the model is optimized using the Adam optimizer (learning rate lr = 0.0001, = 0.9, = 0.9) on an 8GB video memory NVIDIA 3070 Laptop GPU. The experiment uses a fixed random seed, the batch size is 10, and the training is 500 rounds. When training, 5 frames of images are randomly selected from the focus stack each time and are cropped and scaled to a resolution of 224x224. In the training process, random horizontal flipping and random rotation are used to avoid overfitting. In the training process, a depth map with four resolution sizes is formed by upsampling, so the loss function is defined as follows:
[0137] (7)
[0138] wherein is the predicted depth of pixel point numbered j at s sampling level, is the true depth of pixel point numbered j, and M is the total number of pixel points, is the weighting coefficient for balancing different sampling levels, wherein , , , and are 8 / 15, 4 / 15, 2 / 15 and 1 / 15, respectively. The model is trained using the public datasets DoD500 and DDFF-12.
[0139] The FoD500 dataset contains 500 groups of samples. Each group of samples contains five frames of focus stack images with known focus distances and a frame of depth map representing the depth value of each pixel. The resolution of each frame of image is 256x256. Since the dataset is originally constructed for the DFD task, the focus distance range in each group of samples does not completely cover the real depth range, and the data beyond the range needs to be hidden in training and testing. The first 400 groups of samples are selected as the training set, and the last 100 groups of samples are selected as the test set.
[0140] The DDFF-12 dataset is a dataset collected from 12 different real-world scenes using a light field camera. Among them, the kitchen, office A, conference room, social corner, student laboratory and glass room six spatial scenes each contain 100 groups of samples. For the other six spatial scenes: cafeteria, library, living room, office B and laboratory, each scene contains 20 groups of samples. Each group of samples contains 10 frames of focus stack images and corresponding field of view maps. The resolution of each frame of image is 383x552. The samples in the kitchen, conference room, social corner and student laboratory scenes are selected as the training set, and the samples in the glass room and office 1 scenes are selected as the test set.
[0141] The following evaluation indicators are used in the experiment to evaluate the estimation accuracy of the model. The evaluation indicators are: mean square error (MSE), root mean square error (RMS), absolute relative error (absRel) and estimation time. Based on the above evaluation indicators, the trained model is evaluated, and the results are shown in Table 3. It is shown that the model trained by the network architecture proposed by the present application has good depth estimation accuracy.
[0142] Table 3: Test results of the model on the FoD500 and DDFF-12 datasets
[0143]
[0144] 4. Absolute depth estimation experiment in real scene
[0145] Based on the preset experimental scene, the liquid lens-based fast depth estimation method is tested, wherein the tank, the airplane and the truck are respectively placed at 920mm, 740mm and 540mm in front of the camera. Two frames of focus stack images with driving voltages of 61.2V and 63.0V are taken as inputs. The generated model can better evaluate the focusing probability of each pixel in each frame image. The output focusing probability map shows that the model can better predict the best focusing probability of each pixel in each frame image, and the object contour recognition is more accurate.
[0146] The cold and warm color tones in the focusing probability map reflect the focusing state probability of the pixel points. The warmer the color tone, the better the focusing state. When the driving voltage is 61.2V, the imaging system focuses on the tank, and the color tone of the tank area in the focusing probability map output by the model is obviously warmer than other areas. When the driving voltage is 63.0V, the imaging system focuses on the nearest truck, and the color tone of the truck area is obviously warmer than other areas.
[0147] The depth image finally output by the model reflects the distance of each region in the image from the camera lens. The cold and warm color tones of the image region in the depth map reflect the distance of the object from the lens. The warmer the color tone, the larger the object distance, and the colder the color tone, the smaller the object distance. The imaging area corresponding to the truck has a cold color tone, the imaging area corresponding to the airplane has a moderate color tone, and the imaging area corresponding to the tank has a warm color tone, which better reflects the relative depth relationship of the three imaging targets. Although only two frames of focus stack images focusing on the tank target and the truck target are input, based on the powerful feature extraction capability of CNN, the DFV network model still accurately rates the focusing clarity of each pixel point through probability regression, and accurately estimates the object distance of the airplane.
[0148] The driving voltage-object distance conversion module based on the focal length electric addressing characteristics of the liquid lens converts the focusing distance of the two input images into the focusing distance of the image, and then estimates the object distance of the object corresponding to each pixel point in the image, and realizes the reconstruction of the three-dimensional scene in the object space. The color bar on the right side of the depth map identifies the correspondence between the cold and warm color tones in the depth map and the actual object distance. The actual object distance of the three imaging targets is compared with the estimated object distance obtained by the model. The comparison data is shown in Table 4.
[0149] Table 4: Actual object distance of imaging target and estimated object distance obtained by algorithm
[0150]
[0151] The experimental results show that by combining the focal length electrically addressable characteristics of the liquid lens with the DFV depth estimation network model, the object distance corresponding to each pixel point in the focus stack image can be estimated only by two frames of focus stack images. The object distance of the background at the back of the image in the output depth map is estimated incorrectly, which may be due to the interference of the large letters in the background of the image on the pixel definition evaluation, which can be improved by increasing the number of input frames of focus stack images. In summary, the liquid lens-based depth estimation method proposed in the application has high depth estimation accuracy within a certain object distance range, with a maximum error rate of 4.1% and an average error rate of 3.43% for the estimation results of the three imaging targets. And the liquid lens dynamic response time and model running time measured by the experiment show that the imaging system can complete focus stack image acquisition within 200ms and complete depth estimation task within about 50ms, realizing fast depth estimation of the object space scene.
[0152] The application proposes a variety of modes of focus stack fast acquisition method based on the electrically controlled focusing performance of the liquid lens. For the application scenarios of depth of field extension and fast depth estimation, a jump interval sampling mode is proposed to quickly obtain a focus stack image sequence of a larger object distance range. For the application scenarios of super-resolution imaging, a dense sampling mode is proposed to quickly obtain a focus stack image sequence of dense sampling of the object space. For the change of magnification of the focus stack image caused by the change of back focal length during the zooming of the liquid lens, a corresponding focus stack alignment and aberration correction algorithm is proposed. The experimental results show that the all-in-focus image synthesized by the focus stack image sequence after alignment processing effectively avoids image stitching misalignment and has high imaging quality. Based on the excellent electrically controlled focusing performance of the liquid lens, the application proposes corresponding depth of field extension and image fusion algorithms, which increase the depth of field of the imaging system by at least one order of magnitude. The application combines the Lipmann-Yang equation with the differential focus volume network model to estimate the object distance of the imaging target in the object space scene, with a maximum error of 4.1% and an average error of 3.4%.
[0153] As shown in Figure 10 The application also correspondingly provides an electronic device 1000. The electronic device 1000 includes a processor 1001, a memory 1002, and a display 1003. Figure 10 Only part of the components of the electronic device 1000 are shown, but it should be understood that all the shown components are not required, and more or fewer components can be alternatively implemented.
[0154] The memory 1002 can be an internal storage unit of the electronic device 1000, such as a hard disk or a memory of the electronic device 1000 in some embodiments. The memory 1002 can also be an external storage device of the electronic device 1000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, and the like, equipped on the electronic device 1000 in other embodiments.
[0155] Further, the memory 1002 can include both an internal storage unit and an external storage device of the electronic device 1000. The memory 1002 is used to store application software and various data installed on the electronic device 1000.
[0156] The processor 1001 can be a Central Processing Unit (CPU), a microprocessor, or other data processing chip in some embodiments, used to run program codes or process data stored in the memory 1002, such as the focal stack imaging method of the liquid lens in the present application.
[0157] The display 1003 can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, and the like in some embodiments. The display 1003 is used to display information of the electronic device 1000 and to display a visualized user interface. The components 1001-1003 of the electronic device 1000 communicate with each other through a system bus.
[0158] In some embodiments of the present application, when the processor 1001 executes the focal stack imaging program of the liquid lens in the memory 1002, the following steps can be implemented:
[0159] The power supply control module provides a driving voltage for the liquid lens to form an image on the CMOS array;
[0160] An image collected by the CMOS array is acquired, and all collected images are aligned based on the image when the driving voltage is lowest to obtain an image sequence of anastigmatic constant magnification focal stack;
[0161] The image sequence is evaluated for a focusing state, and the image sequence is fused based on a focusing state evaluation result to obtain a full-focus image;
[0162] A subject distance of the collected image is determined, a depth map is generated based on the subject distance and the full-focus image, and a target image is generated based on the full-focus image and the depth map.
[0163] It should be understood that, in addition to the above functions, the processor 1001 can also implement other functions when executing the liquid lens focus stack imaging program in the memory 1002, which can be specifically understood in the description of the corresponding method embodiments.
[0164] Further, the type of the electronic device 1000 mentioned in the embodiments of the present application is not specifically limited, and the electronic device 1000 can be a portable electronic device such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, etc. Exemplary embodiments of the portable electronic device include but are not limited to a portable electronic device running an IOS, android, microsoft or other operating system. The above-mentioned portable electronic device can also be other portable electronic devices, such as a laptop computer having a touch-sensitive surface (e.g., a touch panel), etc. It should also be understood that in some other embodiments of the present application, the electronic device 1000 can also not be a portable electronic device, but a desktop computer having a touch-sensitive surface (e.g., a touch panel).
[0165] In another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the liquid lens focus stack imaging method provided by the above-mentioned methods, and the method comprises:
[0166] The power supply module is controlled to provide a driving voltage for the liquid lens to form an image on the CMOS array;
[0167] An image collected by the CMOS array is acquired, and all collected images are focus stack aligned based on the image when the driving voltage is lowest to obtain an image sequence of anastigmatic constant magnification focus stack;
[0168] The image sequence is evaluated for focus state, and the image sequence is full-focus fused based on the focus state evaluation result to obtain a full-focus image;
[0169] The object distance of the collected image is determined, a depth map is generated based on the object distance and the full-focus image, and a target image is generated based on the full-focus image and the depth map.
[0170] The present application also provides a liquid lens focus stack imaging system, comprising a power supply module, a coaxial fixed focus lens, a liquid lens and a CMOS array, and the above-mentioned electronic device, wherein the electronic device is electrically connected to the power supply module and communicatively connected to the CMOS array.
[0171] Those skilled in the art can understand that all or part of the processes of the above-mentioned embodiment methods can be completed by instructing the relevant hardware by a computer program, and the program can be stored in a computer readable storage medium. The computer readable storage medium is a magnetic disk, an optical disk, a read-only memory, a random access memory, etc.
[0172] The liquid lens focus stack imaging method, system and electronic device provided by the present application are described in detail above, and the principles and implementation modes of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed, and the above description should not be understood as a limitation of the present application.
Claims
1. A focal stack imaging method for a liquid lens, characterized in that: Applied to a focal stack imaging device, the device includes: a power supply module, and a coaxial fixed-focus lens, a liquid lens, and a CMOS array; the method includes: Controlling the power supply module to provide a driving voltage to the liquid lens so as to form an image on the CMOS array; Acquire images captured by the CMOS array, and perform focal stack alignment on all captured images based on the image at the lowest driving voltage to obtain an image sequence of aberration-free constant magnification focal stack; performing a focus state evaluation on the image sequence, and performing full focus fusion on the image sequence based on the focus state evaluation result to obtain a full focus image; determining an object distance of the acquired image, generating a depth map based on the object distance and the all-in-focus image, and generating a target image based on the all-in-focus image and the depth map; Based on the image at the lowest drive voltage, all acquired images are focal stack aligned to obtain an aberration-free constant magnification focal stack image sequence, including: An image captured by the CMOS array when the driving voltage is lowest is used as a reference image, and based on a feature matching algorithm, the focal stack image of the CMOS array is magnified to an image having the same imaging magnification as the reference image, so as to obtain a constant magnification focal stack image sequence without aberration; Wherein, the focal stack comprises images focused at different spatial depths acquired by the CMOS array; The image of the focal stack of the CMOS array is magnified to an image having the same imaging magnification as the reference image based on a feature matching algorithm to obtain a constant magnification focal stack image sequence without aberration, comprising: performing feature point matching on two adjacent frames of images in the focal stack of the CMOS array to determine a transformation relationship between the two adjacent frames of images; Based on the transformation relationship between each frame image in the focal stack of the CMOS array and all two frames of images between the reference image, each frame image is magnified to an image with the same imaging magnification as the reference image to obtain a constant magnification focal stack image sequence without aberration.
2. The focal stack imaging method of liquid lens according to claim 1, characterized in that: The control power supply module provides a driving voltage for the liquid lens, including: determining a driving voltage required for the liquid lens based on the focal depth, target depth of field, and image distance of the focal stack imaging device; Based on the driving voltage required by the liquid lens, the power supply module is controlled to provide the driving voltage to both ends of the liquid lens, so as to form an image of the target depth of field on the CMOS array.
3. The focal stack imaging method of liquid lens according to claim 1, characterized in that: Performing a focus state evaluation on the image sequence, comprising: The pixel clarity of each frame of the image in the image sequence is evaluated, and a focus state evaluation result is obtained based on the evaluation result of the pixel clarity.
4. The focal stack imaging method of liquid lens according to claim 3, characterized in that: Evaluating the pixel clarity of each frame in the image sequence, and obtaining a focus state evaluation result based on the pixel clarity evaluation result, including: Converting each frame of the image in the image sequence into a grayscale image; Calculating the focus and clarity of the grayscale image within a window of a set size; Based on the focus and clarity of the grayscale image in a window of a set size, an evaluation result of the pixel clarity is obtained, and based on the evaluation result of the pixel clarity, a focus state evaluation result is obtained.
5. The focal stack imaging method of liquid lens according to claim 1, characterized in that: Performing full focus fusion on the image sequence based on the focus state evaluation result to obtain a full focus image, including: Based on the focus state evaluation result, the clearest imaging pixels of each frame image in the image sequence are spliced into a fully focused image.
6. The focal stack imaging method of a liquid lens according to any one of claims 1 to 5, characterized in that: Generating a depth map based on the object distance and the all-focus image, comprising: Extracting two-dimensional features of each frame of image, forming a focal volume based on the two-dimensional features, and performing differentiation processing on the focal volume based on the object distance to obtain a differential focal volume; Extracting features from the differential focus volume, and predicting the probability of a pixel being in an optimal focus state based on the extracted features; Based on the predicted probability of the pixel point being in the best focus state, probabilistic regression is performed and combined with the focal length of each frame image to generate a depth map reflecting the image spatial depth information.
7. An electronic device, characterized in that: comprising a memory and a processor, wherein, The memory is used to store programs; The processor is coupled to the memory and is configured to execute the program stored in the memory to implement the steps of the focal stack imaging method of the liquid lens as claimed in any one of claims 1 to 6.
8. A focal stack imaging system of a liquid lens, characterized in that: include: A power supply module, a coaxial fixed-focus lens, a liquid lens and a CMOS array, and the electronic device according to claim 7, wherein the electronic device is electrically connected to the power supply module and is communicatively connected to the CMOS array.