Rapid automatic focusing method, device and equipment based on light field camera and medium

By using microlens arrays in the light field camera for space and angle information reconstruction, filtering the images to be processed and adjusting the focus distance, the problem of time and calculation of the light field camera autofocus is solved, and a fast and efficient autofocus effect is achieved.

CN120050517APending Publication Date: 2025-05-27META-RETINA (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204694.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The automatic focus technology of light field cameras faces problems such as long time, large calculation amount and high computing power required. The traditional camera automatic focus method is not completely applicable to light field cameras.

Method used

By reconstruction based on the spatial and angle information of the microlens array, images with high information richness are selected, and the optimal focus distance is determined based on the fusion degree and clarity of the angle information in the preset focus window. If the optimal focus distance is not met, adjust the distance between the optical system and the light field camera section to achieve optimal focus.

Benefits of technology

It realizes fast automatic focus of the light field camera, reduces time-consuming and computational quantities, and improves focus efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050517A_ABST
    Figure CN120050517A_ABST
Patent Text Reader

Abstract

The invention provides a rapid automatic focusing method, device and equipment based on a light field camera and a medium, and the method comprises the steps: carrying out the space and angle information reconstruction of a first sub-image according to the arrangement mode of a micro-lens array, and obtaining a first image set; screening out a first to-be-processed image according to the information richness corresponding to each first image in the first image set; determining whether the first distance is an optimal focusing distance or not according to the angle information fusion degree and definition of the first to-be-processed image in a preset focusing window; wherein the optimal focusing distance is used for light field reconstruction. The problems of long time consumption, large calculation amount, high required calculation power and the like in the light field focusing process can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of light field image processing. Specifically, it relates to a fast autofocus method, device, equipment and medium based on a light field camera. Background Technique

[0002] Since the light field technology was proposed in the 1990s, with the development of computational optics, the performance of light field cameras has been significantly enhanced. Such cameras use a special microlens array structure to simultaneously record the position and direction of light rays, thus expanding the traditional two-dimensional image to four dimensions. This feature makes light field cameras a core device in many scientific and technological fields and has been widely used in fields such as remote sensing sky surveys, medical equipment, and mobile terminals.

[0003] However, the autofocus technology of light field cameras faces severe challenges. Due to the high-dimensional characteristics of light field data, traditional camera autofocus methods, such as active and passive methods, are not fully applicable to light field cameras. The active method requires auxiliary equipment to calculate the depth of the target area, while the passive method relies on the clarity of the already captured image to iteratively find the natural image plane. However, the light field image data collected by light field cameras cannot be directly iterated in the traditional passive way of calculating clarity, and the mechanical iteration method takes time that is exponentially related to the scene depth, resulting in low focusing efficiency. The characteristic of the light field camera to expand the depth of field greatly increases the iteration time. Summary of the Invention

[0004] In view of this, the purpose of the present application is to provide a fast autofocus method, device, equipment and medium based on a light field camera to overcome the problems in the prior art.

[0005] In a first aspect, an embodiment of the present application provides a fast autofocus method based on a light field camera, which acts on a target light field camera. The target light field camera includes an optical system part and a light field camera part. The light field camera part includes a microlens array and a sensor. Among them, the optical system part is spaced from the light field camera part by a first distance; the target light field camera is used to capture a first image based on the first distance; the first image includes a plurality of first sub-images; the method includes:

[0006] According to the arrangement mode of the microlens array, reconstruct the spatial and angular information of the first sub-images to obtain a first image set;

[0007] According to the information richness corresponding to each first image in the first image set, screen out the first image to be processed;

[0008] Determine whether the first distance is the optimal focusing distance according to the angular information fusion degree and clarity of the first image to be processed within the preset focusing window; wherein, the optimal focusing distance is used for light field reconstruction.

[0009] In some technical solutions of the present application, if the first distance is not the optimal focusing distance, the method further includes:

[0010] Adjust the distance between the optical system part and the light field camera part until the optimal focusing distance is obtained.

[0011] In some technical solutions of the present application, before obtaining the first image set, the method further includes: determining the focusing window; or, after obtaining the first image set, the method further includes: determining the focusing window;

[0012] The method determines the focusing window in the following manner:

[0013] Set a first weight for each of the target areas according to the fullness of the texture information within the preset target area;

[0014] Perform a scaling process with the target area having the largest first weight as the center, and calculate the second weights of each of the target areas after scaling;

[0015] Select the focusing window according to the target weight and the density of the texture information; wherein, the target weight is calculated from the first weight and the second weight;

[0016] Alternatively, determine the focusing window through a pre-trained model.

[0017] In some technical solutions of the present application, the above-mentioned selection of the first image to be processed according to the information richness corresponding to each first image in the first image set includes:

[0018] If the information richness corresponding to the first image is within a preset first range, select a first number of images as the first image to be processed;

[0019] If the information richness corresponding to the first image is within a preset second range, select a second number of images as the first image to be processed;

[0020] If the information richness corresponding to the first image is within a preset third range, select a third number of images as the first image to be processed.

[0021] In some technical solutions of the present application, the method calculates the angular information fusion degree of the first image to be processed in the following manner:

[0022] Calculate the angular information fusion factor between the first images to be processed;

[0023] Calculate the angular information fusion degree of the first images to be processed according to the angular information fusion factor.

[0024] In some technical solutions of this application, calculating the angular information fusion factor between the first images to be processed includes:

[0025] Calculate the mean square error and the structural similarity index between the first images to be processed;

[0026] Calculating the angular information fusion degree of the first images to be processed according to the angular information fusion factor includes:

[0027] Calculate the angular information fusion degree of the first images to be processed according to the mean square error and the structural similarity index.

[0028] In some technical solutions of this application, the above method calculates the sharpness of the first images to be processed in the following manner:

[0029] Calculate the sharpness of the first images to be processed according to the dispersion degree of the high-frequency components of the first images to be processed.

[0030] In a second aspect, an embodiment of this application provides a fast autofocus device based on a light field camera, which acts on a target light field camera. The target light field camera includes an optical system part and a light field camera part. The light field camera part includes a microlens array and a sensor. Among them, the optical system part and the light field camera part are spaced apart by a first distance; the target light field camera is used to capture a first image based on the first distance; the first image includes a plurality of first sub-images. The device includes:

[0031] A reconstruction module, configured to reconstruct the spatial and angular information of the first sub-images according to the arrangement mode of the microlens array to obtain a first image set;

[0032] A screening module, configured to screen out the first images to be processed according to the information richness corresponding to each first image in the first image set;

[0033] A determination module, configured to determine whether the first distance is the optimal focus distance according to the angular information fusion degree and the sharpness of the first images to be processed within a preset focus window; wherein, the optimal focus distance is used for light field reconstruction.

[0034] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned fast autofocus method based on a light field camera are implemented.

[0035] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the above-mentioned fast autofocus method based on a light field camera are executed.

[0036] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:

[0037] The method of the present application includes reconstructing the spatial and angular information of the first sub-image according to the arrangement mode of the microlens array to obtain a first image set; screening out a first image to be processed according to the information richness corresponding to each first image in the first image set; and determining whether the first distance is the optimal focus distance according to the angular information fusion degree and clarity of the first image to be processed within a preset focus window. The optimal focus distance is used for light field reconstruction.

[0038] The present application can solve the problems of long time consumption, large computational amount, and high required computing power in the light field focusing process.

[0039] The above objects, features, and advantages of the present application will become more obvious and understandable. The following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.

[0041] Figure 1 Shows a schematic flowchart of a fast autofocus method based on a light field camera provided by an embodiment of the present application;

[0042] Figure 2 Shows a schematic diagram of a first image provided by an embodiment of the present application;

[0043] Figure 3a Shows a schematic diagram of the arrangement of the first microlens array provided by an embodiment of the present application;

[0044] Figure 3bShows the schematic diagram of the second arrangement of the microlens array provided by the embodiments of the present application;

[0045] Figure 3c Shows the schematic diagram of the third arrangement of the microlens array provided by the embodiments of the present application;

[0046] Figure 3d Shows the schematic diagram of the fourth arrangement of the microlens array provided by the embodiments of the present application;

[0047] Figure 3e Shows the schematic diagram of the fifth arrangement of the microlens array provided by the embodiments of the present application;

[0048] Figure 3f Shows the schematic diagram of the sixth arrangement of the microlens array provided by the embodiments of the present application;

[0049] Figure 3g Shows the schematic diagram of the seventh arrangement of the microlens array provided by the embodiments of the present application;

[0050] Figure 3h Shows the schematic diagram of the eighth arrangement of the microlens array provided by the embodiments of the present application;

[0051] Figure 4 Shows the schematic diagram of the extraction of a first sub-image provided by the embodiments of the present application;

[0052] Figure 5 Shows the schematic diagram of the reconstruction process provided by the embodiments of the present application;

[0053] Figure 6 Shows the schematic diagram of a target area provided by the embodiments of the present application;

[0054] Figure 7 Shows the schematic diagram of a fast autofocus device based on a light field camera provided by the embodiments of the present application;

[0055] Figure 8 Shows the schematic diagram of the structure of an electronic device provided for the embodiments of the present application by the embodiments of the present application. Detailed implementation manners

[0056] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. It should be understood that the accompanying drawings in this application are only for the purposes of illustration and description, and are not used to limit the protection scope of this application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application.

[0057] In addition, the described embodiments are only some embodiments of this application, rather than all embodiments. The components of the embodiments of this application usually described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of this application to be protected, but only represents selected embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.

[0058] It should be noted that the term "including" will be used in the embodiments of this application to indicate the existence of the features stated thereafter, but does not exclude the addition of other features.

[0059] Since the light field technology was proposed in the 1990s of the 20th century, with the development of computational optics, the performance of light field cameras has been significantly enhanced. Such cameras use a special microlens array structure and can record the position and direction of light rays simultaneously, thus extending the traditional two-dimensional image to four dimensions. This characteristic makes light field cameras become core devices in many scientific and technological fields and has been widely applied in fields such as remote sensing sky surveys, medical equipment, and mobile terminals.

[0060] However, the autofocus technology of light field cameras faces severe challenges. Due to the high-dimensional characteristics of light field data, traditional camera autofocus methods, such as active and passive methods, are not fully applicable to light field cameras. The active method requires auxiliary equipment to calculate the depth of the target area, while the passive method relies on the clarity of the already imaged image to iteratively find the natural image plane. However, the light field image data collected by light field cameras cannot be directly iterated in the traditional passive way of calculating clarity, and the mechanical iteration method has a time consumption that is exponentially related to the scene depth, resulting in low focusing efficiency. The characteristic of the extended depth of field of light field cameras greatly increases the iteration time consumption.

[0061] To solve this problem, researchers have proposed various methods for autofocusing of light fields. For example, autofocus is achieved by calculating depth information, but this method requires a large amount of computation, high computing power, and long time consumption. Another method is to model the point spread function and train and learn the blur kernels for different scenes, but this method faces the problem of weak generalization ability and requires re-modeling. In addition, there are methods such as reconstructing image sequences according to different focusing parameters, calculating sharpness, and selecting focusing parameters, as well as methods for calculating the focus degree of the refocus image at different depths. However, these methods all have problems such as long processes, large amounts of computation, or high costs.

[0062] Based on this, the embodiments of the present application provide a fast autofocus method, device, equipment, and medium based on a light field camera, which will be described below through embodiments.

[0063] Figure 1 The flowchart of a fast autofocus method based on a light field camera provided by the embodiments of the present application is shown. The method acts on a target light field camera, and the target light field camera includes an optical system part and a light field camera part. The light field camera part includes a microlens array and a sensor. Among them, the optical system part is spaced from the light field camera part by a first distance; the target light field camera is used to capture a first image based on the first distance; the first image includes a plurality of first sub-images; wherein, the method includes steps S101 - S103; specifically:

[0064] S101. Reconstruct the spatial and angular information of the first sub-images according to the arrangement mode of the microlens array to obtain a first image set;

[0065] S102. Screen out a first image to be processed according to the information richness corresponding to each first image in the first image set;

[0066] S103. Determine whether the first distance is the optimal focusing distance according to the angular information fusion degree and sharpness of the first image to be processed within a preset focusing window; wherein, the optimal focusing distance is used for light field reconstruction.

[0067] The present application can solve the problems of long time consumption, large amount of computation, high computing power required, etc. in the light field focusing process.

[0068] Some embodiments of the present application will be described in detail below. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0069] For ease of description, in the embodiments of the present application, the light field camera that needs to perform focus processing is referred to as the target light field camera. The target light field camera includes an optical system part and a light field camera part. The light field camera part includes a microlens array and a sensor. The microlens array is disposed between the optical system and the sensor. Among them, the optical system part is spaced apart from the light field camera part by a first distance, and the sensor is spaced apart from the microlens array by a second distance. It should be noted that in the present application, the sensor and the microlens array are adjusted for focus as a whole (a single entity) of the light field camera part, that is, the second distance between the sensor and the microlens array in the present application does not change (or the second distance between the sensor and the microlens array is adjusted after the method of the present application is executed). The embodiments of the present application refer to the image captured based on the first distance as the first image. The focusing process of the target light field camera is to record all the information of the light rays, including position, direction, intensity, etc., when the first image is captured, and then select the focus position.

[0070] The first image is as Figure 2 shown. The first image includes a plurality of first sub-images. Since the first image is captured by adding a microlens array between the imaging sensor and the optical system. Such a microlens array can separate the light rays transmitted by the main lens in different directions, and thus capture the four-dimensional light field. This capture method enables the light field camera to record all the information of the light rays, such as position, direction, intensity, etc., rather than just the two-dimensional image information recorded by a traditional camera. When the first image is magnified locally, due to the existence of the microlens unit structure, the true details may not be visible at the original scale. This is because the microlens array divides the light rays into multiple directions, and the light rays in each direction are captured by different microlenses and form corresponding first sub-images on the sensor. Therefore, in the magnified image, a pixelation effect generated by the microlens array will be seen. In order to extract different angular information and spatial information from the light field data, it is necessary to reconstruct the four-dimensional light field information according to the arrangement of the microlens array.

[0071] When performing the reconstruction, it is also necessary to consider the influence of the arrangement of the microlens array. Specifically, the arrangement of the microlens array is square as Figure 3a , Figure 3b , Figure 3c , Figure 3d , Figure 3e , Figure 3f , Figure 3g , Figure 3hAs shown, it includes shapes such as circles, squares, ellipses, hexagons, etc., and can also be composed of lens combinations of various specifications. For microlens arrays with different arrangements, there are certain differences in the reconstruction methods. For example, for the square arrangement, the reconstruction process is to sort in the order from left to right and from top to bottom, and sequentially extract the first pixel, second pixel, etc. of each first sub-image, as Figure 4 shown. After reconstruction, a first image set is obtained, as Figure 5 shown. Each image in the first image set represents image information at different angles and depth information in space.

[0072] In the real data of four-dimensional light field information recording, autofocus can be achieved through the fusion of light rays at different angles and clarity analysis. The imaging of light rays in a certain direction or multiple directions on the sensor can be selected for clarity analysis. By evaluating the clarity of these images, it is determined which position has the clearest imaging, thus completing the autofocus process.

[0073] When screening images in a certain direction or multiple directions, the screening basis in the embodiments of the present application is the information richness corresponding to each first image in the first image set. Here, the information richness represents the quantity and quality of information contained in the first image. In specific implementation, corresponding quantity weights and quality weights can be set for the quantity and quality of information respectively. For example, the quantity weight is set to 1 and the quality weight is set to 1.9. The information richness of the first image can be obtained by performing weighted calculation on the quantity and quality of information contained in the first image through the first image.

[0074] If the information richness is large enough, two images can be selected. If the information richness is not large enough, a larger number of images can be selected. The determination method of whether the information richness is sufficient here is to compare the information richness of the first image with a preset range: if the information richness corresponding to the first image is within the preset first range, a first number of images are screened out as the first images to be processed; if the information richness corresponding to the first image is within the preset second range, a second number of images are screened out as the first images to be processed; if the information richness corresponding to the first image is within the preset third range, a third number of images are screened out as the first images to be processed. Among them, the upper limit value of the first range is less than or equal to the lower limit value of the second range, and the upper limit value of the second range is less than or equal to the lower limit value of the third range; the first number is less than the second number, and the second number is less than the third number. The above different numbers represent different directions. The imaging of light rays in multiple directions on the sensor is selected to judge the fusion degree between different angles. By comparing the fusion degree of light ray imaging at different angles, the general depth of the region of interest can be estimated, which helps to determine the approximate position of the image formed by the optical system at the natural image plane.

[0075] In order to improve efficiency, when determining the degree of fusion of angle information, the embodiment of the present application needs to first determine the focus window, and then calculate the degree of fusion of angle information for the first image to be processed under the focus window. The process of determining the focus window is relatively flexible, and the focus window can be determined before the first image set, or after the first image set is obtained.

[0076] Specifically, the process of determining the focus window in the embodiment of the present application includes setting a first weight for each target area according to the fullness of the texture information in the preset target area; performing a scaling process with the target area with the largest first weight as the center, and calculating the second weight of each target area after scaling; screening the focus window according to the target weight and the density of the texture information; wherein the target weight is calculated by the first weight and the second weight.

[0077] The target area is a specific position of the light field camera. In natural photography scenes, people usually focus on the center area of ​​the image, the focus of the rule of thirds, or other areas manually selected according to composition needs. In order to realize the automatic selection of the focus window, the embodiment of the present application designs a dynamic weight selection method based on 13 targets: Figure 6 As shown, 13 fixed target positions are preset in the image, which cover common compositional focal points and possible areas of interest. Calculate the fullness of texture information in each target area. This can be achieved by analyzing the grayscale changes, edge density, and detail richness in the area. Assign the first weight: assign a different first weight W to each target area according to the fullness of the texture information. The higher the weight, the more likely the area is to be the focus of user attention. Scale the target area and recalculate: scale the target area with the highest first weight as the center to adapt to the size of the focus in different scenes. Recalculate the fullness of texture information in the scaled area and assign the second weight W'. Dynamic weight weighting: weight the first weight W and the second weight W' to obtain the target weight of each target area. Filter the focus window: filter out the area with the largest and densest target weight as the focus window for subsequent use. This focus window will be used in the subsequent natural image plane search and light field reconstruction focus steps.

[0078] Design advantages: Dynamic balance of the attention area: Through dynamic weight selection, it can automatically balance the attention area in different scenarios, ensuring that the focus window can accurately cover the key points that the user is interested in. Control the window size: By scaling the target area and recalculating the weights, it can flexibly control the size of the focus window, reduce unnecessary calculation areas, and thus greatly compress the focusing time. Cover the main details of interest: Discrete area selection can cover the main details of interest in the image, ensuring the accuracy of focusing. Eliminate the influence of depth changes: Through dynamic weight selection and window adjustment, it can effectively eliminate the influence of target defocus and loss caused by depth changes, and improve the stability and reliability of focusing. In summary, the dynamic weight selection method based on 13 targets can realize the automatic selection of the focus window. Through design advantages such as dynamically balancing the attention area, controlling the window size, covering the main details of interest, and eliminating the influence of depth changes, it meets the requirements of focusing accuracy.

[0079] The focus window can also be determined by a pre-trained model: The focus window can be efficiently and stably selected according to certain prior information. These prior information include the type of shooting scene, the characteristics of the object being photographed, the lighting conditions, etc. By using this prior information, it is possible to more accurately predict which areas may contain important focusing targets, thereby optimizing the selection of the focus window.

[0080] In addition, with the development of deep learning technology, algorithms such as YOLO (You Only Look Once) can be introduced for real-time focus window selection. YOLO is an advanced object detection algorithm that can quickly and accurately identify the target area in an image in real-time applications. By applying YOLO to the autofocus system, the target area can be initially determined and used as the focus window. This method not only improves the accuracy and efficiency of focusing, but also enhances the adaptive ability of the system, enabling it to perform well in different shooting scenarios.

[0081] After determining the focus window, calculate the angle information fusion degree of the first image to be processed within the focus window. When calculating the angle information fusion degree, the embodiments of the present application can use different calculation methods (for example, including the first calculation method and the second calculation method). Different calculation methods can be used alone or in combination.

[0082] For example, the first calculation method is: The gradient descent method will be used to optimize the angle fusion degree and predict the focus position. Using the gradient descent method as the optimization strategy aims to find the minimum value of the objective function through iteration. The specific formula is as follows:

[0083] Θ1 = Θ0 – α▽J(Θ)

[0084] Where: Θ represents the parameter vector; Θ0 is the initial value of the parameter; Θ1 is the updated parameter value; α is the learning rate or step size, which controls the amplitude of parameter update; is the gradient of the objective function J(Θ) with respect to Θ.

[0085] To accurately evaluate the fusion degree between different perspectives, the embodiments of the present application use MSE (Mean Squared Error) and SSIM (Structural Similarity Index) in combination as the similarity measurement methods.

[0086] MSE: Calculate the average of the sum of squares of the pixel differences between the predicted image and the target image, which reflects the overall difference between the two:

[0087]

[0088] SSIM: Measure the similarity between the predicted image and the target image from three aspects: brightness, contrast, and structure, which is closer to the human visual perception:

[0089] SSIM(x,y) = ([l(x,y)] α )([c(x,y)] β )([s(x,y)] γ )

[0090] Calculate the angle fusion degree: Use MSE and SSIM to calculate the similarity of the images formed by the light rays in different directions on the sensor as a measure of the angle fusion degree (angle information fusion factor). Optimize the angle fusion degree through the gradient descent method, and iteratively calculate the transformation matrix to make the images between different perspectives more consistent.

[0091] The second calculation method: First extract the feature points (such as corner points, edge points, etc.) in the first image to be processed, and use the matching degree between the feature points as the angle information fusion factor. The higher the matching degree, the higher the similarity between the two images and the better the fusion degree.

[0092] The fusion of the first calculation method and the second calculation method: Use the intersection of the first calculation result obtained by the first calculation method and the second calculation result obtained by the second calculation method as the final fusion result.

[0093] Judge the natural image plane according to the angle information fusion degree: When the imaging similarity of the light rays in different directions on the sensor reaches the highest (i.e., the angle fusion degree is the largest), it is considered that these light rays all come from the same object point, and this object point is located on the natural image plane. Use the gradient descent method to further fine-tune the parameters to ensure that the found natural image plane is optimal.

[0094] Considering that the resolution result of the light field reconstruction of the natural image plane is not optimal, after finding the uniquely determined natural image plane, while keeping the second distance between the microlens array and the sensor unchanged, the microlens array and the sensor are moved as a whole in the direction closer to the optical system (to obtain the third distance, the fourth distance, etc. between the optical system and the sensor, which are determined according to the number of movements). After the movement, repeat the above method to take pictures again, obtain the third image, the fourth image, etc. at the third distance, and determine the sharpness of the third image to be processed, the sharpness of the fourth image to be processed, etc., until the optimal sharpness is determined. After determining the sharpness, the distance corresponding to the optimal sharpness is used as the optimal focusing distance, and the light field reconstruction is performed based on the optimal focusing distance.

[0095] In the embodiment of the present application, when calculating the sharpness of the first image to be processed, the embodiment of the present application calculates the sharpness of the first image to be processed according to the high-frequency details of the first image to be processed. When determining the high-frequency details of the first image to be processed, the embodiment of the present application can use different methods: using the Laplace operator or using the DCT wavelet transform.

[0096] Specifically, using the Laplace operator: Laplace filter:

[0097]

[0098] α is the shape operator, α ∈ [0, 1]. Convolve the spatial information image I(x, y) within any window with the Laplace filter mask to obtain the texture information IL(x, y) of the input value, that is, the information with rapid gray-scale changes. Usually, the sharpness is evaluated by calculating the mean IL-mean(x, y) or the energy entropy IL-entropy(x, y) of the result after convolution. However, there is a problem with this method, that is, there is a large flat area and slow change in the focusing curve, resulting in the inability to quickly and accurately find the focusing position. The embodiment of the present application notices that in the image after convolution, the area with slow gray-scale change usually corresponds to the background, while the area with rapid gray-scale change corresponds to the texture. When the focus is clear, the texture is rich, so it becomes very crucial to count the number of textures in the entire window area. The embodiment of the present application evaluates the sharpness by statistically analyzing the discreteness D(I L (x, y)) of the high-frequency textures in the image IL(x, y) after convolution with the Laplace filter. Specifically, within the statistically analyzed window area, if the original input texture is clear and the gray-scale changes rapidly, then there will be more high-frequency components, and the high-frequency will be more discrete compared to the low-frequency:

[0099] D(I L (x,y))=mean[I L (x,y)-mean(I L(x,y))] 2

[0100] This method can more accurately reflect the clarity of the image and can greatly improve the problem of the flat and slow focusing curve, enabling rapid autofocus to be achieved in a short time.

[0101] Use DCT wavelet transform: First, perform a discrete cosine transform (DCT) on the first image to be processed. DCT is a transformation method commonly used in image compression and feature extraction, which can transform the image from the spatial domain to the frequency domain. After the DCT transformation, the high-frequency components of the image usually correspond to details such as edges and textures in the image, while the low-frequency components correspond to the overall brightness or average gray value of the image.

[0102] Next, perform a wavelet transform on the image after the DCT transformation. The wavelet transform is a multi-scale analysis method that can decompose the image into components of different scales and frequencies. Through the wavelet transform, we can further analyze the detail information of the image at different scales, thereby more accurately evaluating the clarity of the image.

[0103] During the analysis process, focus on the degree of dispersion of the high-frequency components. If there are more high-frequency components in the image and the high-frequency is more dispersed compared to the low-frequency, it usually means that there are more details in the image, and thus the clarity of the image is higher.

[0104] Finally, quantitatively evaluate the clarity of the image based on the degree of dispersion of the high-frequency components, and perform image processing or analysis tasks accordingly.

[0105] The autofocus method of this application directly calculates the four-dimensional light field data, can quickly perform depth positioning and clarity evaluation calculations, has a significant distinction for different focusing situations, and at the same time estimates the best position for potential reconstruction in real time to obtain the focusing position that makes the reconstructed resolution the highest. It has a high tolerance for noise interference and strong adaptability to aberrations in different environments.

[0106] Figure 7 The structural schematic diagram of a fast autofocus device provided by an embodiment of this application is shown, which acts on a target light field camera. The target light field camera includes an optical system part and a light field camera part. The light field camera part includes a microlens array and a sensor. Among them, the optical system part is spaced apart from the light field camera part by a first distance; the target light field camera is used to capture a first image based on the first distance; the first image includes a plurality of first sub-images, and the device includes:

[0107] A reconstruction module, configured to reconstruct the spatial and angular information of the first sub-image according to the arrangement mode of the microlens array, so as to obtain a first image set;

[0108] A screening module, configured to screen out a first image to be processed according to the information richness corresponding to each first image in the first image set;

[0109] A determination module, configured to determine whether the first distance is the optimal focusing distance according to the angular information fusion degree and clarity of the first image to be processed within a preset focusing window; wherein, the optimal focusing distance is used for light field reconstruction.

[0110] If the first distance is not the optimal focusing distance, the apparatus further includes: an adjustment module, configured to adjust the distance between the optical system part and the light field camera part until the optimal focusing distance is obtained.

[0111] Before obtaining the first image set, the method further includes: determining the focusing window; or, after obtaining the first image set, it further includes: determining the focusing window;

[0112] The focusing window is determined by the following method:

[0113] According to the fullness of the texture information within a preset target area, set a first weight for each of the target areas;

[0114] Perform a scaling process with the target area having the largest first weight as the center, and calculate the second weights of each of the target areas after scaling;

[0115] According to the target weight and the density of the texture information, screen out the focusing window; wherein, the target weight is calculated from the first weight and the second weight;

[0116] Alternatively, the focusing window is determined by a pre-trained model.

[0117] The screening out a first image to be processed according to the information richness corresponding to each first image in the first image set includes:

[0118] If the information richness corresponding to the first image is within a preset first range, screen out a first number of images as the first image to be processed;

[0119] If the information richness corresponding to the first image is within a preset second range, screen out a second number of images as the first image to be processed;

[0120] If the information richness corresponding to the first image is within a preset third range, screen out a third number of images as the first image to be processed.

[0121] Calculate the angular information fusion degree of the first image to be processed in the following manner:

[0122] Calculate the angular information fusion factor between the first images to be processed;

[0123] Calculate the angular information fusion degree of the first images to be processed according to the angular information fusion factor.

[0124] The calculation of the angular information fusion factor between the first images to be processed includes:

[0125] Calculate the mean square error and the structural similarity index between the first images to be processed;

[0126] The calculation of the angular information fusion degree of the first images to be processed according to the angular information fusion factor includes:

[0127] Calculate the angular information fusion degree of the first images to be processed according to the mean square error and the structural similarity index.

[0128] Calculate the sharpness of the first images to be processed in the following manner:

[0129] Calculate the sharpness of the first images to be processed according to the dispersion degree of the high-frequency components of the first images to be processed.

[0130] As Figure 8 shown, an embodiment of the present application provides an electronic device for executing the fast autofocus method based on a light field camera in the present application. The device includes a memory, a processor, a bus, and a computer program stored on the memory and executable on the processor. Among them, when the above-mentioned processor executes the above-mentioned computer program, the steps of the above-mentioned fast autofocus method based on a light field camera are implemented.

[0131] Specifically, the above-mentioned memory and processor can be general memory and processor, and no specific limitation is made here. When the processor runs the computer program stored in the memory, it can execute the above-mentioned fast autofocus method based on a light field camera.

[0132] Corresponding to the fast autofocus method based on a light field camera in the present application, an embodiment of the present application also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the above-mentioned fast autofocus method based on a light field camera are executed.

[0133] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is run, it can execute the above-mentioned fast autofocus method based on a light field camera.

[0134] In the embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the system or unit can be in electrical, mechanical or other forms.

[0135] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0136] In addition, each functional unit in the embodiments provided in the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0137] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0138] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0139] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. All should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A fast autofocus method based on a light field camera, characterized in that: Acting on a target light field camera, the target light field camera includes an optical system part and a light field camera part, the light field camera part includes a microlens array and a sensor, wherein the optical system part is separated from the light field camera part by a first distance; the target light field camera is used to capture a first image based on the first distance; the first image includes a plurality of first sub-images; the method includes: Reconstructing the spatial and angular information of the first sub-image according to the arrangement of the microlens array to obtain a first image set; Screening out first images to be processed according to information richness corresponding to each first image in the first image set; According to the fusion degree and clarity of the angle information of the first image to be processed in a preset focus window, it is determined whether the first distance is an optimal focus distance; wherein the optimal focus distance is used for light field reconstruction.

2. The method according to claim 1, characterized in that: If the first distance is not the optimal focusing distance, the method further includes: The distance between the optical system part and the light field camera part is adjusted until the optimal focusing distance is obtained.

3. The method according to claim 1, characterized in that Before obtaining the first image set, the method further includes: determining the focus window; or after obtaining the first image set, the method further includes: determining the focus window; The method determines the focus window by: According to the fullness of texture information in the preset target area, a first weight is set for each of the target areas; Performing a scaling process with the target area having the largest first weight as the center, and calculating a second weight of each of the target areas after scaling; The focus window is selected according to the target weight and the density of the texture information; wherein the target weight is calculated by the first weight and the second weight; Alternatively, the focus window is determined by a pre-trained model.

4. The method according to claim 1, characterized in that The step of selecting first images to be processed according to the information richness corresponding to each first image in the first image set includes: If the information richness corresponding to the first image is within a preset first range, selecting a first number of images as the first images to be processed; If the information richness corresponding to the first image is within a preset second range, selecting a second number of images as the first images to be processed; If the information richness corresponding to the first image is within a preset third range, a third number of images are screened out as the first images to be processed.

5. The method according to claim 1, characterized in that: The method calculates the fusion degree of angle information of the first image to be processed in the following manner: Calculating an angle information fusion factor between the first images to be processed; The angle information fusion degree of the first image to be processed is calculated according to the angle information fusion factor.

6. The method according to claim 5, characterized in that The calculating the angle information fusion factor between the first to-be-processed images includes: Calculating the mean square error and the structural similarity index between the first to-be-processed images; The calculating the angle information fusion degree of the first to-be-processed image according to the angle information fusion factor includes: The angle information fusion degree of the first image to be processed is calculated according to the mean square error and the structural similarity index.

7. The method according to claim 1, characterized in that The method calculates the clarity of the first to-be-processed image in the following manner: The clarity of the first image to be processed is calculated according to the discreteness of the high-frequency component of the first image to be processed.

8. A fast autofocus device based on a light field camera, characterized in that: Acting on a target light field camera, the target light field camera includes an optical system part and a light field camera part, the light field camera part includes a microlens array and a sensor, wherein the optical system part is separated from the light field camera part by a first distance; the target light field camera is used to obtain a first image based on the first distance; the first image includes a plurality of first sub-images, and the device includes: A reconstruction module, configured to reconstruct the spatial and angular information of the first sub-image according to the arrangement of the microlens array to obtain a first image set; A screening module, configured to screen out first images to be processed according to information richness corresponding to each first image in the first image set; A determination module is used to determine whether the first distance is an optimal focus distance according to the fusion degree and clarity of the angle information of the first image to be processed in a preset focus window; wherein the optimal focus distance is used for light field reconstruction.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the fast autofocus method based on a light field camera are performed.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the fast auto-focusing method based on a light field camera as claimed in any one of claims 1 to 7 are executed.