A method, system, device, and storage medium for generating a dense focus stack image
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2026-08-11
AI Technical Summary
[0039] (1) The method for generating dense focal stack images of the present invention involves numbering and dividing a set of focal stack images sparse in the focal length domain according to focal length, dividing blocks with the same coordinates into a block group, and further dividing each block group into block pairs. For each block pair, one block is used as a reference block to fit another block to solve for the best fitting parameter pair. The best fitting parameter pair and the corresponding reference block are used to obtain the fitting parameter pair and reference block of the new focal length image to be predicted, thereby obtaining the prediction block. The prediction blocks corresponding to the block pairs with the same focal length number are stitched together according to the coordinates of the corresponding block pairs to generate the predicted new focal length image, wherein the focal length of the new focal length image is located between the focal lengths of the two blocks in the corresponding block pair. The present invention increases the number of different focal lengths by generating images of unknown focal lengths between sparse focal lengths, thereby obtaining a dense focal stack image. It realizes the generation of a dense focal stack image with a small number of focal stack images (i.e., a set of focal stack images sparse in the focal length domain), which can avoid the defects of hardware requirements for physical methods of shooting dense focal stack images, and reduce the data storage and equipment performance requirements.
Smart Images

Figure CN117670697B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, and more specifically, relates to a method, system, device and storage medium for generating dense focal stack images. Background Technology
[0002] The focus change and depth information contained in focus stack images have important applications in imaging and display across various fields. For example, in the display field, the foreground portion of a scene image can be extracted using the depth information of the focus stack image. By applying Gaussian blurring to different depths of the scene, a visually appealing background blurring effect can be generated. In the field of stereoscopic display, focus stack images can be organized sequentially into a coherent video format, achieving a visual panoramic depth scan simply by playing the video. In the field of immersive multimedia, focus stack images can be used to generate views that better match the human eye's focusing and blurring experience.
[0003] Increasing the density of the focus stack image can improve the smoothness and fluidity of video, thereby enhancing the viewing experience. Current technologies typically capture dense focus stack images using a camera; however, capturing dense focus stack images requires extremely small focal intervals to generate a large number of focus stack images, which places high demands on the camera's performance and data storage capacity. Summary of the Invention
[0004] In view of the shortcomings of the existing technology and the need for improvement, the present invention provides a method, system, device and storage medium for generating dense focal stack images, the purpose of which is to reduce the amount of data storage and the requirements for device performance.
[0005] To achieve the above objectives, according to a first aspect of the present invention, a method for generating a dense focal stack image is provided, comprising:
[0006] S1. Number a set of focal stack images with sparse focal length domain according to focal length.
[0007] S2. In the same coordinate system, divide the numbered focal stack image into blocks, and divide the blocks with the same coordinates into a block group. Two blocks in a block group form a block pair, and the two blocks in a block pair are the first block and the second block, respectively.
[0008] S3. Using the first block as the reference block and the second block as the current block, the corresponding fitting block and fitting parameter pair (m, σ) are obtained by performing mode transformation on the reference block. The fitting parameter pair with the smallest error between the fitting block and the current block is selected as the first fitting parameter pair (m0, σ0) of the current block pair. m represents the mode corresponding to the mode transformation and σ represents the parameter of the corresponding mode transformation.
[0009] S4. Use the first fitted parameter pair (m0, σ0) as the best fitted parameter pair to generate the predicted parameter pair (m p ,σ p ), and use the prediction parameters to (m p ,σ p The reference block is subjected to a corresponding mode transformation to obtain the prediction block;
[0010] S5. The predicted blocks corresponding to the block pairs with the same focal length number are stitched together according to the coordinates of the corresponding block pairs to obtain the predicted new focal length image; the new focal length image and the sparse focal stack image in the focal length domain constitute a dense focal stack image.
[0011] Furthermore, S3 also includes: taking the second block as the reference block and the first block as the current block to obtain the second fitting parameter pair (m1,σ1) of the current block pair;
[0012] S4 further includes: selecting the pair with the smaller error between the first fitting parameter pair (m0, σ0) and the second fitting parameter pair (m1, σ1) as the optimal fitting parameter pair.
[0013] Furthermore, in S3, the mode transformation includes one or more of the following: copy transformation, Gaussian transformation, or Wiener transformation;
[0014] The copy transformation involves using the reference block as the fitting block;
[0015] The Gaussian transform includes: performing a convolution operation between the reference block and a two-dimensional Gaussian function to obtain the fitting block;
[0016] The Wiener transform includes: multiplying the reference block with Wiener deconvolution in the frequency domain and then transforming it to the spatial domain to obtain the fitted block.
[0017] Furthermore, in S4, the prediction parameter pair (m p ,σ p The parameter σ of the mode transformation in ) p The average of the mode transformation parameters in the best-fit parameter pair is taken.
[0018] Furthermore, in S4, if m p If the mode is copy mode, then the corresponding mode transformation is a copy transformation;
[0019] If m p If the mode is Gaussian, then the corresponding mode transformation is a Gaussian transform;
[0020] If m p If it is a Wiener mode, then the corresponding mode transformation is the Wiener transform.
[0021] Further, in S5, the new focal length image is numbered according to the focal length numbers corresponding to the two blocks in the block pair;
[0022] The newly numbered focal length image and the sparse focal stack image in the focal length domain are integrated into the dense focal stack image according to the focal length number.
[0023] Furthermore, S1 also includes: selecting a portion of the focal stack image as the ground truth image;
[0024] S5 is followed by: calculating the peak signal-to-noise ratio between the new focal length image and the corresponding ground truth image with the same focal length, in order to verify the effect of the new focal length image.
[0025] According to a second aspect of the invention, a system for generating densely stacked images is provided for performing the method according to any one of the first aspects, the system comprising:
[0026] Focal length numbering unit is used to number a set of focal length stack images that are sparse in the focal length domain according to the focal length size.
[0027] The block pair division unit is used to divide the numbered focal stack image into blocks in the same coordinate system, and divide the blocks with the same coordinates into a block group. Two blocks in a block group form a block pair, and the two blocks in a block pair are the first block and the second block, respectively.
[0028] The mode transformation unit is used to take the first block as a reference block and the second block as the current block, and to obtain the corresponding fitting block and fitting parameter pair (m,σ) by performing mode transformation on the reference block; m represents the mode corresponding to the mode transformation, and σ represents the parameter of the corresponding mode transformation.
[0029] The first fitting parameter pair calculation unit is used to select the fitting parameter pair with the smallest error between the fitting block and the current block as the first fitting parameter pair (m0,σ0) of the current block pair;
[0030] The prediction block generation unit is used to generate prediction parameter pairs (m0, σ0) using the first fitted parameter pair (m0, σ0) as the best fitted parameter pair. p ,σ p ), and use the prediction parameters to (m p ,σ p The reference block is subjected to a corresponding mode transformation to obtain the prediction block;
[0031] The stitching unit is used to stitch together the prediction blocks corresponding to the block pairs with the same focal length number according to the coordinates of the corresponding block pairs to obtain a new predicted focal length image; the new focal length image and the sparse focal stack image in the focal length domain constitute a dense focal stack image.
[0032] Furthermore, it also includes a second fitting parameter pair calculation unit, used to take the second block as a reference block and the first block as the current block to obtain the second fitting parameter pair (m1,σ1) of the current block pair;
[0033] The prediction block generation unit further includes: selecting the pair with the smaller error between the first fitting parameter pair (m0, σ0) and the second fitting parameter pair (m1, σ1) as the optimal fitting parameter pair.
[0034] According to a third aspect of the present invention, an electronic device is provided, including a computer-readable storage medium and a processor;
[0035] The computer-readable storage medium is used to store executable instructions;
[0036] The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method described in any one of the first aspects;
[0037] Or / and, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method as described in any of the first aspects.
[0038] In summary, the above-described technical solutions conceived in this invention can achieve the following beneficial effects:
[0039] (1) The method for generating dense focal stack images of the present invention involves numbering and dividing a set of focal stack images sparse in the focal length domain according to focal length, dividing blocks with the same coordinates into a block group, and further dividing each block group into block pairs. For each block pair, one block is used as a reference block to fit another block to solve for the best fitting parameter pair. The best fitting parameter pair and the corresponding reference block are used to obtain the fitting parameter pair and reference block of the new focal length image to be predicted, thereby obtaining the prediction block. The prediction blocks corresponding to the block pairs with the same focal length number are stitched together according to the coordinates of the corresponding block pairs to generate the predicted new focal length image, wherein the focal length of the new focal length image is located between the focal lengths of the two blocks in the corresponding block pair. The present invention increases the number of different focal lengths by generating images of unknown focal lengths between sparse focal lengths, thereby obtaining a dense focal stack image. It realizes the generation of a dense focal stack image with a small number of focal stack images (i.e., a set of focal stack images sparse in the focal length domain), which can avoid the defects of hardware requirements for physical methods of shooting dense focal stack images, and reduce the data storage and equipment performance requirements.
[0040] (2) Further, two blocks in a block pair are selected alternately as the reference block and the current block. Through bidirectional prediction, the best fitting parameter pair is calculated by making full use of the information of the two blocks in the block pair. This can be closer to the trend of focal length change between the two blocks in the block pair, so that the focal length of the generated new focal stack image is closer to the middle state of the focal length of the two blocks. This makes the change of the focal length domain of the final dense focal stack image more uniform, which can further improve the visual effect.
[0041] (3) Preferably, considering that the sharpness, image defocus and image focus of some areas in the actual focal length change are constantly changing, the mode transformation selected by the present invention includes copy transformation, Gaussian transformation and Wiener transformation, which respectively correspond to the sharpness change, image defocus change and image focus change of different areas in the focal stack image corresponding to the current block pair. The fitting parameter pair with the smallest error is selected as the current first fitting parameter pair, so that the fitting block calculated by the reference block through the parameter pair is closest to the current block. Using the parameter pair as the best parameter pair can more accurately simulate the actual focal length change from the reference block to the current block.
[0042] (4) As a preferred option, the parameters of the best fitting parameters are averaged to obtain the parameters of the prediction parameters. This can make the focal length of the generated new focal stack image closer to the middle state of the two block focal lengths, thereby making the focal length domain of the final dense focal stack image more uniform and further improving the visual effect.
[0043] In summary, the dense focal stack image generated by this method has excellent visual effects, with very smooth and natural transitions between focal lengths; and it can reduce data storage requirements and device performance requirements. Attached Figure Description
[0044] Figure 1 This is a flowchart of the method for generating dense focal stack images according to the present invention.
[0045] Figure 2 This is a flowchart of how to divide a focal stack image with a sparse focal length domain into blocks and divide them into block pairs in Embodiment 1 of the present invention.
[0046] Figure 3 This is a flowchart of obtaining the corresponding fitting parameters and errors through mode transformation in Embodiment 1 of the present invention.
[0047] Figure 4 This is a flowchart illustrating how the reference block B is copied and transformed to obtain the corresponding fitting parameters and errors in Embodiment 1 of the present invention.
[0048] Figure 5 This is a flowchart illustrating how Gaussian transformation is performed on reference block B to obtain the corresponding fitting parameters and errors in Embodiment 1 of the present invention.
[0049] Figure 6 The flowchart in Embodiment 1 of this invention shows the process of performing Wiener transformation on reference block B to obtain the corresponding fitting parameters and errors.
[0050] Figure 7 This is a flowchart of bidirectional prediction in Embodiment 1 of the present invention.
[0051] Figure 8 This is a flowchart of generating a prediction block in Embodiment 1 of the present invention.
[0052] Figure 9 This is a flowchart of generating corresponding prediction blocks and stitching them together to form a corresponding new focal length image for n block pairs with the same number in Embodiment 1 of the present invention.
[0053] Figure 10 This is a flowchart of Embodiment 1 of the present invention, in which the prediction blocks corresponding to n block pairs with the same focal length number are stitched together according to the coordinates of the corresponding block pairs to obtain the predicted new focal length image.
[0054] Figure 11 This is the verification process in Embodiment 1 of the present invention. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0056] In this invention, the terms "first," "second," etc., used in the invention and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0057] like Figure 1 As shown, the method for generating dense focal stack images according to the present invention mainly includes the following steps:
[0058] S1. Number a set of focal stack images with sparse focal length domain according to focal length.
[0059] S2. In the same coordinate system, divide the numbered focal stack image into blocks, divide the blocks with the same coordinates into a block group, and divide the same block group into block pairs. Each block pair includes two blocks with different numbers. Let one block in each block pair be the first block and the other block be the second block.
[0060] S3. Take the first block in each block pair as the reference block B and the second block as the current block I. By performing mode transformation on the reference block B, obtain the corresponding fitting block F and the fitting parameter pair (m, σ). Select the fitting parameter pair with the smallest error between the fitting block F and the current block I as the first fitting parameter pair (m0, σ0) of the current block pair. In this embodiment of the invention, the mode transformation includes one or more of the following: copy transformation, Gaussian transformation, or Wiener transformation. In the fitting parameter pair, m is the selected mode, for example, m represents the copy mode, Gaussian mode, or Wiener mode, and σ is the standard deviation of the two-dimensional Gaussian function corresponding to the mode. In other embodiments, other mode transformations can also be used, where σ is the parameter corresponding to the selected mode transformation.
[0061] S4. Use the first fitted parameter pair as the best fitted parameter pair to generate the predicted parameter pair (m p ,σ p The best-fit parameters are used to determine the reference block B as the reference block B for the new focal length image to be predicted. p By predicting the parameter pair (m) p ,σ p For reference block B p Perform the corresponding mode transformation to obtain the prediction block P; where the corresponding mode transformation generates the mode corresponding to the best-fit parameter pair, that is, the mode corresponding to parameter m in the best-fit parameter pair, or the prediction parameter pair (m p ,σ p ) parameter m p The corresponding mode;
[0062] S5. The predicted blocks corresponding to the block pairs with the same focal length number are stitched together according to the coordinates of the corresponding block pairs to obtain the predicted new focal length image; wherein, the predicted new focal length image is integrated with the original set of sparse focal stack images in the focal length domain to form a dense focal stack image.
[0063] Specifically, in S1, the input set of focal length sparse stack images have the same resolution and the same shooting viewpoint. Here, focal length sparse means that the focal length difference between adjacent focal stack images exceeds a set threshold, which can be determined according to actual needs.
[0064] In this embodiment of the invention, the focal stack images with sparse focal length domains are numbered 0, 1, ..., k according to their focal length, where k represents the focal length number corresponding to the (k+1)th image. In other embodiments, other numbering methods can also be selected to distinguish focal stack images with different focal lengths.
[0065] Specifically, in S2, such as Figure 2As shown, the focal stack images with sparse focal length domains are divided into blocks in the same coordinate system. Preferably, each focal stack image block is the same size, and the block division method of the group of focal stack images is the same, so that the focal length of the final generated new focal stack image is more uniform. In other embodiments, the size of each focal stack image block can also be different.
[0066] When the resolution of the focal stack image is not an integer multiple of the block size, edge padding can be used to make the blocks meet the set block size. Specifically, in this embodiment of the invention, a set of focal stack images is selected, containing 11 different focal lengths, numbered from 0 to 10. The resolution of each focal stack image is 1920*1350, and the block size is set to 64. Since 1350 is not divisible by 64, edge padding is performed.
[0067] Blocks with the same coordinates are divided into a block group. Assuming that a focal stack image can be divided into n blocks, then the focal stack image with sparse focal length domain can be divided into n block groups. In the embodiment of the present invention, the number of each block is the same as the number of the corresponding focal stack image. That is, in each block group, block pairs are divided according to the focal length number. (0,2), (2,4), (4,6), (6,8), (8,10) constitute the corresponding block pairs.
[0068] Within the same block group, block pairs are divided. A block pair contains blocks corresponding to two different focal stack images. In this embodiment of the invention, two blocks with adjacent focal length numbers form a block pair, such as (0, 1), (1, 2), ..., (k-1, k). In other embodiments, two blocks in the same block group can be arbitrarily made to form a block pair. For example, block pairs in the form of (0, 2), (1, 3), etc., can be set so that the generated new focal lengths are all within the focal lengths of the focal stack images.
[0069] Specifically, in S3, such as Figure 3 As shown, in this embodiment of the invention, preferred mode transformations include copy transformation, Gaussian transformation, and Wiener transformation, which respectively correspond to changes in sharpness, defocus, and focus in different regions of the current block pair's focus stack image. After performing copy transformation, Gaussian transformation, and Wiener transformation on the reference block B, the corresponding fitting parameter pairs (m, σ) are denoted as: (m c ,σ c ), (m g ,σ g ) and (m w ,σ w ); where the error between the corresponding fitted block F and the current block I is denoted as loss. c loss g and loss wFor the current block pair, among the three types of errors obtained, select the smallest error: loss0 = min[loss0]. c loss g loss w The corresponding fitting parameters are used as the current first fitting parameter pair (m0, σ0).
[0070] Specifically, such as Figure 4 As shown, performing a copy transformation on the current reference block B means directly using the current reference block B as the fitting block F, with the corresponding fitting parameter pair (m c ,σ c The mode m is (m, 0). In this embodiment of the invention, the mode m is 0. In other embodiments, it can also be set to other values, as long as the corresponding mode can be distinguished.
[0071] In this embodiment of the invention, m = 0 and σ = 0 during initialization; the fitting error loss is calculated using the error calculation formula, and the calculated error loss is used as the loss. c m c =m,σ c =σ, thus obtaining (m c ,σ c ) and the corresponding error loss c In this embodiment of the invention, mean squared error (MSE LOSS) is used, and the corresponding error calculation formula is as follows:
[0072]
[0073] Where loss represents the error calculated so far, M represents the block resolution, i.e. the block size, I represents the current block, C(B,σ) represents the fitted block obtained by the copy transformation, and B represents the reference block; in other embodiments, other error calculation formulas may also be used.
[0074] Specifically, such as Figure 5 As shown, performing a Gaussian transform on the current reference block B means convolving it with a two-dimensional Gaussian function to obtain the fitted block F. Initialization error loss. g Given the search range [L1, L2] and search step size s, for each search, let σ = σ + s, calculate the error between the fitted block and the current block under the current σ using the error calculation formula, and the corresponding error loss. Compare the loss with the current loss. g Size, if loss is less than loss g Then assign the loss value to loss. g The current σ is copied to σ gThe search continues until the set search range is reached, at which point the search stops, and the minimum error loss during the search process is taken as the final error loss. g And the σ corresponding to the minimum error is taken as the final σ. g In this embodiment of the invention, the search range is set to [0, 20], the search step size s = 0.1, and the loss... g =10000, m=1 and σ=0, the corresponding Gaussian transform formula is:
[0075]
[0076] Where (x,y) represents the coordinates of a pixel.
[0077] The fitted block F obtained after Gaussian transformation is:
[0078]
[0079] Where G(B,σ) represents the fitted block obtained by Gaussian transformation.
[0080] Specifically, such as Figure 6 As shown, performing a Wiener transform on the current reference block B involves multiplying it with Wiener deconvolution in the frequency domain and then transforming it to the spatial domain to obtain the fitted block F. Similar to the Gaussian transform process, the initialization error loss... w Given the search range [L1, L2] and search step size s, for each search, let σ = σ + s, calculate the error between the fitted block and the current block under the current σ using the error calculation formula, and the corresponding error loss. Compare the loss with the current loss. w Size, if loss is less than loss w Then assign the loss value to loss. w The current σ is copied to σ w The search continues until the set search range is reached, at which point the search stops, and the minimum error loss during the search process is taken as the final error loss. w And the σ corresponding to the minimum error is taken as the final σ. w In this embodiment of the invention, the search range is set to [0, 20], the search step size s = 0.1, and the loss... w =10000, m=2 and σ=0, the corresponding Wiener transformation formula is:
[0081]
[0082] Among them, H ω (u,v) represents the frequency domain representation of the Wiener transform, and SNR represents the signal-to-noise ratio of the reference block B.
[0083] The fitted block F obtained after Wiener transform is:
[0084]
[0085] Where W(B,σ) represents the fitted block obtained after Wiener transform. This indicates the inverse Fourier transform operation. This indicates the Fourier transform operation.
[0086] For each block pair, perform the corresponding mode transformation under the three modes mentioned above to obtain the corresponding fitting parameter pair (m). c ,σ c ), (m g ,σ g ) and (m w ,σ w ), and the corresponding error loss c loss g and loss w Select loss c loss g and loss w The fitting parameters corresponding to the minimum mean error are used as the current first fitting parameter pair (m0, σ0).
[0087] As a further design of the present invention, such as Figure 7 As shown, in S3, it also includes: taking the second block in each block as the reference block B and the first block as the current block I, and obtaining the corresponding fitting block F and fitting parameter pair (m,σ) by performing mode transformation on the reference block B, and selecting the fitting parameter pair with the smallest error between the fitting block F and the current block I as the current second fitting parameter pair (m1,σ1).
[0088] S4 also includes: selecting the pair with the smaller error between the first fitted parameter pair (m0, σ0) and the second fitted parameter pair (m1, σ1) as the current best fitted parameter pair to generate the predicted parameter pair (m p ,σ p ).
[0089] The second fitting parameter pair (m1,σ1) and the corresponding error loss1 are calculated in the same way as the first fitting parameter pair (m0,σ0) and the corresponding error loss0.
[0090] Specifically, in S4, the best-fit parameter pair is used to generate the prediction parameter pair (m). p ,σ p ),include:
[0091] After averaging the parameters σ corresponding to the mode transformation in the best-fit parameter pair, we obtain σ. p ; that is, t takes a positive value, σ = argmin[loss0, loss1]; m p This refers to the selected mode m in the best-fit parameter pair.
[0092] As a preferred setting, t=2 can make the focal length of the generated new focal stack image closer to the middle state of the two block focal lengths, thereby making the change in the focal length domain of the final dense focal stack image more uniform and further improving the visual effect.
[0093] It should be noted that the process of solving for the optimal parameter pairs between block pairs does not affect each other. Therefore, the processing of each block pair can be done serially or in parallel, depending on the actual application requirements.
[0094] Specifically, such as Figure 8 As shown, in this embodiment of the invention, the prediction parameter pair (m) p ,σ p For reference block B p Perform the corresponding mode transformation to obtain the prediction block, including:
[0095] If m p For replication mode, use the prediction parameter pair (m) p ,σ p For reference block B p Perform a copy transformation to obtain the prediction block P;
[0096] If m p If it is a Gaussian mode, then use the prediction parameter pair (m) p ,σ p For reference block B p Perform a Gaussian transform to obtain the prediction block P;
[0097] If m p For the Wiener mode, the prediction parameter pair (m) is used. p ,σ p For reference block B p Perform Wiener transform to obtain the prediction block P.
[0098] Specifically, such as Figure 9 and Figure 10 As shown in S5, for each block pair generating a prediction block, the prediction blocks corresponding to n block pairs with the same focal length number are stitched together according to the coordinates of the corresponding block pairs to obtain the predicted new focal length image; the focal length of the new focal length image is numbered according to the focal length numbers corresponding to the two blocks in a block pair. In this embodiment of the invention, the average of the focal length numbers corresponding to the two blocks in a block pair is used as the focal length number of the new focal length image. For example, the focal length number of the new focal length image corresponding to the two images with focal length numbers 0 and 2 is 1.
[0099] In this embodiment of the invention, the focal length numbers of the new focal length image and the numbers of the original set of sparse focal stack images in the focal length domain are sorted from smallest to largest to obtain a dense focal stack image.
[0100] As a further design of the present invention, S1 also includes: selecting a portion of the focal stack image as the ground truth image; wherein the ground truth image does not participate in the subsequent S2-S5 processing;
[0101] Following S5, the method also includes: calculating the peak signal-to-noise ratio (PSNR) between the generated new focal length image and the corresponding ground truth image with the same focal length; where a higher PSNR indicates a better effect.
[0102] In embodiments of the present invention, such as Figure 11 As shown in the figure, the system in Figure 11 refers to the system of the present invention for generating dense focal stack images. The input set of focal length-sparse focal stack images includes 11 different focal lengths, numbered from 0 to 10. Five odd-numbered images are extracted as ground truth images. The remaining six even-numbered images are processed using the method of the present invention to obtain five new focal length images, which are also numbered 1, 3, 5, 7, and 9. Then, the peak signal-to-noise ratio (PSNR) is calculated according to the corresponding labels of the five extracted ground truth images. The calculation results are shown in Table 1. It can be seen that the PSNR is around 42dB, which is very good. Generally, a PSNR exceeding 40 indicates extremely good image quality, very close to the original image.
[0103] Table 1. PSNR results of the new focal length image and the ground truth image.
[0104] PSNR (dB) 42.237 42.353 42.743 42.521 41.920
[0105] The method for generating dense focal stack images of the present invention involves dividing a set of sparse focal stack images in the focal length domain into blocks numbered according to focal length. Blocks with the same coordinates are grouped into a block group, and each block group is further divided into block pairs. For each block pair, one block is used as a reference block to fit the other block, solving for the optimal fitting parameter pair. The optimal fitting parameter pair and the corresponding reference block are used to obtain the fitting parameter pair and reference block for the new focal length image to be predicted, thus obtaining the prediction block. The prediction blocks corresponding to the block pairs with the same focal length number are stitched together according to the coordinates of the corresponding block pairs to generate the predicted new focal length image, wherein the focal length of the new focal length image is located between the focal lengths of the two blocks in the corresponding block pair. The present invention increases the number of different focal lengths by generating images of unknown focal lengths between sparse focal lengths, thereby obtaining a dense focal stack image. It realizes the generation of a dense focal stack image with a small number of focal stack images (i.e., a set of sparse focal stack images in the focal length domain), which can avoid the hardware requirements of physical methods for capturing dense focal stack images, and reduce the data storage and equipment performance requirements. The generated dense focal stack graph can be used to improve the visual effects of focal stack image-based visual applications.
[0106] Preferably, considering that the sharpness, image defocus, and image focus of some areas in the actual focal length change are constantly changing, the mode transformation selected in this invention includes copy transformation, Gaussian transformation, and Wiener transformation, which respectively correspond to the sharpness change, image defocus change, and image focus change of different areas in the focal stack image corresponding to the current block pair. The fitting parameter pair with the smallest error is selected as the current first fitting parameter pair, so that the fitting block calculated by the reference block B through the parameter pair is closest to the current block. Using the parameter pair as the optimal parameter pair can more accurately simulate the actual focal length change from the reference block B to the current block I.
[0107] Furthermore, by selecting two blocks in a block pair and using them alternately as the reference block B and the current block I, the best-fit parameter pair is calculated by making full use of the information of the two blocks in the block pair through bidirectional prediction. This allows the best-fit parameter pair to be closer to the trend of focal length change between the two blocks in the block pair, making the focal length of the generated new focal stack image closer to the intermediate state of the focal length of the two blocks. This results in a more uniform change in the focal length domain of the final dense focal stack image, further improving the visual effect.
[0108] Example 2
[0109] This invention provides a system for generating dense focal stack images, comprising: an image preprocessing module, a bidirectional parameter prediction module, and a new focal length image generation module;
[0110] The image preprocessing module includes: a focal length numbering unit and a block pair division unit;
[0111] The bidirectional parameter prediction module includes: a mode transformation unit and a first fitting parameter pair calculation unit;
[0112] The new focal length image generation module includes a prediction block generation unit and a stitching unit;
[0113] The focal length numbering unit is used to number a set of focal stack images with sparse focal length domains according to the focal length size.
[0114] The block pair partitioning unit is used to divide the numbered focal stack image into blocks in the same coordinate system, divide blocks with the same coordinates into a block group, and divide block pairs in the same block group;
[0115] The mode transformation unit is used to take the first block in each block pair as the reference block B and the second block as the current block I. By performing mode transformation on the reference block B, the corresponding fitting block F and the fitting parameter pair (m,σ) are obtained.
[0116] The first fitting parameter pair calculation unit is used to select the fitting parameter pair with the smallest error between the fitting block F and the current block I as the current first fitting parameter pair (m0, σ0);
[0117] The prediction block generation unit is used to generate prediction parameter pairs (m) using the current first fitted parameter pair as the best fitted parameter pair. p ,σ p The best-fit parameters are used to determine the reference block B as the reference block B for the new focal length image to be predicted. p By predicting the parameter pair (m) p ,σ p For reference block B p Perform the corresponding mode transformation to obtain the prediction block P;
[0118] The stitching unit is used to stitch together the predicted blocks corresponding to the block pairs with the same focal length number according to the coordinates of the corresponding block pairs to obtain the predicted new focal length image; wherein, the predicted new focal length image is integrated with a set of original focal length domain sparse focal stack images to form a dense focal stack image.
[0119] As a further design of the present invention, the system for generating dense focal stack images of the present invention further includes a second fitting parameter pair calculation unit, which is used to take the second block in each block as a reference block B and the first block as the current block I, and obtain the corresponding fitting block F and fitting parameter pair (m,σ) by performing mode transformation on the reference block B, and select the fitting parameter pair with the smallest error between the fitting block F and the current block I as the current second fitting parameter pair (m1,σ1).
[0120] Correspondingly, the prediction block generation unit also includes: selecting the pair with the smaller error between the first fitted parameter pair (m0, σ0) and the second fitted parameter pair (m1, σ1) as the current best fitted parameter pair to generate the prediction parameter pair (m0, σ0). p ,σ p ).
[0121] As a further design of the present invention, a verification module is also included, which is used to select a portion of the focal stack images in the focal length numbering unit as ground truth images; and to calculate the peak signal-to-noise ratio (PSNR) of the generated new focal length image with the corresponding ground truth image of the same focal length to verify the effect of the generated new focal length image; wherein, the higher the PSNR, the better the effect.
[0122] Each of the above units or modules is used to implement the steps corresponding to the method for generating dense focal stack images in Embodiment 1 above, as detailed in Embodiment 1 above. Figures 2-10 The specific implementation details of the corresponding steps will not be elaborated here.
[0123] Example 3
[0124] This invention also provides an electronic device, including a computer-readable storage medium and a processor;
[0125] Computer-readable storage media are used to store executable instructions;
[0126] The processor is used to read executable instructions stored in a computer-readable storage medium and execute the steps corresponding to the method for generating a dense focal stack image in Embodiment 1 above.
[0127] The electronic device can be a computer, mobile phone, television, embedded device, or other electronic device.
[0128] Example 4
[0129] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps corresponding to the method for generating a dense focal stack image as described in Embodiment 1 above.
[0130] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for generating a dense focal stack image, characterized in that, include: S1. Number a set of focal stack images with sparse focal length domain according to focal length. S2. In the same coordinate system, divide the numbered focal stack image into blocks, and divide the blocks with the same coordinates into a block group. Two blocks in a block group form a block pair, and the two blocks in a block pair are the first block and the second block, respectively. S3. Using the first block as the reference block and the second block as the current block, the corresponding fitting block and fitting parameter pair (m, σ) are obtained by performing mode transformation on the reference block. The fitting parameter pair with the smallest error between the fitting block and the current block is selected as the first fitting parameter pair (m0, σ0) of the current block pair. m represents the mode corresponding to the mode transformation and σ represents the parameter of the corresponding mode transformation. S4. Use the first fitted parameter pair (m0, σ0) as the best fitted parameter pair to generate the predicted parameter pair (m p ,σ p ), and use the prediction parameters to (m p ,σ p The reference block is subjected to a corresponding mode transformation to obtain the prediction block; S5. The predicted blocks corresponding to the block pairs with the same focal length number are stitched together according to the coordinates of the corresponding block pairs to obtain the predicted new focal length image; the new focal length image and the sparse focal stack image in the focal length domain constitute a dense focal stack image.
2. The method according to claim 1, characterized in that, S3 further includes: taking the second block as the reference block and the first block as the current block to obtain the second fitting parameter pair (m1,σ1) of the current block pair; S4 further includes: selecting the pair with the smaller error between the first fitting parameter pair (m0, σ0) and the second fitting parameter pair (m1, σ1) as the optimal fitting parameter pair.
3. The method according to claim 1 or 2, characterized in that, In S3, the mode transformation includes one or more of the following: copy transformation, Gaussian transformation, or Wiener transformation; The copy transformation involves using the reference block as the fitting block; The Gaussian transform includes: performing a convolution operation between the reference block and a two-dimensional Gaussian function to obtain the fitting block; The Wiener transform includes: multiplying the reference block with Wiener deconvolution in the frequency domain and then transforming it to the spatial domain to obtain the fitted block.
4. The method according to claim 3, characterized in that, In S4, the prediction parameter pair (m) p ,σ p The parameter σ of the mode transformation in ) p The average of the mode transformation parameters in the best-fit parameter pair is taken.
5. The method according to claim 3, characterized in that, In S4, if m p If it is a copy mode, then the corresponding mode transformation is a copy transformation; If m p If the mode is Gaussian, then the corresponding mode transformation is a Gaussian transform; If m p If it is a Wiener mode, then the corresponding mode transformation is the Wiener transform.
6. The method according to claim 1 or 2, characterized in that, In S5, the new focal length image is numbered according to the focal length numbers corresponding to the two blocks in the block pair. The newly numbered focal length image and the sparse focal stack image in the focal length domain are integrated into the dense focal stack image according to the focal length number.
7. The method according to claim 1 or 2, characterized in that, S1 also includes: selecting a portion of the focal stack image as the ground truth image; S5 is followed by: calculating the peak signal-to-noise ratio between the new focal length image and the corresponding ground truth image with the same focal length, in order to verify the effect of the new focal length image.
8. A system for generating densely focused images, characterized in that, The system for performing the method according to any one of claims 1-7, the system comprising: Focal length numbering unit is used to number a set of focal length stack images that are sparse in the focal length domain according to the focal length size. The block pair division unit is used to divide the numbered focal stack image into blocks in the same coordinate system, and divide the blocks with the same coordinates into a block group. Two blocks in a block group form a block pair, and the two blocks in a block pair are the first block and the second block, respectively. The mode transformation unit is used to take the first block as a reference block and the second block as the current block, and to obtain the corresponding fitting block and fitting parameter pair (m,σ) by performing mode transformation on the reference block; m represents the mode corresponding to the mode transformation, and σ represents the parameter of the corresponding mode transformation. The first fitting parameter pair calculation unit is used to select the fitting parameter pair with the smallest error between the fitting block and the current block as the first fitting parameter pair (m0,σ0) of the current block pair; The prediction block generation unit is used to generate prediction parameter pairs (m0, σ0) using the first fitted parameter pair (m0, σ0) as the best fitted parameter pair. p ,σ p ), and use the prediction parameters to (m p ,σ p The reference block is subjected to a corresponding mode transformation to obtain the prediction block; The stitching unit is used to stitch together the prediction blocks corresponding to the block pairs with the same focal length number according to the coordinates of the corresponding block pairs to obtain a new predicted focal length image; the new focal length image and the sparse focal stack image in the focal length domain constitute a dense focal stack image.
9. The system according to claim 8, characterized in that, It also includes a second fitting parameter pair calculation unit, which is used to take the second block as the reference block and the first block as the current block to obtain the second fitting parameter pair (m1,σ1) of the current block pair; The prediction block generation unit further includes: selecting the pair with the smaller error between the first fitting parameter pair (m0, σ0) and the second fitting parameter pair (m1, σ1) as the optimal fitting parameter pair.
10. An electronic device, characterized in that, Includes computer-readable storage media and processors; The computer-readable storage medium is used to store executable instructions; The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method according to any one of claims 1-7; Or / and, a computer-readable storage medium having a computer program stored thereon, characterized in that, when the program is executed by a processor, it implements the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Focus stack photo fusing method based on image reconstruction
CN104952048A
Optical field focus stack image sequence encoding and decoding method, device and system
CN110996104A