Microscopic defocusing image depth estimation method based on array light spots

Through array spot segmentation and complex-real value joint neural network model, the problem of insufficient depth estimation accuracy of a single defocused image in low texture or complex background is solved, and fast and accurate microscopic measurement is achieved.

CN120634878APending Publication Date: 2025-09-12SHANGHAI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510857389.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In existing microscopic measurement technologies, the depth estimation method of a single defocused image is not accurate enough in low-texture, low-contrast or complex backgrounds, and is ill-posed, making it difficult to meet industrial-grade measurement needs.

Method used

A microscopic defocused image depth estimation method based on array spot is adopted. The array spot image is collected in a microscope device and divided into image sub-blocks. The trained neural network model is used for depth prediction, and a complex-real value joint neural network model is constructed for feature extraction and fusion to generate a complete depth map.

Benefits of technology

It enables accurate measurement of scene depth with a single image, improving measurement speed and accuracy. It is suitable for textureless or highly reflective surfaces, and its accuracy meets industrial-grade requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634878A_ABST
    Figure CN120634878A_ABST
Patent Text Reader

Abstract

The invention relates to a microscopic defocusing image depth estimation method based on array light spots, and the method comprises the steps: collecting an array light spot image at a certain focal plane position of a microscopic device, and segmenting the array light spot image into a plurality of image sub-blocks; performing depth prediction on the plurality of image sub-blocks by using the trained neural network model to obtain depth maps of the plurality of image sub-blocks; and splicing the depth maps of the plurality of image sub-blocks to obtain a complete depth map so as to realize depth estimation of the to-be-measured object. Compared with the prior art, the method is high in measurement speed, high in precision and high in robustness, the measured object can be measured even if the measured object is a texture-free surface or a high-reflection surface, and the application scene of the technology is greatly expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of microscopic scene depth estimation, and in particular to a microscopic defocused image depth estimation method based on array light spots. Background Art

[0002] In microscopic three-dimensional measurement technology, sample depth estimation is crucial for achieving accurate three-dimensional reconstruction of the sample surface. Accurate depth information directly determines the accuracy and reliability of the measurement results, and is of great significance for industrial inspection. Currently, in industrial microscopic inspection, in order to obtain higher measurement accuracy, axial scanning is usually used to obtain more focus information and achieve high-precision depth estimation. However, the scanning imaging speed is slow and the data processing volume is large, which makes it difficult to meet the application requirements of efficient inspection. In comparison, defocus depth recovery (Depth from Defocus, DFD) can analyze defocus information and infer depth based on a single defocused image, which significantly simplifies the data acquisition process and effectively improves imaging efficiency and inspection speed.

[0003] DFD uses the relationship between the degree of defocus blur and the depth of an object to infer the depth of the scene by estimating the amount of blur. Current methods for estimating the defocus depth of a single image in microscopy are mainly divided into three categories: PSF-based estimation, frequency-domain analysis, and edge and gradient analysis. These methods have some limitations in practical applications:

[0004] 1) Depth estimation performance is poor for low-texture or flat areas, as the lack of sufficient image features makes it difficult to accurately infer the degree of blur;

[0005] 2) Sensitive to noise, especially in low-contrast or low-light environments, where noise can affect the accuracy of depth estimation;

[0006] 3) It has poor adaptability to complex scenes or highly non-uniform images. In addition, the lack of prior information makes the depth estimation of a single image ill-posed.

[0007] The deep learning-based method learns the complex mapping relationship between image blur and depth through large-scale data training, enabling the model to automatically extract multi-scale depth features from the defocus information of the image, significantly improving its performance in low-texture and high-noise environments, with higher accuracy and strong robustness. Chinese patent application CN102663721A proposes a method for defocus depth estimation and full-focus image acquisition of dynamic scenes. It adopts a global approach and uses a convolutional model in the imaging model. It first eliminates the radiosity variable of the scene and only estimates the depth of the scene. Then, it uses the image deblurring method to estimate the radiosity of the scene, thereby optimizing the depth estimation result and achieving the acquisition of full-focus images and depth maps of dynamic scenes.

[0008] However, while deep learning-based methods possess a wealth of prior information and are well-suited to various scenarios, they still rely heavily on edge features, making them inapplicable to texture-depleted and low-contrast scenes. Furthermore, depth estimation based solely on a single RGB image suffers from low accuracy, making it difficult to meet the requirements of industrial-grade measurement precision.

[0009] The current depth estimation method based on a single defocused image has insufficient estimation accuracy when dealing with scenes such as texture loss, low contrast or complex background, and has inherent ill-posedness. Therefore, it is necessary to develop a defocused image depth estimation method with the characteristics of fast measurement speed, high accuracy and strong robustness. Summary of the Invention

[0010] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide a method for estimating the depth of a microscopic defocused image based on an array spot.

[0011] The purpose of the present invention can be achieved by the following technical solutions:

[0012] A method for estimating depth of a microscopic defocused image based on an array spot, the method comprising:

[0013] An array spot image is collected at a certain focal plane position of a microscope device and divided into a plurality of image sub-blocks;

[0014] Use the trained neural network model to perform depth prediction on several image sub-blocks respectively to obtain depth maps of several image sub-blocks;

[0015] The depth maps of the plurality of image sub-blocks are spliced ​​together to obtain a complete depth map, thereby realizing depth estimation of the object to be measured.

[0016] Furthermore, the process of dividing the array light spot image into a plurality of image sub-blocks includes dividing the array light spot image into image sub-blocks of preset pixel sizes in sequence from left to right and from top to bottom.

[0017] Furthermore, the training process of the neural network model includes:

[0018] Build a generative neural network model;

[0019] Construct paired datasets for training generative neural network models;

[0020] A generative neural network model is trained according to the paired data set to generate a trained neural network model.

[0021] Furthermore, the process of training the generative neural network model based on the paired data set also includes:

[0022] Inputting the array light spot image sub-block in the paired data set into the generative neural network model, and outputting a predicted depth map of the array light spot image sub-block;

[0023] According to the predicted depth map of each array spot image sub-block and its corresponding real depth map, a depth loss function between the two is obtained. The expression of the depth loss function is:

[0024]

[0025] in, Denotes the predicted depth map, D is the true depth map, K is the number of predicted array spot images, β is the coefficient used to balance the weight relationship between GSIM loss and L1 loss, and GSIM represents the gradient similarity index;

[0026]

[0027] in, and They represent the gradient of the real depth map and the gradient of the predicted depth map respectively, and C3 is a constant term to prevent the zero division problem during the calculation process;

[0028] The network parameters of the generative neural network model are updated according to the depth loss function, and the above process is repeated until the model training is completed.

[0029] Furthermore, the process of constructing the paired dataset includes:

[0030] Calculate the depth map corresponding to the array spot image on different focal planes for different scenes;

[0031] Several positions in the array spot image are randomly selected. At each position, the array spot image sub-block and the corresponding depth ground truth image sub-block are segmented from different focal planes to form a paired data set.

[0032] Furthermore, the process of calculating the depth maps corresponding to the array spot images on different focal planes for different scenes includes:

[0033] With a step size of s μm and a total axial scanning distance of d μm, d / s array spot images are collected to form a focal stack;

[0034] Extract the position coordinates of all array spots; Based on the clarity evaluation operator, evaluate the image contrast change in the preset neighborhood of each spot in the axial direction, perform Gaussian fitting on it to obtain the clarity response curve, and take the peak position of the clarity response curve as the focus position of each spot; Use the triangular interpolation algorithm to interpolate and predict the focus position of the pixels in the remaining non-spot areas of the entire image to obtain a complete focus map z focus;

[0035] According to the optimal focus position z of each pixel focus (x,y), calculate any focal plane z k The axial distance from the focus, i.e. the defocus distance δ k (x, y), calculate the defocus distance of all pixels in the focal plane, and generate the depth map corresponding to the array spot image in the focal plane.

[0036] Furthermore, the expression for calculating the defocus distance of all pixels in the focal plane is:

[0037] δ k (x,y)=|z k -z focus (x,y)|,k=1,2,…,n

[0038] Among them, δ k (x,y) is the defocus distance, z k is the kth focal plane, z focus (x,y) is the focal position of the pixel, and n represents the total number of pixels in the image.

[0039] Furthermore, the neural network model is a complex-real value joint neural network model, including an encoder and a decoder.

[0040] Furthermore, the encoder includes a frequency domain branch and a spatial domain branch, and the encoder performs hierarchical feature extraction by collaborative modeling of the frequency domain and the spatial domain;

[0041] In the frequency domain branch, the input array spot image U (S) Perform fast Fourier transform to obtain complex-valued feature map U (F) Perform complex convolution on the complex-valued feature map, and use the amplitude and phase information of the complex-valued feature map to extract high-frequency features related to defocus blur; after a preset number of downsampling and feature extraction, the frequency domain branch generates complex-valued feature maps of different resolutions, and converts them into real-valued feature maps in the spatial domain through inverse Fourier transform

[0042] In the spatial domain branch, depthwise separable convolution is used to transform the input array spot image U (S) Perform local fuzzy feature extraction to generate multi-scale spatial feature maps

[0043] Finally, the feature map extracted from the frequency domain branch and the feature map extracted from the spatial branch Perform fusion at different scales to obtain the fusion feature map at each scale The fusion formula is:

[0044]

[0045] in, is the fusion feature map, is the real-valued feature map at scale x, is the spatial feature map at the x scale, and λ represents the fusion coefficient for weighted balance between frequency domain features and spatial domain features.

[0046] Furthermore, when the array spot image is divided into a plurality of image sub-blocks, the initial coordinates of each image sub-block are recorded, and when a complete depth map is obtained by splicing, the depth map is spliced ​​according to the initial coordinates of each image sub-block.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] 1. The present invention can accurately measure scene depth using a single array spot image, eliminating the drawback of traditional methods that require scanning and collecting multiple images along the optical axis, which results in slow topography measurement. The present invention divides a single array spot image into multiple sub-blocks and then predicts depth in parallel through a neural network, avoiding the resource bottleneck of full-image calculation and improving overall recognition speed.

[0049] 2. The present invention interpolates and predicts the focal position of pixels in the non-spot pattern portion. By actively projecting array light spots, it assigns features to the object being measured, enabling measurement even on textureless or highly reflective surfaces, greatly expanding the application scenarios of the technology.

[0050] 3. The present invention constructs a complex-real valued joint neural network model, which performs hierarchical feature extraction through collaborative modeling of frequency domain and spatial domain. In the frequency domain branch, fast Fourier transform and complex convolution operations are used to efficiently extract high-frequency features related to defocus blur; in the spatial domain branch, a depth-separable convolution module is used to extract local blur features, and then the frequency domain and spatial domain features are fused at multiple scales, which enhances the network's modeling ability for defocus features, thereby improving the accuracy of depth estimation. In the model, the decoder performs step-by-step upsampling operations and integrates them with the feature maps of the corresponding scale in the encoder through jump connections, which retains the high-frequency structural features in the image, enhances the spatial continuity of the depth map, and further improves the accuracy of depth estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 Flow chart of the method of the present invention;

[0052] Figure 2 This is a flowchart of the neural network model training of the present invention.

[0053] Figure 3 Schematic diagram of a paired dataset consisting of some array spot images and their corresponding real depth maps;

[0054] Figure 4 Schematic diagram of the process of achieving depth estimation for a single defocused image;

[0055] Figure 5 Schematic diagram of the depth map estimated from array spot images on different focal planes in the same scene and the comparison with the true depth map. DETAILED DESCRIPTION

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0057] Example 1

[0058] This embodiment discloses a method for estimating the depth of a microscopic defocused image based on an array spot. Figure 1 As shown, specifically including:

[0059] Step S1, collecting an array spot image at a certain focal plane position of a microscope device;

[0060] Step S2, dividing the array spot image into N image sub-blocks;

[0061] Step S3, using the trained neural network model to perform depth prediction on the N image sub-blocks respectively to obtain depth maps of the N image sub-blocks;

[0062] Step S4: stitching the depth maps of the N image sub-blocks to obtain a complete depth map, thereby realizing depth estimation of the object to be measured.

[0063] In step S2 , the process of dividing the array light spot image into a plurality of image sub-blocks includes dividing the array light spot image into image sub-blocks of 64×64 pixels in order from left to right and from top to bottom.

[0064] In step S3, the training process of the neural network model is as follows Figure 2 Shown, including:

[0065] Step S301, constructing a generative neural network model;

[0066] Step S302, constructing a paired data set for training a generative neural network model;

[0067] Step S303: training a generative neural network model based on the paired data set to generate a trained neural network model.

[0068] In step S301, the generative neural network model is a complex-real value joint neural network model, including an encoder and a decoder.

[0069] The encoder consists of a frequency domain branch and a spatial domain branch. The encoder performs hierarchical feature extraction by collaborative modeling of the frequency and spatial domains.

[0070] In the frequency domain branch, the input array spot image U (S) Perform fast Fourier transform to obtain complex-valued feature map U (F) Perform complex convolution on the complex-valued feature map, and use the amplitude and phase information of the complex-valued feature map to extract high-frequency features related to defocus blur; After five-level downsampling and feature extraction, the frequency domain branch generates five sets of complex-valued feature maps with different resolutions, and converts them into real-valued feature maps in the spatial domain through inverse Fourier transform

[0071] In the spatial branch, depth-wise separable convolution is used to transform the input array spot image U (S) Perform local fuzzy feature extraction to generate multi-scale spatial feature maps

[0072] Finally, the feature map extracted from the frequency domain branch and the feature map extracted from the spatial branch Perform fusion at five different scales to obtain the fusion feature map at each scale The fusion formula is:

[0073]

[0074] in, is the fusion feature map, is the real-valued feature map at scale x, is the spatial feature map at the x scale, and λ represents the fusion coefficient for weighted balance between frequency domain features and spatial domain features.

[0075] In step S302, the process of constructing the paired data set includes:

[0076] Calculate the depth map corresponding to the array spot image on different focal planes for different scenes;

[0077] Randomly select 600 positions in the array spot image;

[0078] At each position, the array spot image sub-block and the corresponding depth ground truth image sub-block are segmented from different focal planes to form a paired data set.

[0079] The process of calculating the depth map corresponding to the array spot image on different focal planes for different scenes includes:

[0080] Axial scanning to obtain focal stack: with a step size of s μm and a total axial scanning distance of d μm, d / s array spot images are collected to form a focal stack;

[0081] Focus position calculation: Extract the position coordinates of all array spots; Based on the clarity evaluation operator, evaluate the image contrast change in the 7×7 neighborhood of each spot in the axial direction, perform Gaussian fitting on it to obtain the clarity response curve, and take the peak position of the clarity response curve as the focus position of each spot; Use the triangular interpolation algorithm to interpolate and predict the focus position of the pixels in the remaining non-spot areas of the entire image to obtain a complete focus map z focus ;

[0082] Depth map generation: Based on the optimal focus position z of each pixel focus (x,y), calculate any focal plane z k The axial distance from the focus, i.e. the defocus distance δ k (x, y), calculate the defocus distance of all pixels in the focal plane, and generate the depth map corresponding to the array spot image in the focal plane. Figure 3 The figure shows the true depth maps corresponding to the array spot image on a focal plane in several different scenarios. In the figure, Ground Truth is the true depth map, and Depth Value on the side refers to the depth value, which refers to the columnar color scale chart on the side of the Ground Truth image. Different colors in the figure correspond to different depth values.

[0083] The expression for calculating the defocus distance of all pixels in the focal plane is:

[0084] δ k (x,y)=|z k -z focus (x,y)|,k=1,2,…,n

[0085] Among them, δ k (x,y) is the defocus distance, z k is the kth focal plane, z focus (x,y) is the focal position of the pixel, and n represents the total number of pixels in the image.

[0086] In step S303, the process of training the generative neural network model based on the paired data set further includes:

[0087] Input the array spot image sub-block in the paired data set into the generative neural network model, and output the predicted depth map of the array spot image sub-block;

[0088] According to the predicted depth map of each array spot image sub-block and its corresponding real depth map, the depth loss function between the two is obtained. The expression of the depth loss function is:

[0089]

[0090] in, Denotes the predicted depth map, D is the true depth map, K is the number of predicted array spot images, β is the coefficient used to balance the weight relationship between GSIM loss and L1 loss, and GSIM represents the gradient similarity index;

[0091]

[0092] Among them, the gradient term and They represent the gradient of the real depth map and the gradient of the predicted depth map respectively, and C3 is a constant term to prevent the zero division problem during the calculation process;

[0093] Update the network parameters of the generative neural network model according to the depth loss function, and repeat the above process until the model training is completed.

[0094] In step S4, when the array spot image is divided into several image sub-blocks, the initial coordinates of each image sub-block are recorded, and when a complete depth map is obtained, the depth map is spliced ​​according to the initial coordinates of each image sub-block.

[0095] The following is a practical example based on the above method.

[0096] When using the microscopic defocused image depth estimation method based on array spot in this embodiment, a 5-megapixel industrial camera is used as the image acquisition device, and a high-brightness green point light source is selected as the projection imaging light source to achieve depth estimation of the sample.

[0097] like Figure 4 As shown in the figure, the array spot image captured on a focal plane of the object to be measured is divided into N image sub-blocks of size 64×64 pixels, which are input into the algorithm of the present invention (i.e., the trained neural network model) to output a depth estimation map of the object to be measured. Compared with the true value, it can be seen that the error of this depth estimation map is small. Figure 5 This figure shows the depth maps estimated using this method for array spot images at different focal planes in the same scene, and their comparison with the true depth map. In both figures, the ground truth is the true depth map, and the depth value on the side refers to the color scale bar graph next to the ground truth image, with different colors corresponding to different depth values.

[0098] High-precision equipment from both domestic and international sources was used for a comparison of measurement accuracy: a confocal probe from a specific company, a confocal microscope from a specific company, and a 3D topography microscanning measurement method based on an array light spot. The heights of standard gauge blocks (cross-grooved and stepped) were measured using 10x and 20x objective lenses, respectively. The heights of the standard gauge blocks were then estimated using both existing technology (the 3D topography microscanning measurement method based on an array light spot) and the aforementioned method for estimating depth from microscopic defocused images based on an array light spot.

[0099] The comparison results are shown in Table 1 below.

[0100] Table 1 Comparison of measurement results

[0101]

[0102] As can be seen in Table 1, the error of the present invention is small. At different objective lens magnifications, the measurement accuracy of the present invention is better than 0.1 μm. The measurement accuracy is similar to that of high-precision equipment both domestically and internationally. Compared with other equipment, the present invention has a significant advantage in measurement time. Depth estimation can be achieved with only a single defocused array spot image, significantly improving measurement efficiency. This demonstrates that the present invention has achieved industrial-grade measurement accuracy and is suitable for most industrial measurements.

[0103] In summary, the present invention can accurately measure scene depth using a single array spot image, eliminating the drawback of traditional methods that require scanning and collecting multiple images along the optical axis, which results in slow topography measurement. Furthermore, by actively projecting array spots, the object being measured is characterized, making it suitable for measuring textureless or highly reflective surfaces, thereby improving measurement accuracy. A complex-real valued joint neural network model is also constructed to fuse frequency and spatial domain information, enhancing the network's ability to model defocused features and further improving depth estimation accuracy.

[0104] Example 2

[0105] Based on Example 1, this embodiment provides an electronic device, including: one or more processors and a memory, wherein the memory stores one or more programs, and the one or more programs include instructions for executing the aforementioned microscopic defocused image depth estimation method based on array spots.

[0106] At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and of course may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned microscopic defocused image depth estimation method based on the array spot. Of course, in addition to software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0107] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0108] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0109] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A method for estimating depth of a microscopic defocused image based on an array spot, characterized in that: The method comprises: An array spot image is collected at a certain focal plane position of a microscope device and divided into a plurality of image sub-blocks; Use the trained neural network model to perform depth prediction on several image sub-blocks respectively to obtain depth maps of several image sub-blocks; The depth maps of the plurality of image sub-blocks are spliced ​​together to obtain a complete depth map, thereby realizing depth estimation of the object to be measured.

2. The method for estimating depth of a microscopic defocused image based on an array spot according to claim 1, wherein: The process of dividing the array light spot image into a plurality of image sub-blocks includes dividing the array light spot image into image sub-blocks of preset pixel sizes in sequence from left to right and from top to bottom.

3. The method for estimating depth of a microscopic defocused image based on an array spot according to claim 1, wherein: The training process of the neural network model includes: Build a generative neural network model; Construct paired datasets for training generative neural network models; A generative neural network model is trained according to the paired data set to generate a trained neural network model.

4. The method for estimating depth of a microscopic defocused image based on an array spot according to claim 3, wherein: The process of training a generative neural network model based on the paired data set also includes: Inputting the array light spot image sub-block in the paired data set into the generative neural network model, and outputting a predicted depth map of the array light spot image sub-block; According to the predicted depth map of each array spot image sub-block and its corresponding real depth map, a depth loss function between the two is obtained. The expression of the depth loss function is: in, Denotes the predicted depth map, D is the true depth map, K is the number of predicted array spot images, β is the coefficient used to balance the weight relationship between GSIM loss and L1 loss, and GSIM represents the gradient similarity index; in, and They represent the gradient of the real depth map and the gradient of the predicted depth map respectively, and C3 is a constant term to prevent the zero division problem during the calculation process; The network parameters of the generative neural network model are updated according to the depth loss function, and the above process is repeated until the model training is completed.

5. The method for estimating depth of a microscopic defocused image based on an array spot according to claim 3, wherein: The process of constructing the paired dataset includes: Calculate the depth map corresponding to the array spot image on different focal planes for different scenes; Several positions in the array spot image are randomly selected. At each position, the array spot image sub-block and the corresponding depth ground truth image sub-block are segmented from different focal planes to form a paired data set.

6. The method for estimating depth of a microscopic defocused image based on an array spot according to claim 5, characterized in that: The process of calculating the depth maps corresponding to the array spot images on different focal planes for different scenes includes: With a step size of s μm and a total axial scanning distance of d μm, d / s array spot images are collected to form a focal stack; Extract the position coordinates of all array spots; Based on the clarity evaluation operator, evaluate the image contrast change in the preset neighborhood of each spot in the axial direction, perform Gaussian fitting on it to obtain the clarity response curve, and take the peak position of the clarity response curve as the focus position of each spot; Use the triangular interpolation algorithm to interpolate and predict the focus position of the pixels in the remaining non-spot areas of the entire image to obtain a complete focus map z focus ; According to the optimal focus position z of each pixel focus (x,y), calculate any focal plane z k The axial distance from the focus, i.e. the defocus distance δ k (x, y), calculate the defocus distance of all pixels in the focal plane, and generate the depth map corresponding to the array spot image in the focal plane.

7. The method for estimating depth of a microscopic defocused image based on an array spot according to claim 6, wherein: The expression for calculating the defocus distance of all pixels in the focal plane is: δ k (x,y)=|z k -z focus (x,y)|,k=1,2,…,n Among them, δ k (x,y) is the defocus distance, z k is the kth focal plane, z focus (x,y) is the focal position of the pixel, and n represents the total number of pixels in the image.

8. The method for estimating depth of a microscopic defocused image based on an array spot according to claim 1, wherein: The neural network model is a complex-real value joint neural network model, including an encoder and a decoder.

9. The method for estimating depth of a microscopic defocused image based on an array spot according to claim 8, wherein: The encoder includes a frequency domain branch and a spatial domain branch, and the encoder performs hierarchical feature extraction by collaborative modeling of the frequency domain and the spatial domain; In the frequency domain branch, the input array spot image U (S) Perform fast Fourier transform to obtain complex-valued feature map U (F) ; Perform complex convolution operation on the complex-valued feature map, and use the amplitude and phase information of the complex-valued feature map to extract high-frequency features related to defocus blur; After a preset number of downsampling and feature extraction, the frequency domain branch generates complex-valued feature maps of different resolutions, and converts them into real-valued feature maps V in the spatial domain through inverse Fourier transform x (S) ; In the spatial domain branch, depthwise separable convolution is used to transform the input array spot image U (S) Perform local fuzzy feature extraction to generate multi-scale spatial feature maps Finally, the feature map extracted from the frequency domain branch and the feature map extracted from the spatial branch Perform fusion at different scales to obtain the fusion feature map T at each scale x (S) , and its fusion formula is: Among them, T x (S) is the fusion feature map, V x (S) is the real-valued feature map at scale x, is the spatial feature map at scale x, and λ represents the fusion coefficient for weighted balance between frequency domain features and spatial domain features.

10. The method for estimating depth of a microscopic defocused image based on an array spot according to claim 1, wherein: When the array spot image is divided into a number of image sub-blocks, the initial coordinates of each image sub-block are recorded, and when the complete depth map is obtained by splicing, the depth map is spliced ​​according to the initial coordinates of each image sub-block.

Citation Information

Patent Citations

  • Defocus depth estimation and full focus image acquisition method of dynamic scene

    CN102663721A