A method for constructing a high-precision three-dimensional model of a microscopic sample
By using a liquid lens microscopy imaging system and a high-precision calibration plate correction algorithm, combined with the SML-FE algorithm and the LaMa restoration model, the problems of large errors, low efficiency and complex lighting in the three-dimensional reconstruction of microscopic samples were solved, and high-precision and stable three-dimensional model construction was achieved.
Patent Information
- Application Number
- CN202411936641.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2026-06-26
AI Technical Summary
Existing methods for 3D reconstruction of microscopic samples suffer from large errors, low efficiency, high cost, difficulty in adapting to easily damaged and deformable objects, and poor depth map reconstruction quality under complex lighting conditions.
A microscopic imaging system was constructed using liquid lens technology. Combined with a magnification correction algorithm based on a high-precision calibration plate, an improved SML-FE algorithm and a LaMa restoration model based on region segmentation were used. Through feature enhancement and second-order partial derivative calculation, depth information extraction and image processing were optimized.
It reduces hardware costs, improves image processing efficiency and reconstruction accuracy, enhances system stability, and enables high-precision 3D reconstruction under complex lighting conditions.
Smart Images

Figure CN122284082A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of microscopic imaging technology, and relates to liquid lenses, particularly a method for constructing a high-precision three-dimensional model of a microscopic sample. Background Technology
[0002] In recent years, reconstructing the 3D shape of objects from 2D images has become a key problem in computer vision research. During the conversion between the real world and images, the mapping of 3D space to a 2D plane leads to the loss of depth information. Traditional focused 3D reconstruction methods mostly obtain the z-axis height layer image by displacing the object layer by layer using high-precision mechanical devices, and finally reconstruct the 3D image. However, this method is not suitable for easily damaged or deformable objects, and it introduces non-negligible errors for high-precision micro-measurements, reducing efficiency. In particular, it is difficult to meet real-time processing requirements in terms of reconstruction accuracy, measurable range, equipment cost, and sample processing, and the introduction of mechanical devices inevitably introduces errors.
[0003] Traditional focusing evaluation functions are often subject to significant noise interference when dealing with real-world scenarios, making it difficult to produce satisfactory depth maps. In practical applications, due to significant differences in the material of the measured surface, such as overexposed areas or smooth, featureless regions, evaluation functions still cannot adequately determine the depth layer.
[0004] The invention disclosed in CN201010199905.8 mentions a method for acquiring and reconstructing a microscopic field of view, including acquiring and recording the light field information of a microscopic sample to obtain an image stack by controlling a slight movement of the microscopic sample along the optical axis; constructing a statistical theoretical model based on the obtained image stack and the corresponding Poisson model, and obtaining the gradient of the microscopic sample based on a compressed sensing theoretical model; and establishing a joint model based on the statistical theoretical model and the compressed sensing theoretical model, and iteratively solving the joint model to obtain the three-dimensional structure of the microscopic sample. In the above invention, although an image stack is acquired by slightly moving the microscopic sample along the optical axis, the specific acquisition method suffers from image distortion or insufficient depth information due to inaccurate focal length, affecting the quality and accuracy of the three-dimensional reconstruction. Summary of the Invention
[0005] The purpose of this invention is to address the aforementioned problems in existing technologies by proposing a method for constructing high-precision three-dimensional models of microscopic samples.
[0006] The objective of this invention can be achieved through the following technical solution: a method for constructing a high-precision three-dimensional model of a microscopic sample, wherein the specific steps of the method for constructing a high-precision three-dimensional model of a microscopic sample are as follows:
[0007] S1 Acquire the image sequence to be processed: Use a microscopic imaging system to acquire a set of image sequences of microscopic samples at different focal lengths; the microscopic imaging system includes a liquid lens, objective lens, tube lens, camera and light source, with the liquid lens positioned at the back focal plane of the objective lens;
[0008] S2 magnification correction processing: Each image in the image sequence is processed using a magnification correction algorithm based on a high-precision calibration plate to ensure that the pixel correspondence of each image is accurate and the magnification is constant;
[0009] S3 constructs a 3D model: For the image sequence after magnification and correction, a focused 3D reconstruction method is used to construct a 3D model of the micro sample;
[0010] In the focused 3D reconstruction method, the image sequence after magnification and correction is processed by the SML-FE algorithm, which is an improvement on the SML algorithm, to obtain a depth information map. In the SML-FE algorithm, a feature enhancement operator is introduced to filter out small noise while preserving the detailed features of the image. The second-order partial derivative is used to screen outliers. The second-order partial derivative is calculated along the Z-axis for each pixel in the focused evaluation image sequence to correct the focused evaluation curve and obtain the true depth value.
[0011] For the acquired depth information map, a LaMa inpainting model based on region segmentation is used for optimization. The design of the LaMa model based on region segmentation is as follows: a fast mask generation algorithm based on focus metric is used to identify and locate the region to be optimized by utilizing the focus characteristics of the image; an FFC module is used; and a comprehensive loss function is used as the loss function.
[0012] Based on the optimized depth information map, the three-dimensional contour of the object is constructed using the depth information of each pixel, resulting in a three-dimensional model of the microscopic sample.
[0013] Preferably, in step S1, the steps for acquiring images in the image sequence using the microscopic imaging system are as follows: the microscopic sample is positioned at the front focal plane of the microscope objective, and light emitted from the light source is uniformly irradiated onto the surface of the microscopic sample; subsequently, the reflected light from the microscopic sample passes through the objective and the liquid lens in sequence, and the light after passing through the liquid lens is further focused by the tube lens, finally forming an image on the camera.
[0014] Preferably, in step S1, by changing the current value entering the regulating liquid lens, the injection or discharge of liquid in the lens chamber is controlled, the shape of the elastic film of the liquid lens is changed, and the surface curvature radius of the liquid lens is changed, thereby realizing the focusing or zooming function.
[0015] Preferably, in step S1, the microscopic imaging system is an infinity microscopic optical system, the light source is an LED light source, and the objective lens is a metallographic semi-apochromatic objective lens; the microscopic sample is positioned at the front focal plane of the objective lens.
[0016] Preferably, in step S2, the specific steps of the magnification correction process are as follows: First, when the liquid lens current value is 0, a calibration plate image is acquired and recorded as I0, which serves as the target image for magnification correction; the current value is adjusted at intervals of step, and then the current value is adjusted once to obtain a new calibration plate image I1. The current values are adjusted sequentially to acquire images, and finally n calibration plate images are obtained: I1, I2, I3…I n After acquiring all the calibration board images, the contour detection algorithm is used to obtain the center coordinates of the top left, top right, bottom left, and bottom right corners of each image. Then, perspective transformation is used to obtain the transformation matrix between the two images. The transformation matrix is used to correct any image to the same magnification as the target image.
[0017]
[0018]
[0019] Where H represents the transformation matrix, and (x,y) and (x′,y′) represent the coordinates of the points before and after the transformation; based on the acquired image sequence, the complete transformation matrix group [H1,H2,H3...H] is obtained. n This allows you to obtain an image with the same magnification as the target image at any current calibration value.
[0020] Preferably, in step S3, the specific formula for the feature enhancement operator is as follows:
[0021]
[0022] Where: Ω is the calculation window parameter; G s It is a spatial Gaussian kernel, ||pq|| represents the distance between pixels, and G r It is a color Gaussian kernel, ||I(p)-I(q)|| represents the difference in pixel values; W p It is a normalization factor;
[0023] After the image sequence is calculated by the SML algorithm, the best focus evaluation map is fitted, and then the feature enhancement operator is calculated on the focus evaluation map.
[0024] Preferably, in step S3, the formula for calculating using the second-order partial derivative is as follows:
[0025]
[0026] The discrete double integral is defined with no specific meaning in its symbol. It is a second-order partial derivative with respect to the z-axis, used to filter out erroneous evaluation values.
[0027] Preferably, in step S3, a threshold T1 is introduced in the fast mask generation algorithm for focus measurement to classify image pixels into two categories: those with valid focus evaluation and those with invalid focus evaluation. For each pixel (i,j,k), when the focus measurement... i,j,k If the value exceeds the threshold T1, then set its focus_value. i,j,k The value is assigned as 1 if the value is not assigned, and 0 otherwise; this is represented as follows:
[0028]
[0029] focus_measure i,j,k This represents the grayscale value at (x, y) in the k-th image, where T1 represents the specified threshold, and focus_value i,j,k This represents the gray value at (x, y) in the k-th image. Here, there are only 0 and 1, which is a binary image. "Otherwises" represents another case.
[0030] Subsequently, the binarized focus_value is accumulated along the z-axis to obtain the total number of effective focus points for each column of pixels. This process is described by the following formula:
[0031]
[0032] The above formula applies to all focus_values. i,j,k The focus_sum is obtained by summing along the z-axis. i,j ;
[0033] Based on the accumulated results, another threshold T2 is introduced to generate the initial mask. If the number of focal points in a column exceeds T2, the column is considered not to need optimization and is marked as 0 in the mask; otherwise, it is marked as 1. The specific formula is as follows:
[0034]
[0035] T2 is the specified threshold, mask i,j This represents the obtained mask image;
[0036] Finally, a dilation operation is performed on this initial mask using a maximum value filter to expand the non-zero regions within the mask. This operation is achieved by taking the maximum value in a local neighborhood, as shown in the following formula:
[0037]
[0038] final_maski,j The final mask image is represented by (m,n)∈B, which represents the neighborhood range of the maximum value filter. The formula means that the maximum value is found in the neighborhood range of (i,j) in the mask image and then assigned to final_mask.
[0039] Preferably, in step S3, the specific steps of the operation in the FFC module are as follows: The input tensor is first transformed to the frequency domain by performing a two-dimensional real fast Fourier transform. The result of the transformation consists of real and imaginary parts, which are connected together as independent channels. Then, specific feature processing operations are performed on the tensor in the frequency domain. Feature extraction is performed through a 1×1 convolution operation, normalization is performed using batch normalization, and nonlinearity is introduced through the ReLU activation function to enhance feature representation. After completing the feature processing in the frequency domain, the inverse transform is finally applied to convert the processed frequency domain information back to the spatial domain: the real number representation is converted to the inverse form, and a two-dimensional inverse real fast Fourier transform is applied to convert the data from the frequency domain back to the spatial domain.
[0040] Preferably, in step S3, the expression for the comprehensive loss function is as follows:
[0041] L final =κL Adv +αL HRFPL +βL DiscPL +γR1
[0042] Among them, the resistance loss L Adv High receptive field perception loss L HRFPL Gradient penalty term R1, and discriminator-based perceptual loss L DiscPL κ, α, β, and γ are hyperparameters used to control the weight of each loss term in the total loss.
[0043] Compared with existing technologies, the method for constructing high-precision three-dimensional models of microscopic samples has the following advantages:
[0044] 1. Reduced hardware costs and improved system stability: This invention utilizes liquid lens technology to achieve focusing or zooming functions, thereby avoiding the use of complex mechanical zoom systems or inefficient Z-axis focusing systems, thus reducing hardware complexity and costs in industrial production. This technology not only improves the stability of the imaging system but also effectively reduces noise interference and enhances the stability of reconstruction results, making it particularly suitable for scenarios with extremely high requirements for image processing accuracy and stability, such as ultra-precision machining.
[0045] 2. Improve image processing efficiency and ensure real-time performance: By introducing a magnification correction algorithm based on a high-precision calibration plate, the pixel correspondence of each image in the image sequence is ensured to be accurate and the magnification is constant. This algorithm only requires one calibration and can quickly perform image correction in subsequent processing, greatly improving the efficiency of image processing and the real-time performance of subsequent algorithms.
[0046] 3. Optimize depth information extraction and preserve detailed features: Employing the SML-FE feature enhancement algorithm, an improvement upon the SML algorithm, effectively removes noise from the image while preserving detailed features, especially in areas with noise or outliers in the depth map. Second-order partial derivatives are used to calculate the depth information for each pixel along the Z-axis, further correcting the focus evaluation curve and ensuring accurate depth information. This algorithm significantly improves the accuracy of the depth map, providing accurate depth data for subsequent 3D reconstruction.
[0047] 4. Solving the problem of depth map optimization under complex lighting conditions: This invention proposes a LaMa repair model based on region segmentation for depth map optimization. Especially when traditional focusing 3D reconstruction methods face scenes with complex lighting (such as shadows, overexposure, etc.), it can accurately identify and repair the region to be optimized through focus measurement and fast mask generation algorithm, effectively overcoming the limitations of traditional methods under complex lighting conditions and ensuring the quality of high-precision 3D reconstruction.
[0048] In summary, through innovations in hardware design and software algorithms, this invention effectively reduces costs, improves accuracy, enhances real-time performance and stability in the 3D reconstruction of microscopic samples, thus meeting the higher precision requirements of industrial and scientific research. Attached Figure Description
[0049] Figure 1 This is a schematic diagram of the objective lens and liquid lens combination;
[0050] Figure 2 This is a dual telecentric optical path diagram;
[0051] Figure 3 Here is a flowchart of the magnification correction algorithm based on a high-precision calibration board;
[0052] Figure 4 Micrographs of the calibration plate at different current values;
[0053] Figure 5 Micrographs of the calibration plate at different current values after correction of magnification;
[0054] Figure 6 Schematic diagram of the focusing method;
[0055] Figure 7 A comparison of the image stack and evaluation function results is shown in the figure.
[0056] Figure 8 A comparison graph of focusing evaluation curves before and after SML enhancement;
[0057] Figure 9 This is a graph showing the results of the second-order partial derivative focused evaluation value curve;
[0058] Figure 10 This is a map of error areas in 3D reconstruction.
[0059] Figure 11 A diagram showing the mask and the region to be optimized;
[0060] Figure 12 This is a schematic diagram of the network structure of the LmMa repair model based on region segmentation;
[0061] Figure 13 This is a schematic diagram of the FFC module structure;
[0062] Figure 14 This is a schematic diagram of the depth map optimization results;
[0063] Figure 15 A schematic diagram of an image acquisition system for a microscopic imaging system;
[0064] Figure 16 A schematic diagram of the process for obtaining a 3D model of a microscopic sample.
[0065] In the image, 1. Microscopic sample; 2. Objective lens; 3. Liquid lens; 4. Tube lens; 5. Camera. Detailed Implementation
[0066] The following are specific embodiments of the present invention, which are described in conjunction with the accompanying drawings. However, the present invention is not limited to these embodiments.
[0067] The method for constructing a high-precision 3D model of a microscopic sample is described below. The specific steps of this method are as follows:
[0068] S1 obtains the image sequence to be processed: such as Figure 16 As shown, a series of images at different focal lengths of a microscopic sample were acquired using a microscopic imaging system. The microscopic imaging system mainly consists of an LED light source, a metallographic semi-apochromatic objective lens, a tube lens, and a camera. To avoid introducing additional aberrations and other errors, an infinity-corrected microscopic optical system was chosen to complete the imaging system. Figure 15 As shown, the steps for acquiring images in an image sequence using a microscopic imaging system are as follows: The microscopic sample 1 is positioned at the front focal plane of the microscope objective 2, and light emitted from the light source is uniformly irradiated onto the surface of the microscopic sample 1; subsequently, the reflected light from the microscopic sample 1 passes through the microscope objective 2 and the liquid lens 3 in sequence, and the light after passing through the liquid lens 3 is further focused by the tube lens 4, finally forming an image on the camera 5.
[0069] Infinity-corrected microscopic optical systems have significant advantages due to their ease of installation and ability to flexibly integrate various auxiliary optical elements, which facilitates the integration of liquid lenses.
[0070] In infinity-corrected microscopy optical systems, the introduction of liquid lenses as components significantly impacts the overall magnification. To achieve multifocal image sequence acquisition while maintaining constant system magnification, a theoretical analysis of the optimal position of the liquid lens within the microscopy system is necessary. Considering that the optical design of the infinity-corrected microscopy optical system creates parallel light paths between the objective and auxiliary objectives, this study focuses solely on the influence of the relationship between the liquid lens and the objective on the magnification. The objective and liquid lens combination system is as follows: Figure 1 As shown.
[0071] right Figure 1 According to Gaussian optics theory, the combined focal length of the system is:
[0072]
[0073] Where f represents the equivalent focal length of the combined system, f1′ and f2′ represent the focal lengths of the corresponding objective lens and liquid lens, respectively, and d represents the distance between the objective lens and the liquid lens. Given that the focal length f2′ of the liquid lens is variable, it can be seen from the above formula for the combined focal length that the focal length f of the combined system composed of the objective lens and the liquid lens mainly depends on the distance d between the principal planes of the two optical groups and the current focal length f2′ of the liquid lens. Specifically, when d = f1′, the focal length f of the combined system will remain unchanged and f = f1′, thus allowing us to obtain f... obj =f, which can be known from the magnification formula of an infinity-based microscopic optical system, f obj and f tube Since both remain unchanged, the magnification M remains constant. Based on the above analysis, we can conclude that when the liquid lens is positioned on the back focal plane of the objective lens, the liquid lens can be considered as an aperture stop, thus enabling the microscopic imaging system to form a dual telecentric optical path configuration, such as... Figure 2 As shown. At this point, by adjusting the focal length of the liquid lens to zoom, the magnification of the system remains constant.
[0074] S2 magnification correction processing: Each image in the image sequence is processed using a magnification correction algorithm based on a high-precision calibration plate to ensure that the pixel correspondence of each image is accurate and the magnification is constant;
[0075] like Figure 3As shown, the magnification correction algorithm based on a high-precision calibration plate ensures accurate pixel correspondence in each image sequence, maintaining a constant magnification. Furthermore, this algorithm requires only one calibration, allowing for rapid image correction based on the calibration results, ensuring real-time performance of subsequent algorithms. The specific process is as follows: First, a calibration plate image, denoted as I0, is acquired when the liquid lens current is 0, serving as the target image for magnification correction. Assuming the current value adjustment interval is step, adjusting the current value once yields a new calibration plate image I1. This process is repeated, adjusting the current value sequentially to acquire images, ultimately resulting in n calibration plate images: I1, I2, I3…I… n During the image acquisition process, it is crucial to pay attention to the degree of defocusing of the calibration board and ensure that the center of the calibration circle remains identifiable at all times. After acquiring all the calibration board images, a contour detection algorithm is used to obtain the coordinates of the center of the top left, top right, bottom left, and bottom right corners in each image. Then, perspective transformation is performed to obtain the transformation matrix between the two images. Using the transformation matrix, any image can be easily corrected to the same magnification as the target image.
[0076]
[0077] Where H represents the transformation matrix, and (x,y) and (x′,y′) represent the coordinates of the points before and after the transformation. Based on the acquired image sequence, the complete transformation matrix set [H1, H2, H3, ..., H...] can be obtained. n Thus, we can obtain an image with the same magnification as the target image at any current value.
[0078] To verify the accuracy of the magnification correction algorithm, a high-precision calibration plate was used as the experimental object. The magnitude of the liquid lens driving current was changed, and the change in magnification of the calibration plate was observed. Specific experimental results are as follows: Figure 4 and Figure 5 As shown: In Figure 4 In the above, (a) the current value is 0mA; (b) the current value is 20mA; (c) the current value is 40mA; (d) the current value is 60mA; (e) the current value is 80mA; and (f) the current value is 100mA. Figure 5 In the image, (a) the current value is 0mA; (b) the current value is 20mA; (c) the current value is 40mA; (d) the current value is 60mA; (e) the current value is 80mA; and (f) the current value is 100mA. It can be seen that there are significant differences in the magnification of the image under different current values, with magnifications of 5, 4.97, 4.89, 4.84, 4.77, and 4.69, respectively. After applying the magnification correction algorithm in this paper, the magnifications are 5, 5.05, 5.08, 5.03, 5.07, and 4.98, respectively. Figure 5It can be seen that the magnification correction algorithm is effective, and the magnification remains basically unchanged.
[0079] S3 constructs a 3D model: For the image sequence after magnification and correction, a focused 3D reconstruction method is used to construct a 3D model of the micro sample;
[0080] In the focused 3D reconstruction method, the image sequence after magnification and correction is processed by the SML-FE algorithm, which is an improvement on the SML algorithm, to obtain a depth information map. In the SML-FE algorithm, a feature enhancement operator is introduced to filter out small noise while preserving the detailed features of the image. The second-order partial derivative is used to screen outliers. The second-order partial derivative is calculated along the z-axis for each pixel in the focused evaluation image sequence to correct the focused evaluation curve and obtain the true depth value.
[0081] For the acquired depth information map, a LaMa inpainting model based on region segmentation is used for processing. The large mask model based on region segmentation is designed as follows: a fast mask generation algorithm based on focus metric is adopted to identify and locate the region to be optimized by utilizing the focus characteristics of the image; an FFC module is adopted; and a comprehensive loss function is adopted as the loss function.
[0082] By adjusting the current through a liquid lens-based microscopic imaging system, a multi-focus image sequence with constant magnification at different focal lengths is obtained. These images cover different depth ranges from near to far, ensuring that the entire target object or scene is captured clearly. To obtain accurate 3D reconstruction data, a focusing-based 3D reconstruction algorithm is used to process the multi-focus image sequence. This method analyzes the image sequence at different focal lengths and uses a focusing evaluation function to extract the depth information of each point in each image, thereby achieving high-precision 3D reconstruction of the object's surface. The basic idea of the focusing 3D reconstruction algorithm is: under ideal imaging conditions, depth data is obtained by processing a set of clear images, thereby restoring the 3D shape of the object. The key to this method is to obtain the sequence value with the highest sharpness for a particular pixel in the image. The focusing method first calculates the corresponding sharpness evaluation value for each pixel. After performing this operation on all images in the sequence, each pixel will have a sharpness function value equal to the number of images. These values can be fitted into a curve, which is a function of sharpness changing with the distance between the camera and the object. Then, by finding the peak of this curve, the image number to which the pixel belongs is determined. Since the change in the relative distance between the object and the camera is known, the displacement of the stage when that image number was captured can be determined. Combined with the relationship between the focal plane and the initial position of the stage, the depth information of that point can be obtained.
[0083] Once the depth data of all pixels is determined, the 3D shape of the target object can be reconstructed. A diagram illustrating the principle of the focused 3D reconstruction algorithm is shown below. Figure 6 As shown.
[0084] In image processing and computer vision, focus evaluation functions play a crucial role, especially in autofocus systems and depth estimation tasks. However, different types of focus evaluation functions have their own characteristics, and their performance varies across different scenarios. This paper compares five of the most typical focus evaluation functions—SML, Brenner, SMD, Tenengrad, and Roberts—and compares the final focus evaluation results and the execution time per image. Figure 7 As shown in Table 1:
[0085] Experimental results show that the evaluation method based on the SML function performs excellently on several key indicators compared to other focus evaluation functions. First, the SML function exhibits excellent unbiasedness, meaning it can provide objective evaluation results in various image scenarios. Second, the function has a strong unimodal characteristic, which enables it to accurately locate the optimal focus position and reduce misjudgments. Simultaneously, the SML operator demonstrates high sensitivity, capable of capturing minute focus changes, providing the possibility for fine-tuning focus. Finally, Table 1 shows the running time of different focus evaluation functions on a single image, demonstrating that the SML focus evaluation function also has a certain advantage in terms of time compared to other functions. Considering the superior performance and theoretical basis of the SML function, it was decided to use it as the basis for further optimization and improvement. Previous research has shown that for cases where focus and defocus cannot be distinguished, filtering operations are often used directly on the focus evaluation map to remove erroneous evaluation values. However, this leads to the loss of feature details while removing erroneous evaluation values. To address this issue, this invention improves the SML function as follows:
[0086] (1) An enhancement operator improves the SML focusing evaluation function. The specific formula of the enhancement operator is as follows:
[0087]
[0088] Where: Ω is the calculation window parameter; G s It is a spatial Gaussian kernel, ||pq|| represents the distance between pixels, and G r It is a color Gaussian kernel, ||I(p)-I(q)|| represents the difference in pixel values; W p It is the normalization factor. After SML calculation, the image sequence can be fitted with the best focus evaluation map. Then, the enhancement operator is calculated on the evaluation map, which can filter out small noise while preserving the feature details of the image.
[0089] like Figure 8As shown in the figure, the SML function enhanced by the enhancement operator is superior to the original SML function in terms of unimodality and robustness, and the curve transition is smoother and less affected by noise compared to the original SML function.
[0090] 2) Introduce second-order partial derivatives to screen outliers:
[0091] In image sequences, individual pixels and their surrounding areas experience a transition from blurry to sharp and back to blurry. This trend aligns with the focus value curve calculated by the focus evaluation function, demonstrating the unimodal characteristic of the focus curve. However, in practice, some areas still exhibit incorrect focus sequence numbers. Observing the focus curve at a point in this area reveals a multi-peak structure. This anomaly makes conventional depth estimation techniques potentially unable to accurately identify the true focus location. Research and analysis of this phenomenon can broadly categorize scenarios with multiple peaks in the focus curve into two types. The first type of multi-peak focus curve phenomenon typically occurs near areas with significant depth differences on the surface being measured. This multi-peak phenomenon can be attributed to points near areas of abrupt depth changes being influenced by adjacent areas of different depths. Experimental observations show that the closer a point is to the fault edge, the more significant the influence. In some cases, erroneous peaks may even exceed correct peaks. This phenomenon makes it difficult to accurately determine the true depth based solely on peak size. It is also noted that the size of the focus evaluation window significantly affects the occurrence of the bimodal phenomenon. A larger evaluation window increases the degree to which a point is influenced by the surrounding area, thus making it easier to generate multiple peaks.
[0092] However, in practical applications, the depth difference of the measured surface is usually not too drastic. Therefore, even if such multi-peak phenomena occur, the deviation between the calculated depth value and the actual depth is generally not too large. Furthermore, based on the enhancement operator proposed above, the ability to retain spatial features while utilizing small window denoising can also prevent this situation from occurring. The second type of multi-peak phenomenon in focus evaluation values occurs in low-frequency regions near high-frequency regions. High-frequency regions typically refer to parts of the image where gradient changes are significant, such as the boundary contours of objects or the junctions of different materials. The first or second derivatives of these regions usually exhibit large values. The design principle of the focus evaluation operator determines that it is more sensitive to high-frequency regions. Therefore, when applying this type of operator, high-frequency regions often produce response values that are significantly higher than those of low-frequency regions. This difference in response is directly reflected in the focus evaluation values, making the focus evaluation values of high-frequency regions usually much higher than those of low-frequency regions. However, in blurred images, edge blurring may produce imaginary contours in low-frequency regions, which makes the sharpness evaluation of these regions in out-of-focus photos potentially higher than in focused images, thus causing multiple peaks in the focus curve. The actual sharpness information is often obscured by false edges in the surrounding high-frequency region, and the multi-peak phenomenon of the second type of focusing curve is specifically caused by the lack of surface texture information of the measured object. Therefore, this study introduces a second-order partial derivative to address this situation of multiple peaks:
[0093]
[0094] The discrete double integral is defined with no specific meaning in its symbol. It is a second-order partial derivative with respect to the z-axis, used to filter out erroneous evaluation values.
[0095] Specifically, the second-order partial derivatives of the data along the z-axis are calculated for each pixel in the focus evaluation image sequence. The theoretical basis of this method is that second-order partial derivatives can more sensitively capture the rate of change in image intensity, thus providing more refined information about the degree of focus. By introducing second-order partial derivative analysis, the ability to identify true focus can be enhanced, while distinguishing between true focus and false focus caused by image noise or false edges. The focus evaluation curve after the second-order partial derivative is then obtained, as shown below. Figure 9 As shown in the image above, after filtering using the second-order partial derivative, the correct evaluation value for point b can be obtained. Comparative analysis leads to the conclusion that, compared to... Figure 9In the depth image presented in (b), erroneous peaks in multiple regions have been significantly corrected. This improvement greatly enhances the accuracy of focus assessment. Furthermore, this result strongly confirms the effectiveness of the SML-FE algorithm. The experimental results clearly demonstrate the excellent performance of the SML-FE algorithm in eliminating erroneous peaks. This optimization is not limited to individual regions but is pervasive throughout the entire depth map. By reducing the frequency and intensity of erroneous peaks, the algorithm significantly improves the reliability and accuracy of focus assessment.
[0096] After processing with the improved focusing evaluation function, the depth information map of the object under test was successfully obtained. Based on this depth map, the three-dimensional contour of the object was constructed, as shown below. Figure 10 The image shows the results of 3D reconstruction using the obtained depth map. However, significant depth errors were observed in some areas.
[0097] Analysis reveals that in specific experimental scenarios, due to uneven lighting, the reflective properties of material surfaces, and occlusion by external objects, some areas are dark and form shadows, while others have high reflectivity and are overexposed. Significant image information is lost in these areas, making it difficult to obtain high-quality images simply by adjusting the lighting system. This paper draws on existing mature image inpainting networks to optimize the depth map. However, the region to be optimized cannot be determined. Therefore, this paper first segments the region to be optimized and then performs depth optimization on that region. This method overcomes the limitations of traditional focus evaluation functions under complex lighting conditions. Before depth map optimization, obtaining an accurate mask for the region to be optimized is a crucial preprocessing step. To achieve this, a fast mask generation algorithm based on focus metrics is designed. The core idea of this algorithm is to utilize the focus characteristics of the image to identify and locate the region to be optimized, and to optimize the accuracy and coverage of the mask through multi-step processing. The first stage of the algorithm involves binarization of the focus metrics. A threshold T1 is introduced to classify image pixels into two categories: those with valid focus evaluation and those with invalid focus evaluation. Specifically, for each pixel (i,j,k), if its focus metric (focus_measure) {i,j,k} If the value exceeds the threshold T1, then set its focus_value. i,j,k Assign a value of 1 if the value is 1, otherwise assign a value of 0. This step can be represented as:
[0098]
[0099] focus_measure i,j,k This represents the grayscale value at (x, y) in the k-th image, where T1 represents the specified threshold, and focus_value i,j,kThis represents the gray value at (x, y) in the k-th image. Here, there are only 0 and 1, which is a binary image. "Otherwises" represents another case.
[0100] Subsequently, the binarized focus_values are accumulated along the z-axis to obtain the total number of effective focus points for each column of pixels. This process can be described by the following formula:
[0101]
[0102] The above formula applies to all focus_values. i,j,k The focus_sum is obtained by summing along the z-axis. i,j ;
[0103] Based on this accumulated result, another threshold T2 is introduced to generate the initial mask. If the number of focal points in a column exceeds T2, the column is considered not to need optimization and is therefore marked as 0 in the mask; otherwise, it is marked as 1. The specific formula can be expressed as:
[0104]
[0105] T2 is the specified threshold, mask i,j This represents the obtained mask image;
[0106] Finally, the algorithm uses maximum filtering to dilate the initial mask, expanding the non-zero regions within it. This operation is achieved by finding the maximum value within a local neighborhood, as shown in the following formula:
[0107]
[0108] final_mask i,j The final mask image is represented by (m,n)∈B, which represents the neighborhood range of the maximum value filter. The formula means that the maximum value is found in the neighborhood range of (i,j) in the mask image and then assigned to final_mask.
[0109] like Figure 11 As shown, the aforementioned fast mask generation algorithm can not only effectively identify the regions that need optimization, but also optimize the shape and size of the mask through dilation operations, while providing more accurate input information for the optimization algorithm. By processing the depth map using the proposed fast mask generation algorithm, a mask for the region to be optimized can be obtained.
[0110] Based on the aforementioned region segmentation algorithm and combined with the deep learning-based Large Mask Inpainting (LaMa) model, a region segmentation-based Large Mask Inpainting model was constructed; the specific network structure diagram is shown below. Figure 12 As shown.
[0111] FFC Module: When facing challenging situations such as large-area missing pixels, generating appropriate repair results requires consideration of global context. Therefore, a good architecture should have units with the widest possible receptive field in the processing flow as early as possible. Traditional fully convolutional models (such as ResNet) suffer from slow effective receptive field growth. Due to the often small convolutional kernels (such as 3×3), receptivity may also be insufficient, especially in the early layers of the network. Therefore, many layers in the network will lack global context, and computation and parameters will be wasted on creating global context. For wide masks, the generator's entire receptive field at a particular location may be within the mask, thus only the missing pixels can be observed. This problem is particularly prominent in high-resolution images. Fast Fourier Convolution (FFC) is an operator that allows the use of global context in early layers. FFC employs a channel-based Fast Fourier Transform, which is characterized by its ability to cover the receptive range of the entire image. This method divides the channel into two parallel processing parts: the first part is a local branch utilizing traditional convolution operations; the second part is a global branch applying real-valued FFT to capture the global context. Real-number FFT is only applicable to processing real-valued signals, while its inverse transform ensures the real-valued nature of the output. Compared to the standard FFT, real-number FFT only requires half of the spectrum for computation. Specifically, FFC performs the following steps: First, it calculates a two-dimensional real-number fast Fourier transform on the input tensor, concatenating the real and imaginary parts as independent channels; then, it applies a convolution block in the frequency domain, including 1×1 convolution, batch normalization (BN), and the ReLU activation function; finally, it applies an inverse transform to convert the processed frequency domain information back to the spatial domain: converting the real number representation to an inverse form, and applying a two-dimensional inverse real-number fast Fourier transform to convert the data from the frequency domain back to the spatial domain. A schematic diagram of the specific FFC module structure is shown in Figure 13.
[0112] Loss Function Design: Simple supervised loss requires the generator to accurately reconstruct the real situation. However, the visible parts of an image often do not contain enough information for the accurate reconstruction of the masked parts. Therefore, using simple supervision can lead to blurry results due to the averaging of multiple possible padding patterns. LaMa defines a loss function L that includes adversarial loss. Adv High receptive field perception loss L HRFPL In addition to the additional term gradient penalty R1 and discriminator-based perceptual loss L DiscPL The comprehensive loss function has the following specific form:
[0113] L final =κL Adv +αL HRFPL +βL DiscPL +γR1
[0114] Adversarial loss aims to ensure that the generated image looks natural, especially in local details. The discriminator can distinguish between "real" and "fake" patches at the local patch level. Only patches that intersect with the masked area are marked as "fake." Adversarial loss encourages the generator f... θ (x′) produces realistic local details, making the generated images more in line with the natural effect expected by the human visual system.
[0115] In general, L Adv and L DiscPL L is responsible for generating natural local details. HRFPL This component is responsible for providing monitoring signals and ensuring the consistency of the overall structure. This multi-component loss function design balances different performance requirements, thereby achieving higher quality repair results.
[0116] Depth map optimization results: The inpainting model described above was applied to optimize the identified depth map regions to be inpainted. To better compare the processing effects of traditional image inpainting techniques and the LaMa inpainting model, the image inpainting algorithm in OpenCV was used to optimize the depth map of the regions to be inpainted as well. The optimized results are as follows: Figure 14 As shown, in Figure 14 In the middle, (a) is the depth map optimization result based on traditional image inpainting algorithms;
[0117] (b) Depth map optimization results of the LaMa inpainting model based on region segmentation; The depth map optimization results show that, compared to previous depth maps, both traditional algorithms and the LaMa inpainting model based on region segmentation demonstrate depth optimization effects. As shown in the figure above, the improvement is more pronounced in the region to be optimized: previously missing or erroneous areas have been properly filled, making the entire depth map more complete and coherent; the clarity of object edges is effectively maintained, avoiding common edge blurring problems; the depth values of the region to be optimized transition naturally with the surrounding areas, maintaining the overall depth consistency of the scene; finally, the optimized region well maintains the geometric shape and spatial relationships of objects in the scene, without introducing obvious distortions or unreasonable structures. These improvements fully demonstrate the effectiveness of the proposed fast mask generation algorithm and the introduction of a depth map inpainting model for depth map optimization. By accurately locating the region to be optimized and applying an efficient image inpainting model to complete the depth map optimization, the overall quality and usability of the depth map are successfully improved.
[0118] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or substitute them in a similar manner, without departing from the spirit of the invention or exceeding its defined scope. Although the invention has been detailed and described in the accompanying drawings and foregoing description, such descriptions are considered illustrative or exemplary rather than restrictive. It should be understood that changes and modifications can be made by those skilled in the art within the scope of the following claims. Specifically, the invention covers additional embodiments having any combination of features from the different embodiments described above. With regard to the use of the expressions “general” or “substantially,” this patent application should be understood to disclose that the disclosure equally fully satisfies these features and values, i.e., without any of the foregoing characterizations as “general” or “substantially.”
Claims
1. A method for constructing a high-precision three-dimensional model of a microscopic sample, characterized in that, The method for constructing a high-precision 3D model of a microscopic sample includes the following steps: S1 Acquire the image sequence to be processed: Use a microscopic imaging system to acquire a set of image sequences of microscopic samples at different focal lengths; the microscopic imaging system includes a liquid lens, objective lens, tube lens, camera and light source, with the liquid lens positioned at the back focal plane of the objective lens; the process of zooming is completed by changing the curvature of the liquid lens by changing the current value of the liquid lens. S2 magnification correction processing: Each image in the image sequence is processed using a magnification correction algorithm based on a high-precision calibration plate to ensure that the pixel correspondence of each image is accurate and the magnification is constant; S3 constructs a 3D model: For the image sequence after magnification and correction, a focused 3D reconstruction method is used to construct a 3D model of the micro sample; In the focused 3D reconstruction method, the image sequence after magnification and correction is processed by the SML-FE algorithm, which is an improvement on the SML algorithm, to obtain a depth information map. In the SML-FE algorithm, a feature enhancement operator is introduced, and the second-order partial derivative is used to screen outliers. For each pixel in the focused evaluation image sequence obtained during the processing, the second-order partial derivative is calculated along the z-axis (image stacking direction) to correct the focused evaluation curve and obtain the true depth value. For the acquired depth information map, a LaMa inpainting model based on region segmentation is used for optimization. The LaMa model based on region segmentation is designed as follows: a fast mask generation algorithm based on focus metric is used to identify and locate the region to be optimized by utilizing the focus characteristics of the image; an FFC module is used; and a comprehensive loss function is used as the loss function.
2. The method for constructing a high-precision three-dimensional model of a microscopic sample as described in claim 1, characterized in that, In step S1, the steps for acquiring images in the image sequence using the microscopic imaging system are as follows: the microscopic sample is positioned at the front focal plane of the microscope objective, and light emitted from the light source is uniformly irradiated onto the surface of the microscopic sample; subsequently, the reflected light from the microscopic sample passes through the objective and the liquid lens in sequence, and the light after passing through the liquid lens is further focused by the tube lens, finally forming an image on the camera.
3. The method for constructing a high-precision three-dimensional model of a microscopic sample as described in claim 1, characterized in that, In step S1, by changing the current value entering the regulating liquid lens, the injection or discharge of liquid in the lens chamber is controlled, the shape of the elastic film of the liquid lens is changed, and the surface curvature radius of the liquid lens is changed, thereby realizing the focusing or zooming function.
4. The method for constructing a high-precision three-dimensional model of a microscopic sample as described in claim 1, characterized in that, In step S1, the microscopic imaging system is an infinity microscopic optical system, the light source is an LED light source, and the objective lens is a metallographic semi-apochromatic objective lens; the microscopic sample is positioned at the front focal plane of the objective lens.
5. The method for constructing a high-precision three-dimensional model of a microscopic sample as described in claim 1, characterized in that, In step S2, the specific steps of the magnification correction process are as follows: First, when the liquid lens current value is 0, a calibration plate image is acquired and denoted as I0, which serves as the target image for magnification correction. The current value is adjusted at intervals of step, and then the current value is adjusted once to obtain a new calibration plate image I1. The current value is adjusted sequentially to acquire images, and finally n calibration plate images are obtained: I1, I2, I3...I4. After acquiring all the calibration plate images, the contour detection algorithm is used to obtain the center coordinates of the upper left, upper right, lower left, and lower right corners of each image. Then, perspective transformation is used to obtain the transformation matrix between the two images. The transformation matrix is used to correct any image to the same magnification as the target image. Where H represents the transformation matrix, (x, y), (x', y') represent the coordinates of the points before and after transformation; according to the obtained image sequence, all transformation matrix groups [H1, H2, H3…H n ] are obtained, and the image with the same magnification as the target image under the current arbitrary calibration current value is obtained.
6. The method for constructing a high-precision three-dimensional model of a microscopic sample as described in claim 1, characterized in that, In step S3, the specific formula for the feature enhancement operator is as follows: Where: Ω is the calculation window parameter; G s It is a spatial Gaussian kernel, |pq| represents the distance between pixels, and G r It is a color Gaussian kernel, ||I(p)-I(q)|| represents the difference in pixel values; W P It is a normalization factor; After the image sequence is calculated by the SML algorithm, the best focus evaluation map is fitted, and then the feature enhancement operator is calculated on the focus evaluation map.
7. The method for constructing a high-precision three-dimensional model of a microscopic sample as described in claim 1, characterized in that, In step S3, the formula for calculating using second-order partial derivatives is as follows:
8. The method for constructing a high-precision three-dimensional model of a microscopic sample as described in claim 1, characterized in that, In step S3, a threshold T1 is introduced in the fast mask generation algorithm for focus measurement to classify image pixels into two categories: those with valid focus evaluation and those with invalid focus evaluation. For each pixel (i,j,k), when the focus measurement... i,j,k If the value exceeds the threshold T1, then set its focus_value. i,j,k The value is assigned as 1 if the value is not assigned, and 0 otherwise; this is represented as follows: focus_measure i,j,k This represents the grayscale value at (x, y) in the k-th image, where T1 represents the specified threshold, and focus_value i,j,k This represents the gray value at (x, y) in the k-th image. Here, there are only 0 and 1, which is a binary image. "Otherwises" represents another case. Subsequently, the binarized focus_value is accumulated along the z-axis to obtain the total number of effective focus points for each column of pixels. This process is described by the following formula: The above formula applies to all focus_values. i,j,k The focus_sum is obtained by summing along the z-axis. i,j ; Based on the accumulated results, another threshold T2 is introduced to generate the initial mask. If the number of focal points in a column exceeds T2, the column is considered not to need optimization and is marked as 0 in the mask; otherwise, it is marked as 1. The specific formula is as follows: T2 is the specified threshold, mask i,j This represents the obtained mask image; Finally, a dilation operation is performed on this initial mask using a maximum value filter to expand the non-zero regions within the mask. This operation is achieved by taking the maximum value in a local neighborhood, as shown in the following formula: final_mask i,j This represents the final mask image. (m,n)∈B This represents the neighborhood range of the maximum value filter. The formula means that the maximum value is found within the neighborhood of (i,j) in the mask image and then assigned to final_mask.
9. The method for constructing a high-precision three-dimensional model of a microscopic sample as described in claim 1, characterized in that, In step S3, the specific steps of the operation in the FFC module are as follows: The input tensor is first transformed to the frequency domain by performing a two-dimensional real fast Fourier transform. The result of the transform consists of real and imaginary parts, which are connected together as independent channels. Then, specific feature processing operations are performed on the tensor in the frequency domain. Feature extraction is performed through a 1×1 convolution operation, normalization is performed using batch normalization, and nonlinearity is introduced through the ReLU activation function to enhance feature representation. After completing the feature processing in the frequency domain, the inverse transform is finally applied to convert the processed frequency domain information back to the spatial domain: the real number representation is converted to the inverse form, and a two-dimensional inverse real fast Fourier transform is applied to convert the data from the frequency domain back to the spatial domain.
10. The method for constructing a high-precision three-dimensional model of a microscopic sample as described in claim 1, characterized in that, In step S3, the expression for the comprehensive loss function is as follows: L final =κL Adv +αL HRFPL +βL DiscPL +γR1 Among them, the resistance loss L Adv High receptive field perception loss L HRFPL Gradient penalty term R1, and discriminator-based perceptual loss L DiscPL κ, α, β, and γ are hyperparameters used to control the weight of each loss term in the total loss.
Citation Information
Patent Citations
Microcosmic optical field acquisition and three-dimensional reconstruction method and device
CN101865673B