Micro-nano structure three-dimensional reconstruction method and system based on scanning electron microscope
By combining multi-view, multi-detector SEM images with neural field 3D representation, shadow areas are automatically separated, solving the accuracy and applicability issues of existing 3D-SEM technology in the reconstruction of complex micro-nano structures, and achieving high-precision 3D model reconstruction.
Patent Information
- Application Number
- CN202511038785.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-18
Smart Images

Figure CN120976424A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of biological micro / nano surface structure analysis and ultra-precision manufacturing, and specifically relates to a method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy. Background Technology
[0002] Scanning electron microscopy (SEM) is a high-resolution imaging device widely used in scientific research and industrial manufacturing. SEM obtains two-dimensional images at the micro- and nano-scale by scanning the sample surface with a focused electron beam, generating electronic signals including secondary electrons (SE) and backscattered electrons (BSE). However, SEM images essentially only reflect the intensity distribution of electron signals and lack a direct representation of the true three-dimensional geometry of the sample surface. In fields such as biological tissue surface morphology analysis and ultra-precision manufacturing, obtaining accurate three-dimensional structural information of samples is crucial. To reconstruct the three-dimensional morphology of samples from two-dimensional SEM images, a technique called 3D-SEM has emerged. Existing 3D-SEM methods are mainly divided into three categories: multi-view methods, single-view methods, and hybrid methods.
[0003] Multi-view methods are primarily based on structure-of-motion (SEM) techniques, performing feature matching and 3D reconstruction on a series of SEM images acquired from different viewpoints. However, these methods often face difficulties in feature matching when dealing with micro / nano samples with smooth surfaces and scarce texture information, resulting in poor reconstruction quality in these areas, which are highly prevalent in micro / nano structures. Another approach is the single-view method, primarily based on photometric stereo, which reconstructs the 3D structure of a sample surface from multiple images taken under different lighting conditions from a fixed viewpoint. This is mainly based on a four-quadrant backscattered electron (4Q-BSE) detector in SEM, typically mounted below the objective lens. Its annular structure is divided into four symmetrical quadrants, each recording BSE signals from different directions, representing the effect of the object being illuminated by light sources from different directions in the SEM image. The single-view method calculates the sample surface gradient based on the Lambertian surface distribution assumption of the BSE signal and the symmetrical distribution between the quadrants of the 4Q-BSE detector, and obtains the 3D structure of the sample surface through gradient integration. However, single-view methods have some inherent limitations: 1. Surface integration methods cannot handle drastic, discontinuous height changes on the sample surface; 2. They require a calibrated sample with known morphology to complete a complex parameter calibration process; 3. They cannot effectively handle shadow areas in the image, as BSE shadows introduce incorrect gradient estimations, leading to deformation in the reconstruction results. Therefore, single-view methods are mainly suitable for smooth samples with small surface undulations, greatly limiting their practicality. To overcome these limitations, some studies have attempted to combine multi-view and single-view methods, proposing so-called hybrid methods. The core idea is to first use multi-view methods to obtain a rough geometric structure of the sample, and then use single-view methods to supplement detailed information. However, existing hybrid methods still have significant shortcomings: 1. Existing hybrid methods use two-dimensional height maps as geometric representations, which are difficult to express complex three-dimensional structures, limiting the applicability of the method; 2. They still rely on the parameter calibration steps of the single-view part; 3. The shadow interference problem has not been effectively solved, and incorrect gradient information in shadow areas still leads to deviations in the final reconstruction results.
[0004] In summary, existing 3D-SEM technologies still suffer from limitations in accuracy, narrow applicability, dependence on parameter calibration, and sensitivity to shadows. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention provides a method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy. This method uses multi-view, multi-detector two-dimensional SEM images as input, combined with a neural field three-dimensional representation. Without requiring additional parameter calibration procedures, it can automatically separate large-area shadows caused by sample geometric occlusion in the image during training, significantly improving the accuracy of three-dimensional reconstruction of complex micro / nano structures. This invention provides reliable technical support for morphology measurement and structural analysis at the micro / nano scale and can be widely applied in fields such as biological surface micro / nano structure analysis, material surface analysis, and ultra-precision manufacturing.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, this invention proposes a method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy, comprising the following steps:
[0008] (1) Collect multi-view, multi-detector SEM images of the micro-nano sample to be tested, each view including secondary electron SE image and four-quadrant backscattered electron BSE image.
[0009] (2) The motion recovery structure algorithm is used to obtain the pose of the virtual camera in space corresponding to the SE image under each viewpoint, and the depth map and pixel-by-pixel depth confidence map corresponding to the SE image under each viewpoint are calculated.
[0010] (3) The neural field model is used as the implicit representation model of the three-dimensional shape, and a differentiable BSE forward model is introduced to achieve reconstruction through a three-stage training strategy:
[0011] Phase 1: Based on the depth map and confidence map generated in step (2), construct a weighted depth loss function to minimize the depth map obtained by sampling the light rays corresponding to the pixels of the SE image in the neural field and rendering differentiable volumes. The neural field model is initially trained based on the difference between the actual depth map z and the actual depth map z.
[0012] Phase 2: Using a differentiable BSE forward model, a four-quadrant backscattered electron BSE image is generated with the surface normal obtained based on the current neural field model as input. The BSE image loss between the generated BSE image and the actual BSE image (mean absolute error between the generated BSE image and the real BSE image) is minimized, and the neural field model and the BSE forward model are trained simultaneously.
[0013] Phase 3: Introduce an iterative shadow removal mechanism, dynamically label shadow regions, update shadow region masks to exclude their weights in BSE image loss, and alternately optimize the neural field model, BSE forward model, and shadow region masks to form a self-reinforcing closed loop.
[0014] (4) After training, the three-dimensional mesh surface is extracted from the output of the neural field model to generate a high-fidelity micro-nano structure three-dimensional model.
[0015] Further, in step (1), the micro-nano sample to be tested is fixed on the tiltable rotating sample stage of the scanning electron microscope and the sample is adjusted to multiple viewing angles; for each viewing angle, a secondary electron detector is used to acquire an SE image and a four-quadrant backscattered electron detector is used to acquire four BSE images to form a multi-view multi-channel image input.
[0016] Further, in step (2), the motion recovery structure algorithm includes:
[0017] Multi-view SE images are treated as images captured by a virtual camera. Feature extraction and matching are performed, and sparse 3D point clouds are reconstructed through triangulation to obtain the pose of the virtual camera in space corresponding to each viewpoint SE image.
[0018] Dense reconstruction of sparse 3D point clouds yields a rough 3D surface model.
[0019] Extract the depth map z and pixel-by-pixel depth confidence map w from each viewpoint of the rough 3D surface model.
[0020] Furthermore, a multilayer perceptron network is used to represent the symbolic distance function SDF in three-dimensional space as a neural field model.
[0021] Furthermore, Phase One specifically includes:
[0022] Based on the pose of the virtual camera in space corresponding to the SE image from each viewpoint, randomly sample pixels in the SE image and convert them into light rays in space;
[0023] Three-dimensional coordinate points are sampled along the light rays, and their three-dimensional coordinates are input into the neural field model to obtain the SDF value corresponding to each sampling point;
[0024] The volume rendering method is used to accumulate SDF values to obtain the depth values corresponding to each pixel, resulting in a generated depth map.
[0025] An SDF regularization term is introduced on the basis of depth confidence-weighted depth loss to simultaneously minimize both the depth confidence-weighted depth loss and the SDF regularization term, and the neural field model is initially trained.
[0026] Furthermore, the depth loss and SDF regularization term, weighted by depth confidence, are as follows:
[0027]
[0028] in, This represents the depth loss weighted by depth confidence. This represents the SDF regularization term, where M is the number of sampled rays during training, j is the index of the sampled ray, and w j It is the depth confidence of the pixel corresponding to the j-th sampling ray. z j These are the depth values of the pixel corresponding to the j-th sampling ray in the generated depth map and the depth value in the real depth map, respectively, where N is the number of sampling points on each ray. SDF values s generated by the neural field model jk gradient, s jk It is the SDF value corresponding to the k-th sampling point on the j-th sampling ray.
[0029] Furthermore, Phase Two specifically includes:
[0030] Construct a differentiable BSE forward model for each quadrant:
[0031]
[0032] Where i is the quadrant number, It is a forward model that maps the surface normal n to the BSE image corresponding to quadrant i, θ n It is the angle between the surface normal n and the Z-axis. It is the angle between the surface normal n projected onto the XY plane and the X-axis. d is the angle between the vector projected from the origin to the center of quadrant i onto the XY plane and the X-axis. i c i e i It is a learnable parameter related to quadrant i, R(θ) n () is the angle θ n The relevant parameters are calculated using a learnable quartic polynomial R(θ) with angle θ as input.
[0033] The gradient of the SDF value generated by the neural field model is calculated to obtain the normal of the sampling point. The surface normal corresponding to the ray is obtained by accumulating the normals of the sampling points using a volume rendering method. With surface normal Calculate for input Generate a four-quadrant backscattered electron (BSE) image;
[0034] A forward model regularization term is introduced based on the BSE image loss. To simultaneously minimize BSE image loss and forward model regularization term With the loss term of Phase 1 as the objective, the neural field model and the BSE forward model are trained simultaneously.
[0035] Furthermore, the forward model regularization term for:
[0036]
[0037] in, is the forward model regularization term, Var(·) is used to calculate the variance of the same set of learnable parameters in the four quadrants, and c, d, and e are respectively a set of learnable parameters in the four quadrants.
[0038] Furthermore, Phase Three specifically includes:
[0039] Based on the error dynamic labeling generated from the BSE image, high-error areas are designated as shadow areas, resulting in the shadow area mask S. i :
[0040]
[0041] Where α is a preset hyperparameter, d i It is in the current training process Trainable parameters, It is the surface normal Mapped to the BSE image intensity value corresponding to quadrant i, b i S is the true BSE image intensity value in quadrant i; i S is the shadow region mask for quadrant i. For each sampled ray, if it belongs to the shadow region, then S... ij =0, otherwise S ij =1;
[0042] A shadow region mask is introduced into the BSE image loss function to exclude its weight in the BSE image loss function, resulting in the BSE image loss function after introducing the shadow region mask.
[0043] In each round of training, the goal is to simultaneously minimize the BSE image loss after introducing the shadow region mask and the forward model regularization term. With the loss term of Phase 1 as the objective, the neural field model and the BSE forward model are trained simultaneously. The BSE image loss after introducing the shadow region mask is calculated based on the shadow region mask updated in the previous round. Then, a new shadow region mask is calculated based on the updated parameters for the next round of loss calculation.
[0044] Secondly, this invention proposes a three-dimensional reconstruction system for micro-nano structures based on scanning electron microscopy, which is used to realize the above-mentioned three-dimensional reconstruction method for micro-nano structures based on scanning electron microscopy.
[0045] The beneficial effects of this invention are:
[0046] 1) This invention uses neural fields as the basis for 3D reconstruction of SEM images, effectively integrating geometric and photometric information in SEM images. Unlike existing methods that mainly rely on discrete representations such as height maps, neural fields have stronger expressive power and can more accurately recover discontinuous micro-nano surface structures with drastic curvature changes.
[0047] 2) This invention eliminates the need for samples with known geometries to calibrate detector parameters, enabling self-calibration during training. This significantly simplifies the reconstruction process and improves the ease of use of the method.
[0048] 3) This invention can automatically detect and remove shadow areas in SEM images, thereby effectively reducing the interference of shadows on surface gradient estimation and morphology reconstruction. Unlike existing methods that produce severe artifacts and reconstruction distortion in shadow areas, this invention can directly process SEM images with obvious shadows. The generated 3D model has high surface accuracy, strong integrity, and significant shadow suppression effect, thus broadening the application boundaries of SEM 3D reconstruction methods under complex sample conditions. Attached Figure Description
[0049] Figure 1 This is an overall framework diagram of the present invention;
[0050] Figure 2 This invention performs three-dimensional reconstruction of pollen grains and compares the results with those of traditional motion-restorative structural reconstruction methods. Detailed Implementation
[0051] The present invention will be further described and illustrated below with reference to specific embodiments. The embodiments described are merely examples of the content of this disclosure and do not limit the scope of the invention. The technical features of each embodiment in the present invention can be combined accordingly, provided that there is no mutual conflict.
[0052] This invention provides a method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy. This method uses multi-view, multi-detector two-dimensional SEM images (including secondary electron images and four-quadrant backscattered electron images) as input, combined with a neural field three-dimensional representation. Without requiring additional parameter calibration, it can automatically separate large-area shadows caused by sample geometric occlusion in the images during training, significantly improving the accuracy of three-dimensional reconstruction of complex micro / nano structures; for example... Figure 1 As shown, the method includes the following steps:
[0053] Step 1: Multi-view, multi-detector SEM data acquisition;
[0054] Step 2: Initial geometric estimation based on the structure-of-motion algorithm;
[0055] Step 3: SEM 3D Reconstruction Method Based on Neural Field Representation.
[0056] In a specific embodiment of the present invention, the multi-view, multi-detector SEM data acquisition in step 1 specifically includes:
[0057] The micro / nano sample to be tested is fixed on the tiltable and rotatable sample stage built into the SEM system. The micro / nano sample to be reconstructed is placed at different angles by the control software of the SEM system so that the information of the sample in all directions can be fully captured. For each viewpoint, one SE image and four BSE images are acquired by using the secondary electron (SE) detector and the four-quadrant backscattered electron (4Q-BSE) detector that are standard in common SEM systems. The multi-channel images of all viewpoints constitute the basis for the subsequent input of this invention.
[0058] In a specific embodiment of the present invention, the initial geometric estimation based on the motion recovery structure algorithm in step 2 is specifically as follows:
[0059] The structure-of-motion (SOG) algorithm is employed, treating multi-view SE images as images captured by a virtual camera. Feature extraction and matching are performed on the multi-view SE images, and a sparse 3D point cloud is reconstructed using triangulation. Simultaneously, the pose of the virtual camera in space corresponding to the SE image in each viewpoint is obtained. Then, dense reconstruction is further performed on this basis to obtain a relatively complete but coarse 3D surface model with micro-nano structures. Subsequently, the depth map z and the corresponding pixel-wise depth confidence map w are extracted from the coarse 3D surface model for each viewpoint as geometric prior inputs for subsequent neural network training.
[0060] In a specific embodiment of the present invention, the SEM three-dimensional reconstruction method based on neural field representation described in step 3 is as follows:
[0061] This invention employs a neural field as an implicit representation model for 3D topography, fusing geometric information provided by the structure-of-motion reconstruction algorithm initialization with photometric information from 4Q-BSE images to achieve high-fidelity 3D reconstruction of complex micro / nano structures. The neural field model represents the signed distance function (SDF) of the entire 3D space through a multilayer perceptron network, taking 3D coordinates as input and outputting the SDF value at that coordinate position. Its network parameters can be obtained through end-to-end training. This invention proposes a three-stage neural field training strategy, combining a differentiable BSE forward generation model and an iterative shadow culling mechanism to effectively fuse information from both geometric and photometric dimensions. The neural field network and the BSE forward generation model undergo the following three-stage training process sequentially.
[0062] Step 3.1 Geometric Initialization Phase:
[0063] Based on the pose of the virtual camera in space corresponding to the SE image from each viewpoint, pixels in the SE image are randomly sampled and converted into rays in space. Three-dimensional coordinate points are sampled along the rays, and these coordinates are input into the neural field model to obtain the SDF value s corresponding to each sampled point. Then, a volume rendering method is used to accumulate the SDF values to obtain the depth value corresponding to each pixel.
[0064] In this stage, only the depth map z generated by the structure of motion recovery algorithm is used as the supervision signal, and a depth loss function weighted with confidence w is constructed. Minimize the depth value obtained by sampling light rays corresponding to image pixels in the neural field and rendering differentiable volumes. The difference between z and z is used to perform preliminary training on the neural field model through backpropagation, enabling the neural field to fit the overall surface contour structure of the micro-nano sample.
[0065] In this embodiment, the loss function for this stage is:
[0066]
[0067] Where M is the number of sampling rays during training, j is the index of the sampling ray, and w j It represents the depth confidence of the pixel corresponding to the j-th sampling ray. In addition to minimizing the depth loss, a regularization term for the SDF value is also introduced during training. This ensures that the neural field learns a reasonable SDF value. In this embodiment, the SDF regularization term... for:
[0068]
[0069] in, It is the SDF value s jk gradient, s jk It is the SDF value corresponding to the k-th sampling point on the j-th sampling ray, N is the number of sampling points on each ray, and k is the index of the sampling point.
[0070] The training process at this stage needs to minimize simultaneously and This provides stable geometric initialization for subsequent training phases.
[0071] Step 3.2 BSE photometric-guided joint optimization phase:
[0072] At this stage, 4Q-BSE images are introduced as additional supervision information; in the network forward inference process, a differentiable BSE forward model is constructed. The surface normal n can be mapped to the pixel value of the 4Q-BSE image, and its complete expression is as follows:
[0073]
[0074] Where i is the quadrant number of the 4Q-BSE detector, and the value includes i∈{A,B,C,D}; It maps the surface normal n to the forward model θ of the BSE image corresponding to quadrant i. n It is the angle between the surface normal n and the Z-axis, which is also the angle between the opposite incident directions of the electron beam; It is the angle between the surface normal n projected onto the XY plane and the X-axis; d is the angle between the vector projected from the origin to the center of quadrant i onto the XY plane and the X-axis; i c i e i All are learnable parameters related to quadrant i; R(θ) n () is the angle θ n The relevant parameters are calculated using a learnable fourth-order polynomial R(θ) with angle θ as input. p1, p2, p3, and p4 are the learnable parameters, and the formula for R(θ) is:
[0075] R(θ) = 1 + p1θ + p2θ 2 +p3θ 3 +p4θ 4
[0076] During training, the gradient of the SDF value of the neural field can be used to obtain the normal of the coordinate point. Along the sampled ray (the same as the sampled ray in step 3.1, because the SE image and 4Q-BSE image at each viewpoint are perfectly aligned, corresponding to the same virtual camera pose and parameters), the volume rendering method can be used to accumulate the coordinate point normal to obtain the surface normal corresponding to that ray. Then input it into the BSE forward model The intensity values of the 4Q-BSE image are calculated, and the difference between these values and the actual acquired 4Q-BSE image b is used to construct the BSE image loss function.
[0077]
[0078] Since the properties of the four quadrants of a real-world 4Q-BSE detector are very similar, a regularization term is also constructed here. To promote the forward model Consistency of parameters in the four quadrants:
[0079]
[0080] in, It is the surface normal corresponding to the j-th sampling ray. It is the surface normal Mapped to the BSE image intensity value corresponding to quadrant i, b ij is the true BSE image intensity value of quadrant i corresponding to the j-th sampled ray, Var(·) represents the calculation of the variance of the same set of learnable parameters in the four quadrants, and c, d, and e are respectively a set of learnable parameters in the four quadrants.
[0081] The training process in this stage uses the training results from step 3.1 as the initial value and needs to minimize the following values simultaneously. and regularization terms and Backpropagation enables simultaneous processing of neural field parameters and the BSE forward model. The learnable parameters in the model are optimized to transfer the photometric information in the 4Q-BSE image to the geometric model of the implicit field representation, thereby achieving effective fusion of geometric and photometric information and significantly improving the ability to reconstruct local details of micro and nanostructure surfaces.
[0082] Step 3.3 Automatic identification and removal of shadow areas:
[0083] The 4Q-BSE image contains numerous shadow regions caused by sample geometric occlusion. The pixel intensity of these shadow regions is not directly related to the surface gradient, negatively impacting training. Therefore, a shadow removal mechanism is introduced in this stage. Based on the error in BSE image generation, high-error regions are dynamically labeled as shadow regions, yielding the shadow region mask S. i :
[0084]
[0085] Where α is the hyperparameter set in the experiment, and d i For the current training process Trainable parameters, It is the surface normal Mapped to the BSE image intensity value corresponding to quadrant i, b i S is the true BSE image intensity value in quadrant i; i S is the shadow region mask for quadrant i. For each sampled ray, if it belongs to the shadow region, then S... ij =0, otherwise S ij =1. In this training phase, shadow regions are excluded from the calculation of the BSE image loss function, and the determination of shadow regions is dynamically updated as the training process progresses. Removing shadow regions from the training can improve the geometry of the neural site representation and the BSE feedforward model. The accuracy.
[0086] On the other hand, more accurate surface geometry and forward modeling can lead to more precise determination of shadow regions, and the two reinforce each other. In this embodiment, the BSE image loss function after introducing shadow region masking is:
[0087]
[0088] The training process in this stage uses the training results from step 3.2 as the initial value and needs to minimize the following values simultaneously. With mask and regularization terms and This stage further improves the accuracy of the reconstructed geometry.
[0089] It should be noted that in each round of training, the goal is to minimize... and regularization terms and Update the parameters of the neural field model and the BSE forward model F for the target, where The shadow region mask is calculated based on the previously updated shadow region mask; then, a new shadow region mask is calculated based on the updated parameters for the next round of loss calculation.
[0090] After completing the above three-stage training, a high-quality three-dimensional mesh surface is extracted from the trained neural field model: specifically, the Marching Cubes algorithm is used to extract isosurfaces from the SDF, which can obtain a micro-nano structure three-dimensional model with arbitrary resolution.
[0091] Example
[0092] To demonstrate the effectiveness of this invention, the invention achieved three-dimensional reconstruction of the sample surface structure on peach blossom pollen particles. The effectiveness of this invention was compared with two-dimensional images captured by SEM and three-dimensional models reconstructed by traditional motion recovery structure methods.
[0093] This embodiment applies the method of the present invention to the ZEISS Gemini 560 SEM system.
[0094] The experimental sample used in this embodiment was Okubo peach pollen grains, which were fixed to an aluminum cylindrical base using carbon conductive adhesive. The sample surface was sputtered with gold at 15 mA for 60 seconds using a Hitachi E-1010 ion sputtering system. The treated sample was then mounted into a SEM system, and multi-view, multi-detector SEM images (including SE and 4Q-BSE images) were acquired.
[0095] First, the motion recovery structure is initialized, and then combined with the neural field training and reconstruction process proposed in this invention, the input SEM image is finally reconstructed into a complete three-dimensional model of the pollen grain surface.
[0096] Figure 2 This paper presents the original SEM image of the sample, the 3D model reconstructed using the Structure-of-Motion (SOM) algorithm, and the 3D model reconstructed using the method of this invention. The comparison shows that the SEM image is essentially a two-dimensional projection of the sample and cannot provide a quantitative 3D representation of the sample surface; the SOM algorithm has poor reconstruction accuracy and cannot reveal the texture details of the pollen surface; in contrast, the surface model reconstructed by the method of this invention not only accurately restores the overall geometry of the pollen but also clearly captures the complex and dense fine texture of the pollen surface, showing a high degree of consistency with the original SEM image, demonstrating the superiority of the method of this invention. This invention can be widely applied in fields such as micro / nano structure analysis of biological samples and precision control in micro / nano manufacturing, and has high practical value.
[0097] Based on the same inventive concept, this invention also proposes a three-dimensional reconstruction system for micro / nano structures based on scanning electron microscopy, the system comprising:
[0098] The multi-view multi-detector SEM data acquisition module is used to acquire multi-view multi-detector SEM images of the micro-nano sample under test. Each view includes a secondary electron SE image and a four-quadrant backscattered electron BSE image.
[0099] The initial geometry estimation module based on the structure-of-motion (SOMO) algorithm is used to obtain the pose of the virtual camera in space corresponding to the SE image at each viewpoint using the SOMO algorithm, and to calculate the depth map and pixel-wise depth confidence map corresponding to the SE image at each viewpoint.
[0100] The 3D reconstruction module based on neural field representation uses a neural field model as an implicit representation model of the 3D topography and introduces a differentiable BSE forward model. Reconstruction is achieved through a three-stage training strategy.
[0101] Phase 1: Based on the depth map and confidence map generated in step (2), construct a weighted depth loss function to minimize the difference between the depth map z^ obtained by sampling the light corresponding to the SE image pixel in the neural field and differentiable volume rendering and the actual depth map z, and initially train the neural field model;
[0102] Phase 2: Using a differentiable BSE forward model, a four-quadrant backscattered electron BSE image is generated with the surface normal obtained based on the current neural field model as input. The BSE image loss between the generated BSE image and the actual BSE image is minimized, and the neural field model and the BSE forward model are trained simultaneously.
[0103] Phase 3: Introduce an iterative shadow removal mechanism, dynamically label shadow regions, update shadow region masks to exclude their weights in BSE image loss, and alternately optimize the neural field model, BSE forward model, and shadow region masks to form a self-reinforcing closed loop.
[0104] After training, the three-dimensional mesh surface is extracted from the output of the neural field model to generate a high-fidelity micro-nano structure three-dimensional model.
[0105] For the system embodiments, since they basically correspond to the method embodiments, relevant details can be found in the descriptions of the method embodiments; the implementation methods of the remaining modules will not be repeated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0106] The system embodiments of the present invention can be applied to any device with data processing capabilities, such as a computer or other similar device. The system embodiments can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution.
[0107] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. Those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy, characterized in that, Includes the following steps: (1) Collect multi-view, multi-detector SEM images of the micro-nano sample to be tested, each view including secondary electron SE image and four-quadrant backscattered electron BSE image. (2) The motion recovery structure algorithm is used to obtain the pose of the virtual camera in space corresponding to the SE image under each viewpoint, and the depth map and pixel-by-pixel depth confidence map corresponding to the SE image under each viewpoint are calculated. (3) The neural field model is used as the implicit representation model of the three-dimensional shape, and a differentiable BSE forward model is introduced to achieve reconstruction through a three-stage training strategy: Phase 1: Based on the depth map and confidence map generated in step (2), construct a weighted depth loss function to minimize the depth map obtained by sampling the light rays corresponding to the pixels of the SE image in the neural field and rendering differentiable volumes. The neural field model is initially trained based on the difference between the actual depth map z and the actual depth map z. Phase 2: Using a differentiable BSE forward model with the surface normal obtained based on the current neural field model as input, generate a four-quadrant backscattered electron BSE image, minimize the BSE image loss between the generated BSE image and the actual BSE image, and train the neural field model and the BSE forward model simultaneously. Phase 3: Introduce an iterative shadow removal mechanism, dynamically label shadow regions, update shadow region masks to exclude their weights in BSE image loss, and alternately optimize the neural field model, BSE forward model, and shadow region masks to form a self-reinforcing closed loop. (4) After training, the three-dimensional mesh surface is extracted from the output of the neural field model to generate a high-fidelity micro-nano structure three-dimensional model.
2. The method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy according to claim 1, characterized in that, In step (1), the micro / nano sample to be tested is fixed on the tiltable rotating sample stage of the scanning electron microscope and the sample is adjusted to multiple viewing angles. For each viewing angle, a secondary electron detector is used to acquire an SE image and a four-quadrant backscattered electron detector is used to acquire four BSE images to form a multi-view, multi-channel image input.
3. The method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy according to claim 1, characterized in that, In step (2), the motion recovery structure algorithm includes: Multi-view SE images are treated as images captured by a virtual camera. Feature extraction and matching are performed, and sparse 3D point clouds are reconstructed through triangulation to obtain the pose of the virtual camera in space corresponding to each viewpoint SE image. Dense reconstruction of sparse 3D point clouds yields a rough 3D surface model. Extract the depth map z and pixel-by-pixel depth confidence map w from each viewpoint of the rough 3D surface model.
4. The method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy according to claim 1, characterized in that, A multilayer perceptron network is used to represent the symbolic distance function SDF in three-dimensional space as a neural field model.
5. The method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy according to claim 1, characterized in that, Phase One specifically includes: Based on the pose of the virtual camera in space corresponding to the SE image from each viewpoint, randomly sample pixels in the SE image and convert them into light rays in space; Three-dimensional coordinate points are sampled along the light rays, and their three-dimensional coordinates are input into the neural field model to obtain the SDF value corresponding to each sampling point; The volume rendering method is used to accumulate SDF values to obtain the depth values corresponding to each pixel, resulting in a generated depth map. An SDF regularization term is introduced on the basis of depth confidence-weighted depth loss to simultaneously minimize both the depth confidence-weighted depth loss and the SDF regularization term, and the neural field model is initially trained.
6. The method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy according to claim 5, characterized in that, The depth loss and SDF regularization term, which are weighted by depth confidence, are as follows: in, This represents the depth loss weighted by depth confidence. This represents the SDF regularization term, where M is the number of sampled rays during training, j is the index of the sampled ray, and w j It is the depth confidence of the pixel corresponding to the j-th sampling ray. z j These are the depth values of the pixel corresponding to the j-th sampling ray in the generated depth map and the depth value in the real depth map, respectively, where N is the number of sampling points on each ray. SDF values s generated by the neural field model jk gradient, s jk It is the SDF value corresponding to the k-th sampling point on the j-th sampling ray.
7. The method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy according to claim 1, characterized in that, Phase Two specifically includes: Construct a differentiable BSE forward model for each quadrant: Where i is the quadrant number, It is a forward model that maps the surface normal n to the BSE image corresponding to quadrant i, θ n It is the angle between the surface normal n and the Z-axis. It is the angle between the surface normal n projected onto the XY plane and the X-axis. d is the angle between the vector projected from the origin to the center of quadrant i onto the XY plane and the X-axis. i c i e i It is a learnable parameter related to quadrant i, R(θ) n () is the angle θ n The relevant parameters are calculated using a learnable quartic polynomial R(θ) with angle θ as input. The gradient of the SDF value generated by the neural field model is calculated to obtain the normal of the sampling point. The surface normal corresponding to the ray is obtained by accumulating the normals of the sampling points using a volume rendering method. With surface normal Calculate for input Generate a four-quadrant backscattered electron (BSE) image; A forward model regularization term is introduced based on the BSE image loss. To simultaneously minimize BSE image loss and forward model regularization term With the loss term of Phase 1 as the objective, the neural field model and the BSE forward model are trained simultaneously.
8. The method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy according to claim 7, characterized in that, Forward model regularization term for: in, is the forward model regularization term, Var(·) is used to calculate the variance of the same set of learnable parameters in the four quadrants, and c, d, and e are respectively a set of learnable parameters in the four quadrants.
9. The method for three-dimensional reconstruction of micro / nano structures based on scanning electron microscopy according to claim 7, characterized in that, Phase Three specifically includes: Based on the error dynamic labeling generated from the BSE image, high-error areas are designated as shadow areas, resulting in the shadow area mask S. i : Where α is a preset hyperparameter, d i It is in the current training process Trainable parameters, It is the surface normal Mapped to the BSE image intensity value corresponding to quadrant i, b i S is the true BSE image intensity value in quadrant i; i S is the shadow region mask for quadrant i. For each sampled ray, if it belongs to the shadow region, then S... ij =0, otherwise S ij =1; A shadow region mask is introduced into the BSE image loss function to exclude its weight in the BSE image loss function, resulting in the BSE image loss function after introducing the shadow region mask. In each round of training, the goal is to simultaneously minimize the BSE image loss after introducing the shadow region mask and the forward model regularization term. With the loss term of Phase 1 as the objective, the neural field model and the BSE forward model are trained simultaneously. The BSE image loss after introducing the shadow region mask is calculated based on the shadow region mask updated in the previous round. Then, a new shadow region mask is calculated based on the updated parameters for the next round of loss calculation.
10. A three-dimensional reconstruction system for micro / nano structures based on scanning electron microscopy, used to implement the three-dimensional reconstruction method for micro / nano structures as described in claim 1, characterized in that, The system includes: The multi-view multi-detector SEM data acquisition module is used to acquire multi-view multi-detector SEM images of the micro-nano sample under test. Each view includes a secondary electron SE image and a four-quadrant backscattered electron BSE image. The initial geometry estimation module based on the structure-of-motion (SOMO) algorithm is used to obtain the pose of the virtual camera in space corresponding to the SE image at each viewpoint using the SOMO algorithm, and to calculate the depth map and pixel-wise depth confidence map corresponding to the SE image at each viewpoint. The 3D reconstruction module based on neural field representation uses a neural field model as an implicit representation model of the 3D topography and introduces a differentiable BSE forward model. Reconstruction is achieved through a three-stage training strategy. Phase 1: Based on the depth map and confidence map generated in step (2), construct a weighted depth loss function to minimize the depth map obtained by sampling the light rays corresponding to the pixels of the SE image in the neural field and rendering differentiable volumes. The neural field model is initially trained based on the difference between the actual depth map z and the actual depth map z. Phase 2: Using a differentiable BSE forward model, a four-quadrant backscattered electron BSE image is generated with the surface normal obtained based on the current neural field model as input. The BSE image loss between the generated BSE image and the actual BSE image is minimized, and the neural field model and the BSE forward model are trained simultaneously. Phase 3: Introduce an iterative shadow removal mechanism, dynamically label shadow regions, update shadow region masks to exclude their weights in BSE image loss, and alternately optimize the neural field model, BSE forward model, and shadow region masks to form a self-reinforcing closed loop. After training, the three-dimensional mesh surface is extracted from the output of the neural field model to generate a high-fidelity micro-nano structure three-dimensional model.