A microtopography analysis method based on wavelet multilayer frequency band decomposition

By employing wavelet multi-band decomposition, combined with multi-focus image sequences and convolutional networks, high-precision 3D topography reconstruction and texture preservation were achieved. This solved the problems of accuracy and texture reconstruction in traditional methods for micro-topography reconstruction, and improved the robustness and accuracy of micro-topography analysis.

CN122115719APending Publication Date: 2026-05-29SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANXI UNIV
Filing Date
2026-02-06
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing 3D topography reconstruction methods struggle to simultaneously achieve high-precision geometric depth information and high-fidelity reconstruction of microscopic surface textures under the same architecture. In particular, traditional methods suffer from low computational efficiency and poor positioning accuracy under fine textures and complex lighting conditions.

Method used

A wavelet-based multi-band decomposition method is adopted. By acquiring multi-focus image sequences, multi-scale features are extracted using a convolutional network, wavelet transform and feature reconstruction are performed, and high-fidelity full-focus fusion images and depth maps are generated by combining high-frequency details and low-frequency constraints.

Benefits of technology

It achieves simultaneous output of three-dimensional spatial geometric structure reconstruction with micron-level precision and high-fidelity texture images, improving the robustness and accuracy of microscopic morphology analysis and avoiding interference from optical artifacts and defocus noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115719A_ABST
    Figure CN122115719A_ABST
Patent Text Reader

Abstract

The application belongs to the field of three-dimensional reconstruction, and particularly relates to a micro-morphology analysis method based on wavelet multi-layer frequency band decomposition. The method comprises the following steps: firstly, a multi-focus image sequence of a micro scene to be measured is collected, and multi-scale spatial features are obtained through feature coding and projection mapping; then, a two-dimensional discrete wavelet transform method is introduced, the spatial features are decoupled into a low-frequency approximate subband representing global geometry and a high-frequency detail subband representing fine texture, and structure feature correction and directional texture enhancement are respectively implemented; on this basis, inverse discrete wavelet transform and residual fusion are used to construct time-frequency enhanced features, a deep regression network is used to analyze pixel-level focus probability distribution to generate a continuous depth map; finally, a cascading mechanism driven by three-dimensional geometric reconstruction for scene morphology perception is established, and depth geometric edge constraints are used to reversely guide the high-fidelity generation of a full-focus image. Through the wavelet multi-layer frequency band decomposition strategy, the application effectively solves the problem of confusion between texture details and out-of-focus noise in micro imaging, and realizes the synchronous robust analysis of high-precision geometric structure and full-depth clear texture of micro morphology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of three-dimensional reconstruction technology, specifically relating to a micro-morphological analysis method based on wavelet multi-band decomposition. Background Technology

[0002] Microscopic 3D topography reconstruction serves as a crucial link between the microscopic physical world and digital intelligent analysis, encompassing two core dimensions: 3D structure reconstruction and scene topography perception. With the deepening development of intelligent computing, 3D reconstruction technology is continuously delving into the ultra-microscopic realm. Achieving high-precision geometric spatial reconstruction while simultaneously restoring the texture of microscopic surface structures remains one of the significant challenges in current microscopic 3D topography reconstruction.

[0003] Existing 3D topography reconstruction methods are generally divided into two major paradigms: active physical measurement and passive computational perception. Active measurement methods (such as laser confocal microscopy and structured light interferometry) typically rely on explicit physical projection mechanisms, calculating spatial depth relationships by projecting laser arrays or modulating light fields onto the scene under test. Although these methods can acquire high-precision geometric depth information through physical ranging, they are limited by a single ranging mechanism and cannot capture microscopic surface texture structures, making it difficult to meet the requirements of high-fidelity 3D topography perception. In addition, the high dependence of these methods on precision instruments also restricts their application in lightweight deployment scenarios. In contrast, passive perception analysis is based on a general optical microscopic imaging architecture and aims to extract implicit visual cues from 2D image sequences to infer depth information. It utilizes the inherent shallow depth-of-field characteristics of optical systems in microscopic scenes as cues, and constructs a mapping relationship between geometric topography and texture imaging by acquiring multi-focus image stacks. Therefore, compared to the limitations of active measurement methods in depth estimation, the passive perception paradigm has the theoretical potential to simultaneously acquire depth information and scene texture under the same architecture, which is more in line with the comprehensive needs of 3D shape reconstruction and scene shape perception collaborative analysis.

[0004] While the passive perception paradigm possesses theoretical advantages in acquiring multidimensional information, its practical applications are often constrained by the difficulty in quantitatively estimating depth range. This makes it challenging to extract precise 3D cues from image sequences, hindering current technologies from simultaneously meeting the dual requirements of high-precision 3D shape reconstruction and high-fidelity scene shape perception within the same architecture. Existing shape analysis methods mainly fall into two categories: spatial domain analysis and frequency domain analysis. Spatial domain analysis algorithms struggle to balance computational efficiency and edge localization accuracy. On one hand, to obtain high-precision focus response, algorithms typically rely on multi-scale cost aggregation processes, significantly increasing computational overhead. On the other hand, when processing fine textures or highly reflective areas, purely spatial domain features are insufficient to effectively decouple surface texture details from defocus noise. This visual feature confusion easily leads to depth localization errors, failing to meet the requirements of high-fidelity reconstruction. While frequency domain analysis algorithms can capture image texture features through spectral analysis, conventional global frequency domain transformations suffer from a lack of spatial localization information, making it difficult to accurately characterize the distribution of frequency features locally in the image. While frequency domain methods can effectively characterize the global sharpness of images, they lack local spatial localization capabilities, making it difficult to achieve high-precision microscopic topography analysis. Furthermore, traditional frequency domain filters are prone to spectral aliasing when dealing with non-stationary microscopic surface signals, leading to the masking or misjudgment of minute topographic features.

[0005] In summary, the main challenges facing existing microscopic 3D topography reconstruction techniques are: while spatial domain analysis algorithms have high computational efficiency, they cannot achieve fine-grained edge reconstruction; frequency domain analytical methods have good global structure preservation characteristics, but they cannot stably reconstruct homogeneous regions. Therefore, how to utilize the complementary information of these two types of reconstruction methods to achieve high-precision 3D topography and highly reliable surface texture preservation remains a technical problem to be solved. Summary of the Invention

[0006] To overcome the shortcomings of existing micromorphological analysis methods, the purpose of this invention is to provide a micromorphological analysis method based on wavelet multi-band decomposition, comprising the following steps:

[0007] Step 1: Acquire a multi-focus image sequence of the microscopic scene to be tested. Suppose the sequence contains Frame image, the size of a single frame image is The number of channels is Each frame in the sequence corresponds to a different physical focus position on the optical axis, and the image sequence is represented in tensor dimension as follows: ;

[0008] Step 2: Process the multifocus image sequence obtained in Step 1 The convolutional network feature extractor in equation (1) extracts multi-scale basic feature sequences. ,

[0009] (1)

[0010] in, It is a feature extractor for convolutional networks;

[0011] Step 3: Calculate the multi-scale basic feature sequence obtained in Step 2. By aligning the channel dimensions and performing nonlinear mapping using equation (2), the corresponding scale projection feature matrix is ​​obtained. ,

[0012] (2)

[0013] in, This represents the convolution operation. Represents a nonlinear activation function;

[0014] Step 4: Apply the multi-scale projection feature matrix obtained in Step 3. The low-frequency approximate subband is obtained through equation (3). and high-frequency detail subband set ,

[0015] (3)

[0016] in, This represents the two-dimensional discrete wavelet transform operation. Represents the low-frequency approximate component. These represent high-frequency texture details in the horizontal, vertical, and diagonal directions, respectively.

[0017] Step 5: Decompose the high-frequency detail subband set obtained in Step 4. Equation (4) is used to stitch together the high-frequency subbands in the three directions along the channel dimension, and multi-scale convolutional units are used to extract texture features under a large receptive field. ,

[0018] (4)

[0019] in, This indicates a splicing operation. Used to capture long-range texture dependencies in the frequency domain. It is a non-linear activation function;

[0020] Step 6: Decompose the low-frequency approximate subband obtained in Step 4 Equation (5) is used to smooth and extract features from low-frequency components using convolution operations, resulting in the corrected low-frequency features. ,

[0021] (5)

[0022] in, This represents the convolution operation. It is a non-linear activation function;

[0023] Step 7: Using the inverse discrete wavelet transform, the processed high-frequency and low-frequency features are reconstructed back into the spatial domain using equation (6) to obtain the frequency domain sensing features. ,

[0024] (6)

[0025] in, This is the inverse discrete wavelet transform operation;

[0026] Step 8: Reconstruct features Compared with the original projection features Residual fusion is performed using equation (7) to generate time-frequency enhancement features. ,

[0027] (7)

[0028] in, These are learnable weighting coefficients. For layer normalization operation;

[0029] Step 9: Apply the time-frequency enhancement features obtained in Step 8 Equation (8) is used to stack and regularize features along the sequence dimension.

[0030] (8)

[0031] in, This indicates stacking along the depth of focus axis. It is a regularized network composed of three-dimensional convolutions, used to smooth the focusing probability distribution between sequences;

[0032] Step 10: Focus the volume obtained in Step 9 The probability distribution of each pixel at different focal positions is calculated using equation (9). ,

[0033] (9)

[0034] Where exp is the natural exponential function. Represents the two-dimensional pixel coordinates on the image plane. , , This represents the index of a discrete focusing layer in a focused image sequence. , For summation index variables, Indicates the pixel position Place, No. The focus response value corresponding to each focus layer. It is a normalized exponential function;

[0035] Step 11: Apply the probability distribution obtained in Step 10 A continuous depth map is obtained using equation (10). ,

[0036] (10)

[0037] in, Indicates the first The physical focus depth value corresponding to the frame image. Represents the calculated pixels The final depth value at that location;

[0038] Step 12: Calculate the focusing probability distribution from Step 10. The original image sequence is processed by equation (11). Weighted fusion is performed to obtain a high-fidelity fully focused fused image. ,

[0039] (11)

[0040] in, Represents the first image in the original image sequence. Frame in pixel Pixel value at that location, Represents pixels Full-focus fused image at the location;

[0041] Step 13: Obtain the depth map from Step 11 The fully focused fused image obtained in step 12 Point-to-point mapping yields the 3D structure of the scene.

[0042] Compared with the prior art, the present invention has the following advantages:

[0043] (1) This invention achieves simultaneous high-precision analysis of two core tasks: three-dimensional shape reconstruction and scene shape perception. Compared with the limitations of traditional passive perception methods in processing fine textures and complex lighting, this invention can not only obtain three-dimensional spatial geometric structures with micron-level precision, but also output high-fidelity fully focused texture images simultaneously. This ensures that the model can still take into account the accuracy of depth topology and the clarity of surface texture without introducing expensive global iteration calculations, which significantly improves the overall robustness of micro-shape analysis.

[0044] (2) This invention accurately eliminates the interference of optical artifacts and defocus noise by using high-frequency subbands, while using low-frequency subbands to provide robust geometric constraints, effectively avoiding the problem of visual feature confusion in traditional spatial domain methods, and improving the high-precision processing performance of three-dimensional shape analysis in complex microscopic scenes. Attached Figure Description

[0045] Figure 1 This is a flowchart of a micromorphological analysis method based on wavelet multi-band decomposition;

[0046] Figure 2 This is a network structure diagram of a micromorphological analysis method based on wavelet multi-band decomposition.

[0047] Figure 3 These are 30 samples from the multi-focus image sequence acquired in step 1 of Embodiment 1 of the present invention;

[0048] Figure 4 This refers to the depth map of the scene obtained in step 11 of embodiment 1 of the present invention;

[0049] Figure 5 This refers to the fully focused fused image of the scene obtained in step 12 of Embodiment 1 of the present invention;

[0050] Figure 6 This refers to the three-dimensional structure of the scene obtained in step 13 of Embodiment 1 of the present invention. Detailed Implementation

[0051] Example 1

[0052] like Figure 1 , Figure 2 As shown, a micromorphological analysis method based on wavelet multi-band decomposition includes the following steps:

[0053] Step 1: Acquire a multi-focus image sequence of the microscopic scene to be tested. Suppose the sequence contains Frame image, the size of a single frame image is The number of channels is Each frame in the sequence corresponds to a different physical focus position on the optical axis, and the image sequence is represented in tensor dimension as follows: ,like Figure 3 As shown, ;

[0054] Step 2: Process the multifocus image sequence obtained in Step 1 The convolutional network feature extractor in equation (1) extracts multi-scale basic feature sequences. ,

[0055] (1)

[0056] in, It is a feature extractor for convolutional networks;

[0057] Step 3: Calculate the multi-scale basic feature sequence obtained in Step 2. By aligning the channel dimensions and performing nonlinear mapping using equation (2), the corresponding scale projection feature matrix is ​​obtained. ,

[0058] (2)

[0059] in, This represents the convolution operation. Represents a nonlinear activation function;

[0060] Step 4: Apply the multi-scale projection feature matrix obtained in Step 3. The low-frequency approximate subband is obtained through equation (3). and high-frequency detail subband set ,

[0061] (3)

[0062] in, This represents the two-dimensional discrete wavelet transform operation. Represents the low-frequency approximate component. These represent high-frequency texture details in the horizontal, vertical, and diagonal directions, respectively.

[0063] Step 5: Decompose the high-frequency detail subband set obtained in Step 4. Equation (4) is used to stitch together the high-frequency subbands in the three directions along the channel dimension, and multi-scale convolutional units are used to extract texture features under a large receptive field. ,

[0064] (4)

[0065] in, This indicates a splicing operation. Used to capture long-range texture dependencies in the frequency domain. It is a non-linear activation function;

[0066] Step 6: Decompose the low-frequency approximate subband obtained in Step 4 Equation (5) is used to smooth and extract features from low-frequency components using convolution operations, resulting in the corrected low-frequency features. ,

[0067] (5)

[0068] in, This represents the convolution operation. It is a non-linear activation function;

[0069] Step 7: Using the inverse discrete wavelet transform, the processed high-frequency and low-frequency features are reconstructed back into the spatial domain using equation (6) to obtain the frequency domain sensing features. ,

[0070] (6)

[0071] in, This is the inverse discrete wavelet transform operation;

[0072] Step 8: Reconstruct features Compared with the original projection features Residual fusion is performed using equation (7) to generate time-frequency enhancement features. ,

[0073] (7)

[0074] in, These are learnable weighting coefficients. For layer normalization operation;

[0075] Step 9: Apply the time-frequency enhancement features obtained in Step 8 Equation (8) is used to stack and regularize features along the sequence dimension.

[0076] (8)

[0077] in, This indicates stacking along the depth of focus axis. It is a regularized network composed of three-dimensional convolutions, used to smooth the focusing probability distribution between sequences;

[0078] Step 10: Focus the volume obtained in Step 9 The probability distribution of each pixel at different focal positions is calculated using equation (9). ,

[0079] (9)

[0080] Where exp is the natural exponential function. Represents the two-dimensional pixel coordinates on the image plane. , , This represents the index of a discrete focusing layer in a focused image sequence. , For summation index variables, Indicates the pixel position Place, No. The focus response value corresponding to each focus layer. It is a normalized exponential function;

[0081] Step 11: Apply the probability distribution obtained in Step 10 A continuous depth map is obtained using equation (10). ,like Figure 4 As shown,

[0082] (10)

[0083] in, Indicates the first The physical focus depth value corresponding to the frame image. Represents the calculated pixels The final depth value at that location;

[0084] Step 12: Calculate the focusing probability distribution from Step 10. The original image sequence is processed by equation (11). Weighted fusion is performed to obtain a high-fidelity fully focused fused image. ,like Figure 5 As shown,

[0085] (11)

[0086] in, Represents the first image in the original image sequence. Frame in pixel Pixel value at that location, Represents pixels Full-focus fused image at the location;

[0087] Step 13: Obtain the depth map from Step 11 The fully focused fused image obtained in step 12 Point-to-point mapping yields the 3D structure of the scene, such as... Figure 6 As shown.

Claims

1. A method for analyzing microstructures based on wavelet multi-band decomposition, characterized in that, Includes the following steps: Step 1: Acquire a multi-focus image sequence of the microscopic scene to be tested. Suppose the sequence contains Frame image, the size of a single frame image is The number of channels is Each frame in the sequence corresponds to a different physical focus position on the optical axis, and the image sequence is represented in tensor dimension as follows: ; Step 2: Process the multifocus image sequence obtained in Step 1 The convolutional network feature extractor in equation (1) extracts multi-scale basic feature sequences. , (1) in, It is a feature extractor for convolutional networks; Step 3: Calculate the multi-scale basic feature sequence obtained in Step 2. By aligning the channel dimensions and performing nonlinear mapping using equation (2), the corresponding scale projection feature matrix is ​​obtained. , (2) in, This represents the convolution operation. Represents a non-linear activation function; Step 4: Apply the multi-scale projection feature matrix obtained in Step 3. The low-frequency approximate subband is obtained through equation (3). and high-frequency detail subband set , (3) in, This represents the two-dimensional discrete wavelet transform operation. Represents the low-frequency approximate component. These represent high-frequency texture details in the horizontal, vertical, and diagonal directions, respectively. Step 5: Decompose the high-frequency detail subband set obtained in Step 4. Equation (4) is used to stitch together the high-frequency subbands in the three directions along the channel dimension, and multi-scale convolutional units are used to extract texture features under a large receptive field. , (4) in, This indicates a splicing operation. Used to capture long-range texture dependencies in the frequency domain. It is a non-linear activation function; Step 6: Decompose the low-frequency approximate subband obtained in Step 4 Equation (5) is used to smooth and extract features from low-frequency components using convolution operations, resulting in the corrected low-frequency features. , (5) in, This represents the convolution operation. It is a non-linear activation function; Step 7: Using the inverse discrete wavelet transform, the processed high-frequency and low-frequency features are reconstructed back into the spatial domain using equation (6) to obtain the frequency domain sensing features. , (6) in, This is the inverse discrete wavelet transform operation; Step 8: Reconstruct features Compared with the original projection features Residual fusion is performed using equation (7) to generate time-frequency enhancement features. , (7) in, These are learnable weighting coefficients. For layer normalization operation; Step 9: Apply the time-frequency enhancement features obtained in Step 8 Equation (8) is used to stack and regularize features along the sequence dimension. (8) in, This indicates stacking along the depth of focus axis. It is a regularized network composed of three-dimensional convolutions, used to smooth the focusing probability distribution between sequences; Step 10: Focus the volume obtained in Step 9 The probability distribution of each pixel at different focal positions is calculated using equation (9). , (9) Where exp is the natural exponential function. Represents the two-dimensional pixel coordinates on the image plane. , , This represents the index of a discrete focusing layer in a focused image sequence. , For summation index variables, Indicates the pixel position Place, No. The focus response value corresponding to each focus layer. It is a normalized exponential function; Step 11: Apply the probability distribution obtained in Step 10 A continuous depth map is obtained using equation (10). , (10) in, Indicates the first The physical focus depth value corresponding to the frame image. Represents the calculated pixels The final depth value at that location; Step 12: Calculate the focusing probability distribution from Step 10. The original image sequence is processed by equation (11). Weighted fusion is performed to obtain a high-fidelity fully focused fused image. , (11) in, Represents the first image in the original image sequence. Frame in pixel Pixel value at that location, Represents pixels Full-focus fused image at the location; Step 13: Obtain the depth map from Step 11 The fully focused fused image obtained in step 12 Point-to-point mapping yields the 3D structure of the scene.