An adaptive image focus stacking method, system and pathology slide scanner
By using an adaptive image focus stacking method, the problems of insufficient sharpness assessment, low alignment efficiency, and poor fusion effect in pathological scanning are solved, generating high-quality full-area sharp focus stacked images that meet the needs of pathological diagnosis.
Patent Information
- Application Number
- CN202610313803.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-16
- Publication Date
- 2026-07-03
AI Technical Summary
Existing focus stacking techniques in pathological scanning suffer from insufficient robustness in clarity assessment, low alignment efficiency, and fusion results that do not meet diagnostic requirements, thus affecting the accuracy of pathological diagnosis.
An adaptive image focus stacking method is adopted, including perceptual preprocessing, multi-scale spatial frequency evaluation, adaptive raster registration, intelligent decision fusion and spatial consistency verification, to generate high-quality full-area clear focus stacked images through multi-scale weight maps and sub-pixel level matching.
It improves the robustness of sharpness assessment, optimizes alignment efficiency, meets the dual sharpness requirements of pathological diagnosis for cell morphology and tissue boundaries, and generates high-quality images with seamless visual transitions.
Smart Images

Figure CN122335608A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital pathology technology, specifically to an adaptive image focus stacking method, system, and pathology slide scanner suitable for digital pathology slide scanners, with a particular focus on full-area sharpening of high-magnification microscopic images in digital pathology diagnostic scenarios. Background Technology
[0002] Digital pathology technology is the core support for modern pathological diagnosis. Pathological slide scanners, as key equipment, must scan and convert solid pathological slides (typically 3-5 μm thick) into high-resolution digital images, providing pathologists with visual diagnostic information. Due to the three-dimensional spatial distribution differences of tissues / cells in pathological slides, and the potential for uneven thickness and wrinkles during slide preparation, the depth of field under high-magnification objectives is extremely shallow, typically only 0.5-2 μm. A single scan image cannot cover a clear focal plane of all microstructures, resulting in some cell structures being clear while others are blurred within the same field of view. This directly affects the accuracy of pathologists' judgments on cell morphology and lesion characteristics.
[0003] Focus stacking technology is the core solution to this problem. Its principle is to acquire multiple images from different Z-axis focal planes within the same field of view and then fuse them to generate a clear digital slice covering the entire area. However, existing focus stacking technology faces several technical bottlenecks in practical applications of pathological scanning:
[0004] Insufficient robustness of clarity assessment: Pathological slide images contain interference factors such as staining differences (e.g., red-blue contrast of HE staining), slide impurities, and blurred tissue edges. This makes assessment methods based on a single algorithm (e.g., gradient magnitude calculation) prone to misjudging staining boundaries as noise signals or failing to identify clear details in thin tissue areas (e.g., single-layer epithelial cells), resulting in inaccurate weight allocation and consequently causing local distortion in the fused image.
[0005] Alignment schemes are not suitable for pathological scanning scenarios: Pathological slide scanning requires both high speed and high precision (micrometer level). Existing whole-image alignment schemes have a large computational load, which leads to a serious reduction in scanning efficiency (single slide scanning time is extended by more than 50%). Furthermore, simple alignment schemes cannot compensate for the slight offset of the mechanical platform, resulting in ghosting artifacts in the fused images, especially in densely celled areas (such as tumor tissue), which affects diagnostic interpretation.
[0006] The fusion effect does not meet the diagnostic requirements: Pathological diagnosis requires clear differentiation of cell morphology (such as cell nucleus size and chromatin distribution) and tissue boundaries; existing methods either form harsh splicing boundaries (clear division of cell structures at different focal planes) or excessively smooth key details (such as blurring of cell nuclear heterogeneity features) during the fusion process, which cannot simultaneously meet the dual requirements of diagnosis for boundary sharpness and detail preservation.
[0007] The aforementioned problems collectively restrict the quality and diagnostic efficiency of digital pathology images, and there is an urgent need for a focus stacking technology solution that can improve assessment robustness, optimize alignment efficiency, and meet diagnostic characteristics. Summary of the Invention
[0008] The purpose of this invention is to provide an adaptive image focus stacking method, system, and pathological slide scanner, which has the advantages of being able to adaptively process high dynamic range images, improve the robustness of sharpness assessment, optimize alignment efficiency, and improve fusion effect, thereby generating high-quality, clear, full-area focus stacked images.
[0009] This invention provides an adaptive image focus stacking method, comprising the following steps: Acquire multiple frames of original images from different focal planes within the same field of view; Perceptual preprocessing is performed on each frame of the original image. The perceptual preprocessing includes: automatically identifying the image bit depth and performing brightness normalization and sensor noise suppression. Multi-scale spatial frequency evaluation is performed on each preprocessed frame image, a Gaussian Laplacian pyramid is constructed, the Laplacian response of each layer is calculated, and then upsampled to the original size and weighted and accumulated to obtain the multi-scale weight map of the frame image. Adaptive raster registration based on multi-scale weighted graphs includes: dividing each frame image into N×N grids, calculating the energy integral of each grid, selecting the grid with the highest energy integral as the registration anchor point, performing sub-pixel level normalized cross-correlation matching on the anchor point region to obtain the global offset vector, and performing translation correction on all frame images based on the global offset vector. Intelligent decision fusion is performed on each frame of the corrected image, including: calculating the local weight variance pixel by pixel based on the multi-scale weight map; when the local weight variance is less than the first threshold, the weighted fusion mode is used for output; when the local weight variance is greater than the second threshold, the hard-cut selection mode is used for output. The first threshold and the second threshold are adaptively determined based on the statistical characteristics of the weight variance of the whole image. Spatial consistency verification is performed on the fused image, including: generating a pixel source index map based on the multi-scale weight map of each frame image, detecting boundary pixels with inconsistent neighborhood sources in the index map, performing adaptive smoothing processing only on boundary pixels, and outputting the final full-area sharp focus stacked image.
[0010] Furthermore, this invention also proposes that the steps for constructing a Gaussian-Laplace pyramid and calculating a weight map in multi-scale spatial frequency evaluation specifically include: Preprocessed image Gaussian smoothing and downsampling are performed sequentially to construct its... The first layer of the pyramid; Layer Image ,in =0,1,… ; For images Execute the Laplace operator Calculate and then apply Gaussian smoothing. A single-layer response is obtained. ;in, ; After upsampling the responses of each layer to the original image size, a weighted accumulation strategy is used to generate the final weight map. ;in, Weight coefficients of each layer =1; The number of pyramid layers Protected by minimum resolution threshold Limited, when Stop building higher levels at that time; among them, Representing the The width of the layer image, Representing the The height of the layer image.
[0011] Furthermore, this invention proposes a dynamic size adjustment strategy for grid division in adaptive grid registration: setting a minimum grid size threshold. =128 pixels, number of grids The adaptive determination is as follows: ;in , These are the width and height of the image, respectively; The energy integral of the grid is defined as: ; Select the grid with the highest energy integral. Before using them as registration anchor points, perform blank area filtering: If the energy density of the grid If the region is blank or has low feature, the current anchor point is skipped and the next highest energy raster is selected. Furthermore, when using the Laplace method to calculate the weights, the energy integral value needs to be multiplied by an amplification factor of 10000.
[0012] Furthermore, the present invention proposes that the sub-pixel level normalized cross-correlation matching specifically involves: performing normalized cross-correlation calculations within the selected anchor point grid region, locating the correlation peak to sub-pixel accuracy through quadratic surface fitting, and obtaining the offset vector. ;when ≤ ,and ≤ When this is determined to be a minor offset, global translation correction is skipped; where =3 pixels.
[0013] Furthermore, this invention also proposes a multi-anchor-point candidate mechanism: when the optimal grid is located in the image edge region, the energy integrals of all grids are globally sorted in descending order, and the top several grids are selected as candidate anchor points, and the normalized cross-correlation scores are calculated for each. If the current candidate anchor point If the registration fails, the image is considered to have failed and the system automatically switches to the next candidate anchor point. If all candidate anchor points fail to register, the original image is used instead of the original image, and no correction is performed.
[0014] Furthermore, this invention also proposes that, in intelligent decision fusion, local weight variance... Defined as: Where N is the number of image frames, For the current pixel in the th... Weights on each image This represents the average weight of the current pixel across all frames of the image. The first threshold With the second threshold Based on the mean of the variance of the total graph weights with standard deviation Adaptive computation: .
[0015] Furthermore, the present invention also proposes to set pixels... The final fusion value is ; The weighted fusion mode uses normalized weight accumulation: ; The hard-cut selection mode directly takes the pixel corresponding to the maximum weight as the output: .
[0016] Furthermore, this invention also proposes a pixel source index map in spatial consistency verification. Generates a multi-scale weighted map by comparing each frame of the image pixel-by-pixel: The variance of the pixel source index map in the neighborhood is detected, and the current pixel is determined to be a boundary pixel based on the variance value. The adaptive smoothing process is only applied to pixels determined to be boundaries. Gaussian convolution kernels are applied to boundary pixels for smoothing, while non-boundary pixels retain their original pixel values. The size and standard deviation of the Gaussian convolution kernel are adaptively adjusted according to the maximum weight difference on both sides of the boundary.
[0017] In addition, the present invention also proposes an adaptive image focus stacking system for implementing the above method, which includes: The perception preprocessing module is used to automatically identify the bit depth of the input image and perform brightness normalization and sensor noise suppression. The multi-scale spatial frequency evaluation module is used to construct the Gaussian Laplacian pyramid, calculate the Laplacian response of each layer and accumulate it with weights, and output a multi-scale weight map of each frame image. The adaptive raster registration module is used to divide the image into a raster, select registration anchor points based on energy integral, perform subpixel-level normalized cross-correlation matching, and perform translation correction on all frame images based on the obtained global offset vector. The intelligent decision fusion module is used to calculate the local weight variance pixel by pixel based on the multi-scale weight map, adaptively switch between weighted fusion mode and hard-cut selection mode, and output a preliminary fused image. The spatial consistency verification module is used to generate a pixel source index map, detect and locate boundary pixels, perform adaptive smoothing only on boundary pixels, and output the final full-area sharp focus stacked image.
[0018] Furthermore, the present invention also proposes a pathological section scanner, comprising: Stage for holding pathological slides; An optical imaging system, including at least one objective lens and an image sensor, is used to acquire raw images of pathological sections at different focal planes; And as mentioned above, the adaptive image focus stacking system is connected to an image sensor to perform focus stacking processing on the acquired raw images and output a clear digital pathological image across the entire area.
[0019] As can be seen from the above, the adaptive image focus stacking method, system, and pathological slide scanner provided by this invention first automatically identifies the image bit depth (8 / 16 bits) through a perception preprocessing module, retains the complete 16-bit dynamic range, and performs brightness normalization and noise suppression; then, a Gaussian spectrum is constructed through a multi-scale spatial frequency evaluation module. The Laplacian pyramid utilizes multi-level response weighted accumulation to generate a weighted map that can simultaneously characterize cellular-level texture details and tissue-level macroscopic contours. Based on this, the adaptive raster registration module uses the region with the richest raster energy integral localization features as the registration anchor point, achieving efficient and high-precision alignment through sub-pixel normalized cross-correlation matching. The intelligent decision fusion module adaptively switches between weighted fusion and hard-cut selection modes based on local weight variance, ensuring smooth transitions in homogeneous regions while preserving cell edge sharpness. Finally, the spatial consistency verification module detects seams using a pixel source index map and performs adaptive smoothing only on boundary pixels to eliminate visual artifacts.
[0020] Compared with the prior art, the present invention has the following significant advantages: (1) Native support for 16-bit high dynamic range images: The perception preprocessing module automatically identifies the bit depth and retains the complete dynamic range without down-bit conversion, avoiding the loss of staining details and cell microstructure information, and directly adapting to the output format of pathological slide scanners.
[0021] (2) Robust multi-scale clarity assessment: The Gaussian-Laplace pyramid multi-scale response accumulation strategy is adopted to capture high-frequency textures of cell nuclei at the bottom layer and low-frequency features of tissue contours at the top layer, effectively distinguishing staining boundaries from true details and making weight allocation more accurate.
[0022] (3) High efficiency and high precision registration: The anchor point positioning method based on grid energy integration reduces the computational load of full-map registration by more than 90%, while ensuring micron-level alignment accuracy through sub-pixel cross-correlation; multiple anti-misselection mechanisms and multiple anchor point alternative mechanisms ensure registration reliability.
[0023] (4) Diagnostic-level fusion effect: The intelligent decision fusion module dynamically selects the fusion mode based on the local weight variance. In homogeneous areas, weighted fusion is used to achieve a natural transition, while in high-contrast edge areas, hard cutting is used to retain sharpness, thus meeting the dual clarity requirements of pathological diagnosis for cell morphology and tissue boundaries.
[0024] (5) Seamless visual transition: The spatial consistency verification module only performs adaptive smoothing at the seams, and the details of non-boundary areas are fully preserved, which eliminates layer switching artifacts and avoids texture blurring caused by global smoothing.
[0025] (6) No hardware modifications are involved: It is implemented entirely through algorithm optimization and can be directly integrated into existing digital pathology slide scanner software systems, possessing extremely high industrial application value. Attached Figure Description
[0026] Figure 1 This is a schematic diagram of the overall execution flow of the system of the present invention; Figure 2 This is a schematic diagram of the multi-scale weight calculation process of the present invention; Figure 3 This is a schematic diagram of the adaptive raster registration subprocess of the present invention; Figure 4 This is a schematic diagram of the integrated decision-making process of the present invention; Figure 5 This is a schematic diagram of the boundary smoothing process of the present invention. Detailed Implementation
[0027] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of this invention. The components of this invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.
[0028] Traditional focus stacking techniques for pathological slide scanning suffer from several drawbacks, including poor image format compatibility leading to loss of key information, insufficient robustness in sharpness assessment resulting in unreasonable weight allocation, high computational cost of alignment schemes affecting scanning efficiency or failing to compensate for minor shifts causing ghosting, and fusion effects that do not meet diagnostic requirements, resulting in blurred details or abrupt stitching. These issues directly impact the accuracy of pathological diagnosis.
[0029] To address this, the present invention proposes an adaptive image focus stacking method, comprising the following steps: Acquire multiple frames of original images from different focal planes within the same field of view; Perceptual preprocessing is performed on each frame of the original image. This perceptual preprocessing includes: automatically identifying the image bit depth and performing brightness normalization and sensor noise suppression. Multi-scale spatial frequency evaluation is performed on each preprocessed frame image, a Gaussian Laplacian pyramid is constructed, the Laplacian response of each layer is calculated, and then upsampled to the original size and weighted and accumulated to obtain the multi-scale weight map of the frame image. Adaptive raster registration based on the multi-scale weight map includes: dividing each frame image into N×N grids, calculating the energy integral of each grid, selecting the grid with the highest energy integral as the registration anchor point, performing sub-pixel level normalized cross-correlation matching on the anchor point region to obtain the global offset vector, and performing translation correction on all frame images based on the global offset vector. Intelligent decision fusion is performed on each frame of the corrected image, including: calculating the local weight variance pixel by pixel based on the multi-scale weight map; when the local weight variance is less than the first threshold, the weighted fusion mode is used for output; when the local weight variance is greater than the second threshold, the hard-cut selection mode is used for output. The first threshold and the second threshold are adaptively determined based on the statistical characteristics of the weight variance of the whole image. Spatial consistency verification is performed on the fused image, including: generating a pixel source index map based on the multi-scale weight map of each frame image, detecting boundary pixels with inconsistent neighborhood sources in the index map, performing adaptive smoothing processing only on the boundary pixels, and outputting the final full-area sharp focus stacked image.
[0030] For ease of understanding, the following explains some key terms in this embodiment: Multiple raw images refer to a series of images with different sharp focuses acquired by an image sensor within the same field of view by adjusting the position of the focal plane. These images together constitute the input data used for focus stacking processing.
[0031] Perceptual preprocessing refers to the preliminary adjustment and optimization of the original image in the image processing workflow, simulating the characteristics of human visual perception, to improve the accuracy and robustness of subsequent processing. Its purpose is to eliminate or reduce non-ideal factors introduced during image acquisition, such as uneven brightness and noise.
[0032] Multi-scale spatial frequency assessment refers to quantifying the sharpness or detail richness of an image by analyzing its frequency information at different spatial scales. High-frequency information is typically related to image edges and details, while low-frequency information is related to the overall structure of the image. Multi-scale assessment can more comprehensively capture the sharpness features of an image.
[0033] The Laplacian pyramid of Gauss is a multi-scale representation of an image. While a Gaussian pyramid constructs a low-resolution version of an image through layer-by-layer Gaussian smoothing and downsampling, the Laplacian pyramid stores the detailed differences between each layer of the Gaussian pyramid image, which can be used to evaluate the local sharpness of the image.
[0034] A multi-scale weighted map is a method that assigns a weight value to each pixel or region in an image using multi-scale spatial frequency evaluation. This weight value reflects the sharpness or importance of the pixel or region at different scales. The weighted map is the foundation for subsequent image registration and fusion.
[0035] Adaptive raster registration refers to dividing an image into multiple local regions (rasteres) and performing registration operations on these local regions. This method improves the accuracy and efficiency of registration by adaptively selecting registration anchor points and performing local matching to cope with non-uniform deformations or small offsets that may exist in the image.
[0036] Subpixel-level normalized cross-correlation matching is a high-precision image matching technique. It calculates the normalized cross-correlation coefficient between two image regions and further uses interpolation or fitting methods to pinpoint the matching peak to an accuracy of less than one pixel, thereby achieving subpixel-level image alignment.
[0037] Intelligent decision-making fusion refers to the dynamic selection of different fusion strategies based on the local characteristics of an image (such as sharpness and texture) during the image fusion process. This decision-making mechanism enables the fusion process to adaptively adjust to changes in image content, thereby achieving better fusion results.
[0038] Spatial consistency verification refers to checking the fusion result after image fusion is completed to ensure that the image is spatially continuous and natural. This verification process aims to discover and correct inconsistencies or artifacts that may be introduced during the fusion process, such as abrupt transitions between different focal plane regions.
[0039] First Implementation Method like Figure 1 As shown, this embodiment provides an adaptive image focus stacking method.
[0040] First, acquire multiple frames of original images from different focal planes within the same field of view.
[0041] This can be achieved in several ways. For example, by adjusting the Z-axis position of the optical imaging system, after each adjustment, the image sensor acquires multiple images of different Z-axis focal planes within the same field of view, and then fuses them to generate a clear digital slice covering the entire area. Alternatively, by using an imaging device with variable depth of focus, image data at different depths of focus can be acquired in a single scan. These images are typically stored in digital formats such as JPEG, TIFF, or BMP.
[0042] Furthermore, perceptual preprocessing is performed on each frame of the original image. This preprocessing aims to optimize image quality, making it more suitable for subsequent sharpness assessment and fusion. Specifically, perceptual preprocessing may include processing the image bit depth, for example, converting all images to 8-bit or 16-bit format. Additionally, brightness normalization can be performed, for example, by adjusting the overall brightness distribution of the image through linear stretching or gamma correction to eliminate potential brightness differences between different frames. Simultaneously, sensor noise suppression is also a crucial step in the preprocessing; for example, traditional noise reduction algorithms such as mean filtering, median filtering, or Gaussian filtering can be used to reduce random noise in the image.
[0043] Subsequently, a multi-scale spatial frequency evaluation is performed on each preprocessed frame to generate a multi-scale weighted map reflecting its local sharpness. One implementation is to construct an image pyramid, for example, by generating image levels of different resolutions through successive downsampling operations. At each level, spatial frequency analysis operators, such as the Laplacian operator, Sobel operator, or gradient operator, can be applied to calculate the local sharpness response of the image. The responses of each level are then upsampled to the original image size and weighted and accumulated to obtain the multi-scale weighted map of that frame. These weight values reflect the sharpness of different regions of the image.
[0044] Based on the generated multi-scale weight map, adaptive raster registration is performed on each frame of the image. This step first divides each frame of the image into multiple raster cells, for example, using fixed-size rectangular raster cells. Next, the energy integral of each raster is calculated, which can be simply defined as the sum of the weight values of all pixels within the raster. Then, the raster with the highest energy integral is selected as the registration anchor point. Within the selected anchor point region, sub-pixel-level normalized cross-correlation matching is performed to accurately calculate the offset of that region relative to the reference frame, thereby obtaining the global offset vector. Finally, translation correction is performed on all frame images based on this global offset vector to eliminate any minor displacements that may have been introduced during image acquisition.
[0045] Furthermore, intelligent decision fusion is performed on the corrected frames. This fusion process calculates the local weight variance pixel by pixel based on a multi-scale weight map. The local weight variance can be easily obtained by calculating the variance of the weight values of the current pixel across all frames. The fusion mode is dynamically selected by comparing this local weight variance with a preset first threshold and a second threshold. For example, when the local weight variance is small, it indicates that the sharpness difference in that region is not significant between different frames. In this case, a weighted fusion mode can be used, averaging the pixel values of each frame according to their weights. When the local weight variance is large, it indicates that there is a significant sharpness difference in that region. In this case, a hard-cut selection mode can be used, directly selecting the pixel corresponding to the frame image with the largest weight value as the output. The first and second thresholds can be determined based on empirical values or through statistical analysis of the weight variance across the entire image.
[0046] Finally, spatial consistency is checked on the fused image to ensure the naturalness and continuity of the final output image. This check involves generating a pixel source index map based on the multi-scale weight map of each frame, which records which original frame each pixel in the fused image originates from. Then, boundary pixels with inconsistent neighborhood sources in the index map are detected, for example, by comparing whether a pixel has the same source index as its surrounding neighboring pixels. Finally, adaptive smoothing is performed only on these detected boundary pixels; for example, local mean filtering or Gaussian filtering can be applied to smooth these boundaries, while non-boundary pixels retain their original fused values, thus outputting a final, fully sharp, focused stacked image.
[0047] The method provided in this embodiment effectively adapts to various image formats output by pathological slide scanners and suppresses noise through perceptual preprocessing, ensuring the integrity of the original image information. Multi-scale spatial frequency evaluation combined with Gaussian-Laplace pyramids can robustly evaluate the complex sharpness features of pathological images. Adaptive raster registration combined with sub-pixel level matching achieves high-precision and efficient image alignment, effectively avoiding ghosting. Intelligent decision fusion dynamically selects a fusion strategy based on local characteristics, avoiding scrambled stitching and blurred details. The final spatial consistency check further ensures the natural continuity of the fused image, thereby outputting a clear digital pathological image that meets the requirements of pathological diagnosis.
[0048] In some embodiments of the present invention described above, a perceptual preprocessing of the original image is proposed to automatically identify the image bit depth and perform brightness normalization and sensor noise suppression. However, if the image bit depth identification and brightness normalization processing are not precise or adaptive enough, it may lead to the loss of high dynamic range information or poor brightness adjustment effect in different image scenes, thereby affecting the accuracy of subsequent focus stacking and image quality.
[0049] In response, this invention further proposes that the automatic identification of image bit depth in the aforementioned perception preprocessing specifically involves: detecting the number of bits per pixel in the original image; when the bit depth is 16 bits, preserving its complete dynamic range and not performing bit-down conversion; and the brightness normalization is performed based on global histogram statistics.
[0050] Specifically, image bit depth refers to the number of bits used to represent the color or grayscale information of each pixel in an image. A higher bit depth means more hues or grayscale levels can be represented, thus providing a wider dynamic range and finer image detail. The number of bits per pixel in the raw image is detected, for example, by reading the image file's metadata (such as the header information of TIFF or DICOM formats) or analyzing the distribution range of pixel values. When a bit depth of 16 bits is detected, the system ensures that the full dynamic range of the 16-bit image is preserved throughout the entire processing flow, without performing any form of downsizing (e.g., from 16 bits to 8 bits). This means that image data will be stored and processed in 16-bit integer or corresponding floating-point format to avoid information loss due to reduced bit depth, which is particularly critical for scientific or medical images requiring high precision and a wide dynamic range.
[0051] Meanwhile, brightness normalization aims to adjust the overall brightness and contrast of the image to maintain consistency under different acquisition conditions, thereby improving the robustness of subsequent image processing. Brightness normalization based on global histogram statistics means that the system calculates a histogram of pixel intensity distribution across the entire image. By analyzing this global histogram, statistical information such as the overall brightness, contrast, and pixel value range of the image can be obtained. For example, histogram equalization, histogram matching, or linear stretching based on global minimum / maximum values can be used. This global statistical method ensures that brightness adjustment is based on the overall image content, avoiding the adverse effects of outliers in local areas on the overall brightness balance. This effectively corrects brightness differences between different frames, providing consistent input for subsequent fusion processing.
[0052] Through the above technical solution, this invention can accurately identify the bit depth of the original image and adopt a strategy to preserve the complete dynamic range of high bit depth images (such as 16-bit images), effectively avoiding information loss that may occur in the preprocessing stage, thus ensuring that subtle brightness changes and rich detail information in the image are fully preserved. Simultaneously, through brightness normalization based on global histogram statistics, unified and robust brightness adjustment can be performed on multiple frames of original images at different focal planes, effectively eliminating brightness deviations caused by uneven illumination or sensor differences, resulting in high consistency in brightness across all image frames. This not only provides high-quality, highly consistent input for subsequent multi-scale spatial frequency evaluation, significantly improving the accuracy of weight map calculation, but also lays a solid foundation for adaptive raster registration and intelligent decision fusion, ultimately contributing to the generation of full-area sharp focus stacked images with richer details, more realistic colors, and better overall visual effects.
[0053] like Figure 2 As shown, this invention further proposes that the steps for constructing a Gaussian-Laplace pyramid and calculating a weight map in multi-scale spatial frequency assessment specifically include: Preprocessed image Gaussian smoothing and downsampling are performed sequentially to construct its... The first layer of the pyramid; Layer Image ,in =0,1,… ; For images Execute the Laplace operator Calculate and then apply Gaussian smoothing. A single-layer response is obtained. ;in, ; After upsampling the responses of each layer to the original image size, a weighted accumulation strategy is used to generate the final weight map. ;in, Weight coefficients of each layer =1. This multi-scale weighting evaluation mathematical model ensures that both cellular-level texture details (such as nuclear heterogeneity) and tissue-level macroscopic contours are accurately weighted during fusion.
[0054] This invention chooses the Gaussian-Laplace pyramid as a multi-scale analysis tool, rather than wavelet transform or other methods, based on the following technical considerations: (1) Edge response characteristics match pathological image requirements: The core diagnostic features of pathological slide images include cell nucleus edges (high frequency) and staining region transitions (medium frequency). The Laplacian operator, as a second-order differential operator, exhibits zero-crossing characteristics in its edge response, enabling precise localization of cell nucleus contours. ; When combined with Gaussian smoothing, the Gaussian-Laplacian operator can enhance edge response while suppressing noise: ; (2) Computational efficiency advantage: Compared with wavelet transform, which requires multiple convolution operations, Gaussian pyramid adopts a downsampling strategy, halving the image size of each layer, resulting in an exponential decrease in computational complexity. .
[0055] (3) Clear scale separation: Bottom layer (k=0): Captures details of cell nuclear texture (high frequency, wavelength <5μm); Middle layer (k=1,2): Capture tissue staining boundaries (intermediate frequency, wavelength 5-20μm); High-level (k≥3): Capture tissue contour features (low frequency, wavelength>20μm).
[0056] This invention employs an equal-weight accumulation strategy, that is, the weight coefficients of each layer... Because pyramid downsampling causes the resolution of each layer to decrease geometrically, higher-level features naturally exhibit weaker response strength after being upsampled back to their original size. Therefore, equal-weighted accumulation is equivalent to implicitly assigning higher weights to lower layers.
[0057] The number of pyramid layers Protected by minimum resolution threshold In this embodiment, the minimum resolution protection threshold is specified. =10000 pixels; when Stop building higher levels at that time; among them, Representing the The width of the layer image, Representing the The height of the layer image.
[0058] The kernel size of the Gaussian-Laplacian operator The default setting is 33 pixels. This size has a pixel resolution of approximately 0.25 μm under a 40x objective lens, which corresponds to covering a typical cell nucleus diameter of 8-15 μm.
[0059] Specifically, Gaussian smoothing and downsampling are performed sequentially on the preprocessed image to construct its layer pyramid. Gaussian smoothing aims to effectively suppress high-frequency noise and smooth image details by applying a weighted average, providing a more stable input for subsequent downsampling. Downsampling, on the other hand, reduces the image resolution to generate image layers at different scales, thus forming a multi-scale image pyramid structure. For example, 2x2 average pooling or max pooling can be used for downsampling, or pixels can be skipped entirely during sampling. This layer-by-layer processing approach allows the image to be analyzed at different scales, thereby capturing focal features of varying sizes.
[0060] Building upon the pyramid structure, this invention further proposes performing Laplacian operator operations on the image and then superimposing Gaussian smoothing to obtain a single-layer response. The Laplacian operator is a second-order differential operator, highly sensitive to edge and texture variations in an image, effectively highlighting the focal region. By applying the Laplacian operator to the image, regions with drastic brightness changes can be detected; these regions typically correspond to sharp focal points. Superimposing Gaussian smoothing further reduces the impact of noise on the response value during Laplacian response calculation, making focus detection more stable and accurate. For example, the image can be Gaussian smoothed first, and then its Laplacian response calculated, or a Gaussian kernel can be convolved with a Laplacian kernel to form a Gaussian Laplacian operator that is directly applied to the image.
[0061] To integrate focus information at different scales, this invention proposes upsampling the responses of each layer to the original image size and then using an equal-weight accumulation strategy to generate the final weight map. The upsampling operation ensures that response maps at all scales have the same resolution for pixel-by-pixel stacking. The equal-weight accumulation strategy means that the response of each layer in the pyramid contributes equally to the final weight map, which helps balance the importance of features at different scales and prevents any particular scale feature from excessively dominating the final focus determination. For example, upsampling can be performed using bilinear interpolation, bicubic interpolation, or nearest-neighbor interpolation. The final weight map comprehensively reflects the focus information of the image at different scales.
[0062] To ensure the effectiveness and computational efficiency of the pyramid, this invention further proposes limiting the number of pyramid layers to a minimum resolution protection threshold. Specifically, when the width or height of a certain layer of the pyramid is less than the preset minimum resolution protection threshold, the construction of higher-level pyramid layers will cease. For example, the minimum resolution protection threshold can be set to 16 pixels or 32 pixels. This constraint avoids generating image layers that are too small and lack sufficient information. These excessively small layers not only incur high computational costs but may also introduce noise or inaccurate focus information. In this way, the depth of the pyramid can be effectively controlled, ensuring that each layer contains meaningful image information, thereby optimizing the performance and efficiency of multi-scale spatial frequency evaluation.
[0063] Through the above technical solution, this invention provides a specific, systematic, and efficient multi-scale spatial frequency evaluation method. By sequentially performing Gaussian smoothing and downsampling to construct a multi-layer pyramid, it can capture focus information of the image at different scales from coarse to fine, effectively handling focus features of different sizes. Performing Laplacian operator operations on each layer of the image and superimposing Gaussian smoothing can accurately detect edge and texture details in the image. These details are key indicators for judging image sharpness. At the same time, Gaussian smoothing reduces noise interference and improves the robustness of focus detection. Upsampling the response of each layer to the original size and using an equal-weight accumulation strategy ensures the comprehensive integration of focus information at different scales, avoiding the bias of single-scale information, so that the final multi-scale weight map can more accurately and comprehensively reflect the local sharpness of the image. In addition, limiting the number of pyramid layers by using a minimum resolution protection threshold effectively avoids invalid calculations and noise introduction, improving the efficiency and stability of the algorithm. This refined pyramid construction and weight calculation method provides high-quality input for subsequent adaptive raster registration and intelligent decision fusion, significantly improving the accuracy of focus stacking and the sharpness and consistency of the final image.
[0064] like Figure 3 As shown, this invention further proposes an optimization for sub-pixel level normalized cross-correlation matching, specifically: performing normalized cross-correlation calculations within the selected anchor point grid region, and using quadratic surface fitting to locate the correlation peak to sub-pixel accuracy to obtain the offset vector. ;when ≤ ,and ≤ When this is determined to be a minor offset, global translation correction is skipped; where =3 pixels. This threshold can be flexibly adjusted according to the alignment accuracy requirements of the actual application scenario to balance registration accuracy and computational efficiency.
[0065] In the aforementioned adaptive raster registration process, subpixel-level normalized cross-correlation matching is a crucial step in achieving high-precision image alignment. Specifically, this matching process first performs a normalized cross-correlation operation within the selected anchor point raster regions. This means that the system extracts the anchor point raster regions selected by the adaptive raster registration module based on energy integration from the reference image (e.g., the first frame image) and the image to be registered. Subsequently, normalized cross-correlation is performed on these two regions to quantify their similarity. The normalized cross-correlation operation effectively eliminates the influence of image brightness differences, thereby more accurately reflecting the structural similarity of image content.
[0066] To further improve matching accuracy, this invention uses quadratic surface fitting to locate relevant peaks to sub-pixel precision, thereby obtaining an accurate offset vector. After performing normalized cross-correlation, a cross-correlation matrix is obtained, whose peak positions indicate the optimal matching point. To surpass integer pixel precision, this invention utilizes the cross-correlation coefficients of this peak point and its neighborhood to construct a continuous quadratic surface using mathematical methods (such as parabolic or Gaussian fitting), and calculates the precise peak position of this surface. The difference between this precise peak position and the center position of the template region is the sub-pixel-level offset vector between the images. This method significantly improves the fineness of registration, ensuring that even minute image misalignments are accurately captured.
[0067] Through the above technical solution, this invention can intelligently identify minute shifts between images and optimize subsequent image processing accordingly. When the absolute values of both the horizontal and vertical components of the global offset vector obtained by sub-pixel-level normalized cross-correlation matching are less than a preset threshold, the system will determine it as a minute shift and skip the global translation correction step. This avoids performing unnecessary translation operations on images with almost no shift, significantly reducing the consumption of computational resources and improving the overall image processing efficiency. At the same time, by avoiding correction of minute shifts, it also reduces the slight image distortion that may be introduced by interpolation and other operations, thereby optimizing the overall image processing flow while ensuring image alignment accuracy, making the focus stacking results more stable and reliable. This intelligent judgment mechanism ensures that complex translation operations are only performed when correction is truly needed, thus achieving a good balance between computational efficiency and image quality.
[0068] This invention further proposes a multi-anchor point candidate mechanism: when the optimal grid is located in the image edge region, the energy integrals of all grids are globally sorted in descending order, and the top several grids are selected as candidate anchor points, and the normalized cross-correlation scores are calculated for each. If the current candidate anchor point If registration fails at any candidate anchor point, the registration is deemed unsuccessful, and the system automatically switches to the next candidate anchor point. If registration fails at all candidate anchor points, the original image is used without correction. This mechanism aims to enhance the robustness and reliability of the image registration process.
[0069] Traditional registration methods typically rely on a single "optimal" anchor point for matching. However, when this single optimal anchor point fails to match due to various reasons (such as insufficient features or its location at the image edge), the entire registration process is interrupted or produces errors. Multi-anchor alternative mechanisms provide multiple alternative anchor points, ensuring that even if the preferred anchor point is unavailable, the system can try other potentially effective regions for registration, thus significantly improving the registration success rate. Specifically, when the optimal raster is located in an image edge region, the feature information in the image edge region is often incomplete or asymmetrical, making accurate cross-correlation matching in this region difficult. For example, a raster located at the image edge may only contain part of the target structure, resulting in an indistinct peak or spurious peaks in its cross-correlation function, thus affecting the accuracy of sub-pixel-level matching. Therefore, when the optimal raster is located at the edge, its reliability as a registration anchor point is reduced.
[0070] To address the aforementioned challenges, this invention performs a global descending sort of the energy integrals of all grid cells. After determining the energy integrals of each grid cell, the system arranges them from highest to lowest based on these integral values: The purpose of this step is to systematically identify regions with high feature density besides the highest-energy raster. These second-highest-energy rasteres, while not "optimal," may be more suitable as registration anchors in certain situations than the "optimal" raster in edge regions. Based on this, the system sequentially selects a number of rasteres as candidate anchors. "A number" typically refers to a preset small integer, such as 2, 3, or 5. The system selects these rasteres one by one as potential registration anchors, starting with the high-energy raster, according to the descending order of the results. This step-by-step approach ensures that the system fully utilizes all regions with high feature density in the image, rather than relying on just one.
[0071] For each selected candidate anchor point, the system calculates a normalized cross-correlation score. Specifically, for each selected candidate anchor point, the system performs sub-pixel-level normalized cross-correlation matching and calculates a cross-correlation score: This score measures the similarity between two image regions; a higher score indicates a better match. By calculating the score, the system can quantitatively evaluate the effectiveness of each candidate anchor point as a registration benchmark. If registration fails, the original image is used without correction.
[0072] Furthermore, this invention incorporates a fault-tolerance mechanism: if all candidate anchor points fail to register, the original image is used without correction. This is a crucial fault-tolerance mechanism. If the system tries all preset candidate anchor points, but their normalized cross-correlation scores fail to reach a preset threshold, it indicates that a reliable registration reference cannot be found in the current image frame. In this case, to avoid introducing erroneous translation correction, the system chooses not to perform any correction operation and directly uses the original image for subsequent processing. This ensures that, in extreme cases, the quality of the final focus stacked image will not degrade due to incorrect registration.
[0073] Through the above technical solution, this invention effectively solves the problem that a single optimal anchor point may lead to registration failure in image edges or areas with insufficient features. The multi-anchor-candidate mechanism significantly improves the robustness and success rate of image registration by systematically evaluating multiple high-energy grids as candidate anchor points. Even if the preferred anchor point is unreliable, the system can automatically switch and try other candidate anchor points, thus ensuring that an effective registration benchmark can still be found in complex or challenging scenarios. Furthermore, when all candidate anchor points fail to provide a reliable match, the strategy of backtracking to the original image avoids introducing erroneous corrections, ensuring the overall quality and spatial consistency of subsequent focus-stacked images. This enables this method to obtain more stable and accurate focus-stacked results when processing images with complex structures or many edge regions.
[0074] This invention further proposes a precise definition of local weight variance and an adaptive calculation method for the threshold in intelligent decision fusion. Specifically, local weight variance... Defined as: Where N is the number of image frames, For the current pixel in the th... Weights on each image This is the average weight of the current pixel across all frames of the image; this definition allows for the accurate quantification of the sharpness difference or focus information richness of each pixel in different focal plane images.
[0075] When the local weight variance is small, for example Using a weighted fusion mode typically means that the pixel is either relatively blurry or relatively sharp across multiple frames, with insignificant changes in focus information, such as in homogeneous regions; however, when the local weight variance is large, for example... If the hard-cut selection mode is used, it indicates that the sharpness of the pixel varies significantly in different frames, and there may be obvious focus transitions, such as in edge areas.
[0076] Meanwhile, the first and second thresholds used for mode switching are adaptively calculated based on the statistical characteristics of the variance of the full-graph weights. Specifically, the first threshold... With the second threshold Based on the mean of the variance of the total graph weights with standard deviation Adaptive computation: This means the system first calculates the weighted variance of all pixels in the entire image, and then dynamically determines the first and second thresholds based on the mean and standard deviation of these variances. For example, the first threshold can be set as the mean of the weighted variance of the entire image minus a certain multiple of the standard deviation, while the second threshold can be the mean of the weighted variance of the entire image plus another multiple of the standard deviation. Alternatively, other statistical methods can be used, such as determining the thresholds based on percentiles or quantiles, to ensure that the thresholds can adapt to different image content and acquisition conditions.
[0077] By precisely defining the local weight variance through the above technical solution, the focus information difference of each pixel in multiple frames of images can be accurately quantified, providing a reliable basis for subsequent fusion decisions. Simultaneously, by adaptively calculating the first and second thresholds based on the mean and standard deviation of the overall image weight variance, the switching threshold for the fusion mode is no longer a fixed value but can dynamically adapt to different image scenes and acquisition conditions. This avoids fusion errors caused by improper threshold settings, such as misusing the hard-cut mode in areas with gentle focus changes or misusing the weighted fusion mode in areas with rapid focus changes. This adaptive threshold determination mechanism ensures the robustness and accuracy of the intelligent decision-making fusion process, thereby improving the overall sharpness, detail preservation, and spatial consistency of the final focus-stacked image. Especially in complex or variable image content, it can more effectively handle focus transition areas and avoid artifacts and blurring.
[0078] like Figure 4 As shown, the present invention further proposes a final fusion value for pixels, the specific calculation method of which depends on the result of intelligent decision fusion.
[0079] When the system determines to use the weighted fusion mode, this mode employs a normalized weight accumulation method. Specifically, for any pixel in the image, its fused pixel value can be obtained by multiplying the original value of that pixel in each frame image with its corresponding weight value in the multi-scale weight map, and then summing the products of all frames. To ensure the brightness consistency of the fusion result, these weights are usually normalized, meaning the sum of the weights for each pixel is 1. This method can smoothly transition between different focal plane regions, reducing the abruptness of the fusion boundary, thus achieving a natural and continuous fusion effect.
[0080] When the system determines to use the hard-cut selection mode, this mode directly takes the pixel corresponding to the maximum weight as the output. Specifically, the system iterates through all frames of the image at the current pixel position, finds the frame with the highest weight, and then uses the original pixel value of that frame at that pixel position as the final fused output. The advantage of this mode is that it can preserve the sharpness of local areas to the maximum extent and avoid the slight blurring that may be caused by weighted averaging. It is especially suitable for scenes with clear boundaries and high contrast in the focus area, ensuring the sharpness of the focus area.
[0081] In this embodiment, let pixels be... The final fusion value is ; The weighted fusion mode uses normalized weight accumulation: ; The hard-cut selection mode directly takes the pixel corresponding to the maximum weight as the output: .
[0082] By clearly defining the specific calculation methods for the weighted fusion mode and the hard-cut selection mode, this invention ensures the accuracy and stability of the intelligent decision-making fusion process. When the local weight variance is small, indicating that the sharpness difference between different focal plane images in this region is not significant, the weighted fusion mode with normalized weight accumulation can achieve a smooth transition of pixel values, effectively avoiding abrupt fusion boundaries and preserving more image detail information. Conversely, when the local weight variance is large, meaning there is a clear focal area, the hard-cut selection mode is directly used, selecting the pixel with the largest weight as the output. This maximizes the preservation of the highest sharpness in the local area and avoids detail loss or blurring that may result from weighted averaging. The precise definition of these two fusion modes, combined with the intelligent switching mechanism, enables the final output of a fully sharp, stacked image to effectively improve the sharpness and clarity of local details while maintaining overall visual consistency. This significantly improves common problems in traditional focal stacking methods, such as blurring, artifacts, or unnatural transitions, thus providing a high-quality visual foundation for subsequent image analysis and diagnosis.
[0083] In the aforementioned adaptive image focus stacking method, the corrected frames are fused using an intelligent decision fusion module to obtain a clear image across the entire region. However, in the transition areas between the weighted fusion mode and the hard-cut selection mode, or when the contributions of different focal plane images change drastically, unnatural seams, artifacts, or local inconsistencies may appear in the fused image, affecting the visual quality of the final image and the accuracy of subsequent analysis. This local spatial inconsistency is due to the discrete decisions made by the fusion strategy at the pixel level.
[0084] To address this, the present invention further proposes to perform spatial consistency verification on the fused image. In response to the "seam" problem caused by layer switching, the present invention introduces a boundary verification mechanism.
[0085] Specifically, such as Figure 5 As shown, in spatial consistency verification, the pixel source index map Generates a multi-scale weighted map by comparing each frame of the image pixel by pixel. A pixel source index map is an auxiliary information map that records which frame of the original image each pixel in the fused image primarily originates from. During the generation process, for each pixel position in the fused image, the system iterates through all input frame images to find the multi-scale weight values corresponding to that position. By comparing these weight values, the index of the frame with the largest weight value (e.g., 1 for the first frame, 2 for the second frame, and so on) is assigned to the current pixel position, thus forming the pixel source index map. This index map intuitively reflects the "origin" of each pixel in the fused image, providing basic data for subsequent seam detection.
[0086] Based on this, detection Within a neighborhood, if different sources exist, the pixel is marked as a boundary pixel. Seam detection uses a 3×3 neighborhood window, calculating the variance of the index values within the window. If the variance is greater than zero, the current pixel is determined to be a boundary pixel. Seam detection aims to identify transition boundaries between regions with different sources in a fused image. Specifically, the system defines a 3×3 neighborhood window centered on each pixel in the fused image. Within this window, the pixel source index values of all pixels are calculated. If these index values differ, meaning the window contains pixels from different frames, the variance of these index values is calculated. When the calculated variance is greater than zero, it indicates the presence of at least two different pixel sources within the neighborhood, thus the center pixel is determined to be a boundary pixel. This detection method based on neighborhood index variance can effectively locate regions in a fused image that may have discontinuities, the so-called "seams."
[0087] Furthermore, the adaptive smoothing process only applies to boundary pixels. Gaussian convolution kernels are applied to the boundary pixels for smoothing, while the original pixel values of non-boundary pixels remain unchanged.
[0088] To avoid unnecessary blurring of the entire image, the smoothing process of this invention is highly selective. Only regions identified as boundary pixels in the seam detection step are smoothed. For non-boundary pixels, i.e., those located within a single source region, their original pixel values remain unchanged, thereby preserving image detail and sharpness to the maximum extent. This selective processing ensures the locality and specificity of the smoothing operation, avoiding over-processing of sharp regions. Simultaneously, the size and standard deviation of the Gaussian convolution kernel are adaptively adjusted based on the maximum weight difference on both sides of the boundary. The size and standard deviation of the Gaussian convolution kernel are key parameters affecting the smoothing effect. To achieve adaptive smoothing, the system dynamically adjusts these parameters based on the degree of image source difference on both sides of the boundary pixel. Specifically, for a detected boundary pixel, the system analyzes a multi-scale weight map of different source frame images within its neighborhood. The maximum weight difference on both sides of the boundary (e.g., the region contributed by frame A and the region contributed by frame B) is calculated. If the difference is large, it indicates that the transition at the boundary may be more abrupt, requiring a stronger smoothing effect; in this case, a Gaussian convolution kernel with a larger size and / or a larger standard deviation is selected. Conversely, if the differences are small, a smaller kernel size and / or a smaller standard deviation is used to achieve finer smoothing. This adaptive adjustment ensures that the smoothing intensity matches the degree of boundary inconsistency, thereby minimizing the loss of image detail while eliminating artifacts.
[0089] By employing the aforementioned spatial consistency verification technique, this invention effectively addresses the potential local spatial inconsistency issues that may arise after intelligent decision fusion. By generating a pixel source index map, the original contribution frame of each pixel in the fused image is clearly identified. A seam detection method based on neighborhood window variance accurately identifies transition boundary pixels between different source regions. Adaptive smoothing is applied only to these boundary pixels, while non-boundary pixels retain their original sharpness, thus eliminating fusion artifacts and unnatural seams while preserving image details to the maximum extent. The size and standard deviation of the Gaussian convolution kernel are adaptively adjusted based on the maximum weight difference on both sides of the boundary, ensuring that the smoothing intensity matches the actual degree of inconsistency. This refined and localized processing significantly improves the overall visual quality and spatial consistency of the focus-stacked image, resulting in a more natural and smooth final output of a clear, full-area focus-stacked image.
[0090] Second embodiment (system) Corresponding to the aforementioned methods, this invention proposes an adaptive image focus stacking system, which aims to implement the above methods efficiently and stably. The system includes: a perception preprocessing module, a multi-scale spatial frequency evaluation module, an adaptive raster registration module, an intelligent decision fusion module, and a spatial consistency verification module.
[0091] Specifically, the perception preprocessing module performs preliminary processing on the input image to ensure the quality and consistency of input data for subsequent algorithms. This module automatically identifies the bit depth of the input image, for example, by reading image file header information or analyzing pixel value ranges to determine the number of bits per pixel. When a bit depth of 16 bits is detected, the module preserves its complete dynamic range without performing any downsampling to avoid information loss. Furthermore, the module performs brightness normalization, for example, by adjusting the overall brightness distribution of the image based on global histogram statistics to achieve a standardized level, thereby eliminating the influence of brightness differences under different acquisition conditions. Simultaneously, the module also performs sensor noise suppression, for example, by applying algorithms such as Gaussian filtering, median filtering, or nonlocal mean filtering to effectively reduce random noise generated during image acquisition and improve the image signal-to-noise ratio.
[0092] The multi-scale spatial frequency evaluation module is used to quantify the sharpness or focus of different regions in an image. This module analyzes the multi-scale frequency information of the image by constructing a Gaussian-Laplacian pyramid. Specifically, it sequentially performs Gaussian smoothing and downsampling operations on the preprocessed image to generate a series of image layers with different resolutions, forming a pyramid structure. Then, it performs a Laplacian operator operation on each layer of the pyramid and superimposes Gaussian smoothing to obtain a single-layer response reflecting the high-frequency details (i.e., sharpness) of that layer. These single-layer responses are then upsampled to the original image size and weighted using a strategy such as equal-weighted accumulation to generate the final multi-scale weight map. Each pixel value in this weight map represents the focus sharpness at the corresponding location, providing a basis for subsequent image fusion. The number of pyramid layers can be adaptively adjusted according to the image size and the required granularity of detail analysis; for example, the construction depth of the pyramid can be limited by setting a minimum resolution protection threshold.
[0093] The adaptive raster registration module is used to accurately correct minute translational deviations between multiple frames of images.
[0094] This module first divides each frame of image into N×N grids, where the size of the grids can be dynamically adjusted according to the overall size and feature distribution of the image, for example, by setting a minimum grid size threshold. =128 pixels, number of grids The adaptive determination is as follows: ;in , These represent the width and height of the image, respectively.
[0095] Next, the module calculates the energy integral for each grid cell. The energy integral for each grid cell is defined as: ; Select the grid with the highest energy integral. Before using them as registration anchor points, perform blank area filtering: If the energy density of the grid If the region is blank or has low feature, the current anchor point is skipped and the next highest energy raster is selected. Furthermore, when using the Laplace method to calculate the weights, the energy integral value needs to be multiplied by an amplification factor of 10000.
[0096] Before selecting the grid with the highest energy integral as the registration anchor point, this module performs blank region filtering. If the energy density of a grid is too low, it is considered a blank or low-feature region and skipped, instead selecting the grid with the second highest energy as the anchor point. Once the anchor point is determined, the module performs sub-pixel level normalized cross-correlation matching within that anchor point region. For example, it uses quadratic surface fitting to locate correlation peaks to sub-pixel accuracy, thereby obtaining a precise offset vector. When the offset... Less than the tolerance threshold When the value is a pixel, it is considered a minor offset, and global translation correction is skipped.
[0097] Finally, translation corrections are performed on all frame images based on these global offset vectors to ensure precise spatial alignment. To improve robustness, the module can also include a multi-anchor-point candidate mechanism. When the optimal raster is located in the image edge region or registration fails, the energy integrals of all rasters are globally sorted in descending order, and the top few rasters are selected as candidate anchor points for trial.
[0098] The intelligent decision fusion module is used to fuse corrected multi-frame images into a single image with full clarity. This module calculates the local weight variance pixel-by-pixel based on a multi-scale weight map. Based on this local weight variance, and combined with preset first and second thresholds (these thresholds can be adaptively calculated based on the mean and standard deviation of the overall image weight variance, for example), the module can adaptively select the fusion mode. When the local weight variance is small (less than the first threshold), it indicates that the focus difference between frames in that region is not significant. In this case, a weighted fusion mode is used, for example, by accumulating normalized weights to smoothly fuse pixel values. When the local weight variance is large (greater than the second threshold), it indicates that there is a significant focus difference in that region. In this case, a hard-cut selection mode is used, directly selecting the pixel corresponding to the maximum weight as the output to ensure the clarity of the optimal focus area. Furthermore, this module may also include a zero-weight protection mechanism: when the sum of the multi-scale weights of all frames at the current pixel position... At that time, the output is forced to use the pixels corresponding to the last frame of the image, that is... ,in The average weight threshold per pixel is used to define whether an image is blank. Images below this threshold are considered blank or nearly blank, and no artifact removal is required. Users can select an optimal value for this threshold based on the actual sample characteristics to address situations where no effective weight is assigned to a particular pixel across all frames, ensuring the integrity of the output image.
[0099] The spatial consistency verification module is used to further optimize the fused image, eliminate potential fusion boundary artifacts, and ensure visual consistency of the image. This module first generates a pixel source index map by comparing the multi-scale weight maps of each frame pixel by pixel. Record which frame of the original image each fused pixel originated from; among them, Next, seam detection uses a 3×3 neighborhood window to detect boundary pixels with inconsistent neighborhood sources in the index image. This involves calculating the variance of the index values within the window; if the variance is greater than zero, the current pixel is considered a boundary pixel. Finally, this module performs adaptive smoothing only on these boundary pixels, for example, by applying a Gaussian convolution kernel. Non-boundary pixels retain their original pixel values, thus eliminating artifacts while preserving image details to the maximum extent. The size and standard deviation of the Gaussian convolution kernel can be adaptively adjusted based on the maximum weight difference between the two sides of the boundary to achieve the best smoothing effect.
[0100] The aforementioned system architecture decomposes the complex focus stacking method into clearly defined independent modules, effectively solving problems such as unclear interfaces, difficulties in process coordination, and low system integration that may arise during method implementation. The perceptual preprocessing module ensures the quality and consistency of the input image; the multi-scale spatial frequency evaluation module provides accurate focus information for subsequent fusion; the adaptive raster registration module guarantees accurate alignment of multiple frames; the intelligent decision-making fusion module adaptively selects the optimal fusion strategy based on local characteristics, avoiding the limitations of a single fusion mode; and the spatial consistency verification module further eliminates boundary artifacts that may occur during the fusion process, ensuring the smoothness and naturalness of the final image. This modular design not only improves the system's execution efficiency and stability but also greatly simplifies system maintenance, upgrades, and expansion, making the entire focus stacking process more robust and reliable, ultimately outputting a high-quality image with full-area clarity.
[0101] In some embodiments of the present invention described above, although an advanced adaptive image focus stacking method and its implementation system are provided, in practical applications, especially in high-throughput, high-resolution imaging fields such as digital pathology, simply possessing this system cannot completely solve the challenge of obtaining clear images of the entire field of view. For example, pathological slides often have uneven thickness, making it difficult to obtain clear images of the entire field of view under a single focal plane, which brings inconvenience and efficiency problems to subsequent diagnostic analysis.
[0102] Third embodiment (pathology slide scanner) The present invention also proposes a pathological slide scanner, which includes a stage, an optical imaging system, and an adaptive image focus stacking system as described above.
[0103] Specifically, this pathology slide scanner features a stage whose primary function is to hold and hold the pathology slides. The stage is typically designed as a high-precision moving platform, ensuring the stability of the slides throughout the imaging process and allowing the optical imaging system to accurately scan different areas of the slide. To accommodate various sizes and types of pathology slides, the stage can be configured with adjustable clamps or adapters to ensure stable placement and precise alignment of the samples.
[0104] The scanner also includes an optical imaging system consisting of at least one objective lens and an image sensor. The optical imaging system is a key component for digitizing the microstructure of pathological slides. The objective lens magnifies the microstructure on the pathological slide and focuses it onto the image sensor. The image sensor then converts the received light signals into electrical signals, thereby generating a digital image. To support focus stacking processing, the optical imaging system has the ability to accurately acquire raw images at different focal planes. This is typically achieved by controlling the objective lens or stage to move precisely along the Z-axis at the micrometer level, thereby acquiring a series of images with different depths of focus.
[0105] Furthermore, the aforementioned pathology slide scanner also integrates an adaptive image focus stacking system, as described above. This focus stacking system connects to the image sensor in the optical imaging system via a data interface, ensuring that multiple frames of raw images acquired by the image sensor can be transmitted to the focus stacking system in real-time or near real-time for subsequent processing. This tight integration enables the focus stacking system to perform efficient and accurate focus stacking processing on raw images acquired from different focal planes.
[0106] Through the above technical solution, this invention effectively solves the problem of localized image blurring caused by uneven tissue thickness or limited depth of field in traditional pathological slide scanning. The stable support of the stage ensures the precise position of the pathological slide during scanning, while the optical imaging system efficiently and accurately acquires multiple frames of raw images from different focal planes. Subsequently, the seamless connection between the focus stacking system and the image sensor allows for immediate and precise focus stacking processing of the acquired raw images. This integrated solution enables the scanner to automatically generate clear digital pathological images across the entire area, avoiding the tedious manual refocusing and significantly improving scanning efficiency and image quality. The final output digital pathological image has high resolution and full-focus clarity, providing pathologists with more reliable and comprehensive diagnostic information, thereby improving the accuracy and efficiency of pathological diagnosis.
[0107] This invention further proposes a zero-weight protection mechanism. This mechanism aims to handle the special case during focus stacking where all input image frames fail to provide effective focus information at a specific pixel location (i.e., multi-scale weights are extremely low or zero). Its core idea is that, in the absence of effective focus information, instead of relying on conventional fusion algorithms that may produce erroneous results, a pre-defined, robust backoff strategy is employed to ensure the integrity and reasonableness of the output image. This mechanism avoids fusion failure or artifacts caused by missing or invalid weight information.
[0108] Specifically, the zero-weight protection mechanism includes: when the sum of the multi-scale weights of all frames at the current pixel position is less than or equal to a preset minimum threshold. At this time, the output is forced to use the pixel corresponding to the last frame image. During focus stacking, each pixel in each frame image is assigned a multi-scale weight, which reflects the sharpness or focus of the pixel in the corresponding frame. To determine whether a pixel lacks effective focus information, the multi-scale weights at that pixel position in all input frames need to be accumulated. If this accumulated value is lower than a preset minimum threshold... If the summation is insufficient, the pixel is considered to be out of focus or featureless in all frames. This summation operation is the key criterion for determining whether the zero-weight protection mechanism is triggered. When the zero-weight protection mechanism is triggered, i.e., when the sum of the multi-scale weights of all frames at the current pixel position is extremely low, the system will no longer perform the regular intelligent decision fusion, but will directly select the pixel value of the last frame at that pixel position as the final output. Selecting the last frame as the backoff source is a common, simple, and usually effective strategy because it avoids complex calculations when effective information is lacking and provides a relatively stable benchmark. In many application scenarios, the last frame is usually a valid frame in the acquisition sequence, and its pixel value can serve as a reasonable default padding. This is an extremely small positive threshold used to determine whether the sum of the multi-scale weights at the current pixel position across all frames is "close to zero". When the sum of the weights is less than or equal to this threshold, the pixel is considered to lack effective focus information, thus triggering the zero-weight protection mechanism. This threshold needs to be set small enough to ensure that protection is only triggered when the weights are indeed very low, avoiding false positives. Its specific value can be fine-tuned based on the actual application scenario and the characteristics of the weight calculation method, but it is usually an extremely small value close to zero.
[0109] By introducing a zero-weight protection mechanism, this invention effectively solves the problem that conventional intelligent decision fusion may lead to a decrease in output image quality or the appearance of artifacts when the multi-scale weights at a specific pixel position in all frames are extremely low or close to zero. When the sum of the multi-scale weights at the current pixel position in all frames is detected to be less than a preset minimum threshold... In this case, the system forces the output to use the pixels corresponding to the last frame, avoiding inaccurate fusion calculations when effective focus information is lacking. This ensures that even in extremely out-of-focus areas, low-contrast areas, or areas severely affected by noise, the final focus-stacked image maintains reasonable pixel values and visual consistency, significantly improving the robustness of the entire focus-stacking method and the overall quality of the output image.
[0110] The following example will provide a more detailed explanation of the above technical solution: In a digital pathology slide scanning scenario, a pathology slide scanner performs a high-magnification scan on an HE-stained pathology slide with varying thickness. Due to the small depth of field of the objective lens, a single image cannot capture all cellular structures in sharp focus. To obtain a clear digital pathology image of the entire area, the scanner acquires multiple frames of raw images at different focal planes along the Z-axis within the same field of view. These raw images are 16-bit images with a large dynamic range, containing staining details and microstructural information.
[0111] First, the system performs perceptual preprocessing on each frame of the original image. This step automatically detects the number of bits per pixel, identifies the image bit depth as 16 bits, and preserves its full dynamic range without performing bit-down conversion. This avoids the problem of loss of staining gradient and fine cellular structure information commonly found in existing 8-bit processing schemes. Subsequently, the system performs brightness normalization based on global histogram statistics to ensure brightness consistency between different frames and suppress sensor noise, providing high-quality input for subsequent processing.
[0112] Next, the system performs multi-scale spatial frequency evaluation on each preprocessed frame to identify sharp regions in the image. Specifically, the system constructs a Gaussian-Laplacian pyramid. First, Gaussian smoothing and downsampling are performed sequentially on the image to construct its layer pyramid, where the number of layers is limited by a minimum resolution protection threshold; for example, construction of higher levels is stopped when the image width or height is less than 16 pixels. Then, the Laplacian operator operation is performed on each layer, and Gaussian smoothing is superimposed to obtain a single-layer response. Here, the kernel size of the Gaussian-Laplacian operator is set to 3×3 pixels by default. This size corresponds to a pixel resolution of approximately 0.25μm under a 40x objective lens, which can cover the typical cell nucleus diameter of 8-15μm, thus avoiding misjudging stained boundaries as noise and capturing details in thin tissue regions. After upsampling each layer response to the original image size, an equal-weight accumulation strategy is used to generate a multi-scale weight map of the frame image. When Laplacian edge detection is used as the basis for weight calculation, the system also performs an exponential enhancement transform on the multi-scale weight map: The enhanced weighted graph directly replaces the original weighted graph for subsequent intelligent decision fusion, further improving the robustness of clarity assessment.
[0113] After acquiring multi-scale weight maps, the system performs adaptive raster registration based on these weight maps to correct for minor plateau shifts that may occur during scanning. The system divides each frame of image into N×N grids, where the grid division employs a dynamic size adjustment strategy, setting a minimum grid size threshold of 128 pixels, and the number of grids is adaptively determined. For example, for a 2048x1536 pixel image, the number of grids... The adaptive determination is as follows: The system calculates the energy integral of each grid cell, where the energy integral is defined as the sum of the weights of all pixels within that grid cell. Before selecting the grid cell with the highest energy integral as the registration anchor point, the system performs blank area filtering: if the energy density of the grid cell is less than 0.01, it is determined to be blank or have few feature regions, and the current anchor point is skipped, with the next highest energy grid cell selected sequentially. This avoids registration in regions lacking features, improving registration accuracy. Subsequently, the system performs sub-pixel level normalized cross-correlation matching on the selected anchor point region, using quadratic surface fitting to locate the correlation peak to sub-pixel accuracy, obtaining a global offset vector. When the absolute values of both the X and Y components of the offset vector are less than 3 pixels, it is determined to be a small offset, and global translation correction is skipped, thereby improving processing efficiency. In addition, the system includes a multi-anchor point candidate mechanism: when the optimal grid cell is located in the image edge region, the system globally sorts the energy integrals of all grid cells in descending order, sequentially selecting the top few grid cells as candidate anchor points, and calculating the normalized cross-correlation score for each. If the score of the current candidate anchor point is less than 0.7, the registration is determined to have failed, and the system automatically switches to the next candidate anchor point. If registration fails for all candidate anchor points, the original image is used instead of correction. This registration strategy effectively solves the problems of high computational cost and insufficient accuracy leading to ghosting in existing whole-image alignment, ensuring the high precision requirements of pathological scans.
[0114] Next, the system performs intelligent decision fusion on the corrected frames. The system calculates the local weight variance pixel-by-pixel based on the multi-scale weight map, where the local weight variance is defined as the average of the sum of squares of the differences between the current pixel's weight and the mean weight across all frames. The system adaptively determines a first threshold and a second threshold based on the mean and standard deviation of the overall image weight variance. When the local weight variance is less than the first threshold, it indicates that the pixel's sharpness varies little across different focal planes. The system outputs a weighted fusion mode, accumulating the normalized weights to obtain the final pixel value, thus preserving details of smooth transitions. When the local weight variance is greater than the second threshold, it indicates that the pixel has sharpness differences across different focal planes. The system outputs a hard-cut selection mode, directly selecting the pixel corresponding to the maximum weight as the output, ensuring that the sharpest details are preserved. This adaptive fusion strategy avoids the problem of blurred details caused by harsh or overly smooth stitching edges in existing solutions, meeting the needs of pathological diagnosis for distinguishing cell morphology and tissue boundaries. Furthermore, the system includes a zero-weight protection mechanism: when the sum of the multi-scale weights at the current pixel position across all frames is less than... At this time, the output is forced to use the pixels corresponding to the last frame of the image, ensuring that there is image output even when the feature area is very small.
[0115] Finally, the system performs spatial consistency verification on the fused image. The system generates a pixel source index map by comparing the multi-scale weight maps of each frame pixel by pixel, recording which original image each pixel ultimately originates from. The system uses a 3×3 neighborhood window to detect seams and calculates the variance of the index values within the window. If the variance is greater than zero, the current pixel is determined to be a boundary pixel. This means that there are pixels from different focal planes in the neighborhood of this pixel, which may lead to unnatural transitions. The system only performs adaptive smoothing on pixels determined to be boundaries, applying a Gaussian convolution kernel to smooth boundary pixels while keeping the original pixel values unchanged for non-boundary pixels. The size and standard deviation of the Gaussian convolution kernel are adaptively adjusted based on the maximum weight difference on both sides of the boundary, thereby eliminating inconsistent seams while preserving the original details of non-boundary areas to the maximum extent. This avoids the problem of detail loss caused by over-smoothing in existing methods, ultimately outputting a clear, fully focused stacked image, providing diagnostic evidence for pathologists.
[0116] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An adaptive image focus stacking method, characterized in that, Includes the following steps: Acquire multiple frames of original images from different focal planes within the same field of view; Perceptual preprocessing is performed on each frame of the original image. The perceptual preprocessing includes: automatically identifying the image bit depth and performing brightness normalization and sensor noise suppression. Multi-scale spatial frequency evaluation is performed on each frame of the preprocessed image. A Gaussian-Laplacian pyramid is constructed. The Laplacian response of each layer is calculated, then upsampled to the original size and weighted and accumulated to obtain the multi-scale weight map of the frame image. Adaptive raster registration based on the multi-scale weight map includes: dividing each frame image into N×N rasteres, calculating the energy integral of each raster, selecting the raster with the highest energy integral as the registration anchor point, performing sub-pixel level normalized cross-correlation matching on the anchor point region to obtain a global offset vector, and performing translation correction on all frame images based on the global offset vector. Intelligent decision fusion is performed on each frame of the corrected image, including: calculating the local weight variance pixel by pixel according to the multi-scale weight map; when the local weight variance is less than the first threshold, the weighted fusion mode is used for output; when the local weight variance is greater than the second threshold, the hard cut selection mode is used for output. The first threshold and the second threshold are adaptively determined according to the statistical characteristics of the weight variance of the whole image. Spatial consistency verification of the fused image includes: generating a pixel source index map based on the multi-scale weight map of each frame image, detecting boundary pixels with inconsistent neighborhood sources in the index map, performing adaptive smoothing processing only on the boundary pixels, and outputting the final full-area sharp focus stacked image.
2. The adaptive image focus stacking method according to claim 1, characterized in that, In the multi-scale spatial frequency assessment, the steps of constructing the Gaussian-Laplace pyramid and calculating the weight map specifically include: Preprocessed image Gaussian smoothing and downsampling are performed sequentially to construct its... The first layer of the pyramid; Layer Image ,in =0,1,… ; For images Execute the Laplace operator Calculate and then apply Gaussian smoothing. A single-layer response is obtained. ;in, ; After upsampling the responses of each layer to the original image size, a weighted accumulation strategy is used to generate the final weight map. ;in, Weight coefficients of each layer =1; The number of pyramid layers Protected by minimum resolution threshold Limited, when Stop building higher levels at that time; among them, Representing the The width of the layer image, Representing the The height of the layer image.
3. The adaptive image focus stacking method according to claim 1, characterized in that, In the adaptive raster registration, the raster division adopts a dynamic size adjustment strategy: a minimum raster size threshold is set. Number of grid cells The adaptive determination is as follows: ;in , These are the width and height of the image, respectively; The energy integral of the grid is defined as: ; Select the grid with the highest energy integral. Before using them as registration anchor points, perform blank area filtering: If the energy density of the grid If the current anchor point is not found, it is determined to be a blank or low-feature region. The next highest energy raster is then selected.
4. The adaptive image focus stacking method according to claim 3, characterized in that, The subpixel-level normalized cross-correlation matching specifically involves: performing a normalized cross-correlation operation within the selected anchor point grid region, locating the correlation peak to subpixel precision using quadratic surface fitting, and obtaining the offset vector. ;when ≤ ,and ≤ If the offset is small, the global translation correction is skipped.
5. The adaptive image focus stacking method according to claim 3, characterized in that, It also includes a multi-anchor-point candidate mechanism: when the optimal raster is located in the image edge region, the energy integrals of all rasters are globally sorted in descending order, and the top few rasters are selected as candidate anchor points in turn, and the normalized cross-correlation scores are calculated for each. If the current candidate anchor point If the registration fails, the image is considered to have failed and the system automatically switches to the next candidate anchor point. If all candidate anchor points fail to register, the original image is used instead of the original image, and no correction is performed.
6. The adaptive image focus stacking method according to claim 1, characterized in that, In the intelligent decision fusion, the local weight variance Defined as: Where N is the number of image frames, For the current pixel in the th... Weights on each image This represents the average weight of the current pixel across all frames of the image. The first threshold With the second threshold Based on the mean of the variance of the total graph weights with standard deviation Adaptive computation: .
7. The adaptive image focus stacking method according to claim 6, characterized in that, Let pixels The final fusion value is The weighted fusion mode adopts normalized weight accumulation: ; The hard-cut selection mode directly takes the pixel corresponding to the maximum weight as the output: .
8. The adaptive image focus stacking method according to claim 1, characterized in that, In the spatial consistency check, the pixel source index map Generates a multi-scale weighted map by comparing each frame of the image pixel-by-pixel: The variance of the pixel source index map in the neighborhood is detected, and the current pixel is determined to be a boundary pixel based on the variance value. The adaptive smoothing process is only applied to the boundary pixels. A Gaussian convolution kernel is applied to the boundary pixels for smoothing, while the original pixel values of non-boundary pixels remain unchanged. The size and standard deviation of the Gaussian convolution kernel are adaptively adjusted according to the maximum weight difference on both sides of the boundary.
9. An adaptive image focus stacking system for implementing the method of any one of claims 1 to 8, characterized in that, include: The perception preprocessing module is used to automatically identify the bit depth of the input image and perform brightness normalization and sensor noise suppression. The multi-scale spatial frequency evaluation module is used to construct a Gaussian-Laplacian pyramid, calculate the Laplacian response of each layer and sum them up in weights, and output a multi-scale weight map of each frame image. The adaptive raster registration module is used to divide the image into a raster, select registration anchor points based on energy integral, perform subpixel-level normalized cross-correlation matching, and perform translation correction on all frame images based on the obtained global offset vector. The intelligent decision fusion module is used to calculate the local weight variance pixel by pixel based on the multi-scale weight map, adaptively switch between weighted fusion mode and hard-cut selection mode, and output a preliminary fused image. The spatial consistency verification module is used to generate a pixel source index map, detect and locate boundary pixels, perform adaptive smoothing only on boundary pixels, and output the final full-area sharp focus stacked image.
10. A pathological slide scanner, characterized in that, include: Stage for holding pathological slides; An optical imaging system, including at least one objective lens and an image sensor, is used to acquire raw images of the pathological slides at different focal planes; And the adaptive image focus stacking system as described in claim 9, the system being connected to the image sensor for performing focus stacking processing on the acquired raw images to output a clear digital pathological image across the entire area.