Deep learning-based automated quality control tool for identifying out-of-focus regions in microscopic imaging
An automatic quality control tool using a deep learning model segments and classifies microscopic images to assess focus quality, addressing focus challenges in microscopy techniques, improving image analysis accuracy and reliability.
Patent Information
- Application Number
- PCT/US2025/037784
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-29
- Filing Date
- 2025-07-15
- Publication Date
- 2026-02-05
Smart Images

Figure US2025037784_05022026_PF_FP_ABST
Abstract
Description
DEEP LEARNING-BASED AUTOMATED QUALITY CONTROL TOOL FOR IDENTIFYING OUT-OF-FOCUS REGIONS IN MICROSCOPIC IMAGINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the priority and benefit of U.S. Provisional Patent Application No. 63 / 676,835, filed on July 29, 2024. The entire disclosure of the aforementioned application is incorporated by reference herein in its entirety for all purposes.BACKGROUND
[0002] Microscopy techniques are widely used in scientific and medical research by enabling the visualization of microscopic structures. Among these techniques, brightfield imaging is a commonly used method for its ability to illuminate specimens from below, revealing intricate details through variations in light absorption and scattering. Its applications may span diverse fields including pathology, biology, and materials science, offering clear and focused images for accurate analysis. However, maintaining focus in brightfield microscopy may present challenges influenced by various factors. Mechanical misalignments within the microscope setup, such as those affecting the objective lens or stage, can result in improper focusing during image capture. Moreover, discrepancies in specimen thickness, surface irregularities, or differences in refractive indices between the sample and its surroundings can lead to defocusing issues. Environmental conditions, such as vibrations, temperature fluctuations, or changes in humidity, may further complicate matters by destabilizing the imaging system and exacerbating focus-related issues. Additionally, in multiplex imaging where multiple dyes or chromogens are utilized to label different cellular components or biomarkers, maintaining focus may become increasingly difficult. Each dye or chromogen may have unique optical properties, including differing refractive indices and absorption spectra, which can introduce complexities in achieving optimal focus across channels. Variations in staining intensity or penetration depth of different dyes within the sample can also contribute to focus inconsistencies, resulting in suboptimal image quality and potential misinterpretation of results.
[0003] Similar challenges may be encountered in other microscopy techniques, such as darkfield microscopy, fluorescence microscopy, and differential interference contrast (DIC) microscopy. Darkfield microscopy, which uses oblique lighting to selectively illuminate specimens against a dark background, can suffer from focus issues due to variations in specimenthickness or refractive index. Fluorescence microscopy, while offering high sensitivity and specificity, may require precise focus to capture the emission signals from fluorophores accurately. Differential interference contrast microscopy, which enhances contrast in transparent specimens, may rely on precise optical alignment to maintain focus and produce clear images. These techniques, like brightfield microscopy, may demand meticulous attention to detail and environmental control to mitigate focus-related challenges and accurate interpretation of microscopic structures across various research disciplines.
[0004] Overcoming the challenges of out-of-focus (OOF) images in microscopy techniques may demand an approach involving precise equipment calibration, meticulous specimen preparation, and stringent environmental control measures. In the realm of digital pathology software like HALO and Visipharm, where accurate image analysis is paramount, strategies often rely on manual procedures to identify and exclude artifact regions from downstream processing. However, while current quality control (QC) protocols involve manual identification of these artifacts, these may be susceptible to human error and subjectivity, potentially leading to inconsistencies or overlooking minor details in image analysis.SUMMARY
[0005] Certain embodiments of the present disclosure relate to an automatic quality control (QC) tool assessing focus quality on various regions of the microscopic images by leveraging a deep learning model. The disclosed tool may receive a microscopic image depicting a slide that includes a slice of a sample that may be placed in a scanner (e.g., microscopes, confocal scanners, brightfield or fluorescence scanner) designed to capture high-resolution images of the samples or other specimens. In various microscopy techniques, including brightfield and darkfield microscopy, this sample may be preserved by different fixatives and embedded in a medium for providing structural support during sectioning. The sample may be then cut into thin slices and placed onto glass slides. Additionally, staining protocols may be applied for improving contrast and visualization. The captured microscopic images e.g., histopathology images, immunohistochemical (IHC) images, fluorescence images may be segmented by spatially sliding (i.e., horizontally and / or vertically) a window over the entire microscopic image with a stride to extract a set of patches. The window size and stride value may be selected considering factors such as avoiding unnecessary or redundant information that may occur for overlapping windowsand acquiring a suitable classification resolution for each patch. In one example, segmentation results in non-overlapping and uniform sized patches.
[0006] The extracted set of patches may be fed into the machine-learning (ML) / deep learning (DL) model that is configured to generate a predicted probability indicating an extent to which a patch corresponds to one of a (predefined) label. The set of labels may include in-focus (InF), out-of-focus (OOF), and non-tissue area (NTA). The deep learning model may be trained on a labeled dataset comprising a set of microscopic images, alternatively, a set of patches extracted from these microscopic images, captured for a range of focus depths (e.g., -3 pm to 3 pm with non-uniform or uniform spacing of 1 pm). These patches (or microscopic images) with z-offsets between a determined z-offset (or focus depth) margin may be labeled as in-focus (InF) images while the rest of the patches (or microscopic) may be labeled as out-of-focus (OOF). These labeled patches (i.e., InF and OOF) including additional NTA patches may be fed into the ML / DL model for the prediction task.
[0007] The predicted patches corresponding to the microscopic image may be aggregated to make decisions at the image level. For aggregation, various statistical techniques may be used that, based on a metric, combine the independent predictions on patches corresponding to each label within a single image. For example, the statistical approach may involve computing the metric such as percentage of patches that fall into each class label (by counting total number of patches belonging to each class), area covered by patches (specifically useful for varying patch sizes), or weighted sum of number of patches belonging to each class where weights may be assigned by a confidence score (or certainty) of the classifier. Based on the computed metric, the statistical technique may assign a set of scores to each predicted label. If it is determined that the score assigned to OOF region is above a predefined threshold indicating regions with more focus problems, an alert may be triggered. In response to this determination, the slide may be assigned back to the scanner for the rescanning, alternatively, it may be validated by a user (or an expert whether to be rescanned or accepted). Similarly, if the OOF score is below the predefined threshold, the image may be considered as an in-focus image, which can be sent for further analysis or processing.
[0008] The disclosed QC tool may further include reconstructing the microscopic image by placing each patch back into its original position, creating a reconstructed image where different classes may be represented by the assigned scores. Using image processing techniques, thesecontiguous regions can be highlighted with different colors and labeled (e g., by predicted labels such as InF, OOF or NTA, and / or the scores such as percentages, areas etc.) and shown in a graphical user interface (GUI). This method may be particularly useful for visualizing the spatial distribution and extent of each class within the image.
[0009] In one instance, the focus depth margin within which a patch is assigned a label as infocus or out-of-focus, can be approached in various ways including image processing, statistical techniques, quantitative methods, qualitative methods, or a combination thereof. For example, performing qualitative assessment such as manual inspection by domain experts or quantitatively by analyzing boundaries of the specimen within the image. For example, determining averages of the cell counts at different focus depths for a subset of microscopic images by leveraging image processing (that segments the boundaries of identified cells) and counting techniques (e.g., automatic, or manual to count number of cells in each image). Based on the analysis of cell count data, an acceptable focus depth margin, representing the range of z-offsets within which the cell counting accuracy meets a predetermined criteria or quality thresholds, can be identified.
[0010] In one example of the disclosed QC tool, a densely connected convolutional neural network also known as DenseNet model fine-tuned for the three-class classification problem is leveraged as the deep learning model. This deep model may include one or more dense blocks that comprise of one or more dense layers such that each dense layer may receive feature maps from one or more preceding layers. Additionally, Pytorch is leveraged for implementing the deep learning model.
[0011] In some aspects, to prepare dataset for the ML / DL model (prior to or after segmentation), one or more preprocessing techniques may be applied to the microscopic images. For example, the microscopic images may undergo normalization such as z-score, min-max and / or other similar normalizations to adjust the intensity ranges across the images of varying pixel intensities. Additional preprocessing may involve sampling (e.g., upsampling or downsampling) and / or data augmentation that may include affine transformation e.g., scaling, translation, random rotations, random flips, and other transformations such as image cropping, reshaping, image resampling and resizing. Leveraging these preprocessing techniques may be consequent in consistency of the dataset along with enhanced robustness and generalization ability of the model.
[0012] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein.
[0013] In some embodiments, a computer-program product tangibly embodied in a non- transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods or processes disclosed herein.
[0014] In some embodiments, a system is provided that includes one or more means to perform part or all of one or more methods or processes disclosed herein.
[0015] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. The present disclosure is described in conjunction with the appended figures:
[0017] FIG. 1 is a block diagram illustrating an example workflow of an automatic quality control (QC) tool assessing focus quality of microscopic images.
[0018] FIG. 2 shows an example illustration of in-focus and out-of-focus images when a sample is scanned with a microscope.
[0019] FIG. 3 illustrates an example architecture of the automated QC tool that performs classification of the microscopic images.
[0020] FIG. 4 illustrates an example architecture of a densely connected convolutional neural network (DenseNet) performing classification of the microscopic images.
[0021] FIG. 5 illustrates an exemplary method of determining a z-offset (focus depth) margin for in-focus (InF) and out-of-focus (OOF) images.
[0022] FIG. 6 illustrates one or more plots for determining z-offset margin for the exemplary method of the FIG. 5.
[0023] FIG. 7 illustrates one or more graphs showing performance of the disclosed technique in accordance with an example implementation.
[0024] FIG. 8 illustrates an example showing a QC image generated by the disclosed QC tool for a first microscopic image.
[0025] FIG. 9 illustrates an example showing performance of the QC tool for a second microscopic image.
[0026] FIG. 10 illustrates an example showing performance of the QC tool for a third microscopic image.
[0027] FIG. 11 illustrates an example process flow for performing QC on microscopic images.DETAILED DESCRIPTION
[0028] Some embodiments of the present disclosure relate to an automatic quality control (QC) tool that may perform assessments on different regions of the microscopic images and detect quality of focus. The disclosed QC tool may perform focus assessments by leveraging a deep learning (DL) model configured to take a microscopic image that is segmented into patches and generate a predicted probability for mapping each patch to a predefined label e.g., in-focus (InF), non-tissue area (NTA) or out-of-focus (OOF). The deep learning model may be trained on a dataset including non-tissue area along with different focus depths or z-offsets (e.g., -2.5pm to 2.5 m with a uniform spacing of 0.5) images that are labeled by assigning InF to the patches with z-offset between a determined z-offset margin while the rest of the patches may be labeled as out-of-focus (OOF). Based on a metric, a statistical technique may assign scores to each predicted label by aggregating or reconstructing the predicted patches associated with each label within a single microscopic image. Upon determination of the aggregation scores, the OOF images may be assigned back to the scanning device e.g., microscope for the rescanning while the validated in-focus images may be sent for further analysis or processing. The disclosed toolcan be leveraged for assessing the focus quality in various microscopy applications, including brightfield microscopy, darkfield microscopy, differential interference, or phase contrast microscopy, where the accuracy and reliability of images may contribute to precise diagnosis and analysis.
[0029] In microscopic techniques, sample preparation techniques aimed at maintaining the structural integrity and enhancing visibility of the specimen. This may involve processes such as fixation using substances like formaldehyde or alcohol-based solutions to preserve biological samples, followed by embedding in a substance e.g., a historical wax or paraffin and / or resins to provide support during sectioning. The samples (e g., a sample of tumor) are then sliced into thin sections having thickness e.g., ranging from 4 to 10 micrometers, using e.g., a microtome or a vibratome. These sections are mounted onto glass slides, serving as a platform for examining and analyzing tissue structures. In brightfield microscopy, similar preparation methods are employed to achieve high-quality images. Tissue samples undergo fixation, embedding, and sectioning, followed by staining techniques using dyes like hematoxylin and eosin (H&E) to enhance contrast and visualization. These staining protocols can highlight specific cellular components, facilitating the detailed observation of tissue structures under the microscope.
[0030] Once prepared, the slides are subjected to scanning using specialized imaging devices such as microscopes or confocal scanners. These scanners may employ different mechanisms, such as white light illumination in brightfield scanners or laser light and pinholes in confocal scanners, to capture high-resolution images of the stained tissue samples. The captured images may be then transmitted to an automatic quality control (QC) tool running on a computer system embedded within the scanner or connected externally through a communication network. The scanned microscopic images may comprise of a grid-like arrangement, often forming a rectangular matrix, composed of individual picture elements known as pixels. Each pixel represents a digital value capturing specific characteristics of the image, correlating to a precise position within the array corresponding to its location in the image.
[0031] Following the scanning procedure, a dataset may be collected that comprises of microscopic images for a set of different working distances or z-offsets e.g., -5pm to 5pm with a non-uniform or uniform spacing of 1 pm. This dataset may be labeled as in-focus (InF) or out-of- focus (OOF) based on a determined z-offset (or focus depth) margin such that the images that lie beyond this margin are labeled as OOF while the images that lie within this margin are labeled asInF. The determining of an acceptable z-offset margin for the microscopic image datasets may be based on characteristics of the dataset, for example, the thresholds defined for lung tissues of a cancer subject stained with three specific dyes may not be applicable for the kidney tissues from another subject with chronic kidney disease stained with two other dyes. To address these variations among different datasets, context-aware thresholding may be adopted that may involve collecting a representative dataset that includes microscopic images captured for specific regions within the one or more reference slides. These reference slides may be captured for a range of z- offsets such that the dataset includes perfect in-focus images with zero offsets and the corresponding variations with non-zero z-offsets.
[0032] On this representative dataset, quantitative analysis, qualitative analysis, or a combination thereof may be performed to determine the z-offset margin. The quantitative analysis might involve identifying directly relevant variables e.g., presence of a particular biomarker, expected structure such as cells, nuclei, or vessels or indirectly, e g., by leveraging ML-based methods such as autoencoder, Fourier transform, texture analysis, or Laplacian filter. For these identified variables, a statistical analysis may be performed that may involve (combining or averaging) a quantitative measure (e.g., average of power spectrum of each image, cell count or percentage) computed for the images associated with each z-offset. Based on the quantitative measure, a percentage drop (e.g., 5% to 20%) from the value where the quantitative measure is maximum may be estimated. The z-offset margin may correspond to the z-offsets that fall above this percentage drop. Alternatively, or additionally, the domain experts can visually examine the images in the representative dataset (or corresponding to the z-offsets within the percentage drop) and provide assessments regarding focus quality and / or the accuracy of the estimated percentage drop. Based on these quantitative and / or qualitative assessments, z-offset margins may be determined to label or annotate the images corresponding to each z-offset or focal depth within the dataset, classifying them as either InF or OOF.
[0033] In some instances, the z-offset margin may be determined automatically by leveraging an autoencoder that is trained on the representative dataset and configured to generate a reconstructed in-focus image from an image with non-zero z-offset. When this autoencoder is provided with an out-of-focus image, the autoencoder may struggle to reconstruct it accurately. The reconstruction error may be higher for out-of-focus images compared to well-focused ones. Once the autoencoder is trained, it can be used to estimate the reconstruction error (differencebetween input non-zero offset image and output zero offset image) suggesting how far the input image is from its perfect focus. If the reconstruction error for a given microscopic image is below a specified threshold, the corresponding z-offset may be included in the z-offset margin.
[0034] In addition to these images featuring different z-offsets, this dataset may also include images labeled as non-tissue areas (NTA), representing sections within the microscopic image devoid of biological tissue. The identification of these areas may hold importance for precise analysis and enhancement of the classification accuracy of microscopic images. Non-tissue areas within microscopic images can encompass cell-free regions, portions displaying the glass slide surface, and / or vacant spaces surrounding the tissue specimen. These regions often exhibit uniform brightness or darkness, contingent upon imaging conditions and staining techniques.
[0035] Prior to input into an ML / DL-based QC tool, various preprocessing techniques may be applied to the microscopic images within the dataset. These techniques may apply different transformation to address problems such as noise and artifacts, inconsistent image resolution and / or varying intensities due to different acquisition conditions. For example, intensity normalization or z-score normalization may standardize the pixel intensity values, handling variations from different focus levels and z-offsets. Other transformations may include data augmentation (e.g., scaling, translation, random rotations, flips, cropping, reshaping and / or resizing), denoising to reduce noise and improve signal -to-noise ratio, histogram equalization to adjust contrast, and Gaussian filtering to enhance edges and fine details. These preprocessing may enable consistency in image size, resolution, and quality, making features more visible and easier for ML / DL models to learn. Additionally, sampling (or resampling) techniques may be used to prepare and manipulate image resolution. This may involve upsampling or downsampling to either increase or decrease resolution for maintaining spatial uniformity across the dataset. Different interpolation methods such as nearest neighbor, linear, or cubic interpolation can be employed to match the resolution of different images, reducing computation overhead or obtaining high-resolution images as needed. Incorporating these preprocessing techniques may enhance the dataset quality and consistency for ML or DL model training, fostering robustness and generalization.
[0036] Microscopic image segmentation that may be performed before or after preprocessing, may yield patches of equal or varying lengths. Using fixed or variable-size windows, patches may be extracted from the labeled microscopic images with consideration forrelevant information and minimal overlap. These labeled patches including NTA patches may be fed into an ML / DL model to classify each patch as either InF, OOF, or NTA. The ML / DL model may process each patch and generate probabilities indicating a likelihood of a patch belonging to a specific class. By passing patches through its layers, the model can extract and transform features while the output layer may convert raw scores into probabilities using functions e.g., SoftMax or sigmoid. The training process typically optimizes these probabilities to match true labels for accurate classification. To aggregate the classified patches for image-level decisions, statistical techniques may be employed that, based on a metric, combine predictions on patches corresponding to each label within a single image. For example, these methods may compute the metric such as calculating the percentage of patches assigned to each class label that can be determined by counting the total number of patches belonging to each category. Another method may consider the area covered by patches as a metric that can be estimated by multiplying the count of patches in each class by the area of the patch, which proves particularly useful when dealing with patches of different sizes. Alternatively, a weighted sum approach may be employed, where the number of patches assigned to each class is multiplied by a weight determined by the confidence score or certainty of the classifier.
[0037] Based on the computed metric, the statistical technique may assign scores to each predicted label. These classified patches may be placed back to the original image layout for reconstructing the microscopic image where, via a graphical user interface (GUI), different classes may be represented by the assigned scores and / or labels or metrics. Visual representations using color mapping or overlaying in a graphical user interface (GUI) can aid in interpreting the spatial distribution of the classified regions within the microscopic image. If the score attributed to the out-of-focus (OOF) region exceeds a predefined threshold, indicating significant focus issues, an alert may be activated. In response, the slide might be designated for rescanning by the scanner or validated by a user (or an expert) to determine whether rescanning is required or if it can be accepted. Similarly, if the OOF score falls below the predefined threshold, the image may be classified as in-focus and can be forwarded for further analysis or processing.
[0038] FIG. 1 is a block diagram illustrating an exemplary workflow 100 of an automatic quality control (QC) tool assessing focus quality of microscopic images. In various microscopy applications, including brightfield microscopy, darkfield microscopy, differential interference orphase contrast microscopy, the accuracy and reliability of images may contribute to precise diagnosis and analysis. Sample preparation techniques enable the acquisition of high-quality images in which samples 102 undergo processes to preserve their structural integrity and prevent degradation, which may include fixation of biological specimens using fixatives like formaldehyde or alcohol -based solutions. Subsequently, samples 102 may be embedded in a medium, commonly paraffin, to provide structural support during sectioning. The sample 102 may be cut into thin sections (e.g., with thickness of around 4-10 micrometers cut using microtome) and then placed onto glass slides. These slides serve as a supporting medium for examining and analyzing tissue sections. As an example, in brightfield microscopy, similar sample preparation techniques may be employed for achieving the quality of captured images. Tissue samples are fixed, embedded, and sectioned as described above. Additionally, staining protocols may be applied to enhance contrast and visualization. For instance, staining techniques using dyes like hematoxylin and eosin (H&E) are commonly utilized in brightfield microscopy to highlight specific cellular components, enabling the visualization of tissue structures under the microscope.
[0039] The prepared slides may be placed in a scanner 104, which can encompass various imaging devices such as microscopes, confocal scanners, brightfield or fluorescence scanners. These scanners are specialized devices designed to capture high-resolution images of stained tissue samples or other specimens. For example, a brightfield scanner illuminates the sample on the glass slide with white light and utilizes objective lenses to magnify the image, enabling detailed visualization of tissue and cellular morphology. Alternatively, confocal scanners utilize laser light and pinholes to selectively focus on specific planes within the specimen, offering improved optical sectioning and three-dimensional imaging capabilities. The scanner 104 may capture microscopic images that are transmitted further through a communication network 105 to the disclosed automatic QC tool 106 that may be run on one or more computer systems 108. These computer systems 108 may include various input and output devices (not shown) such as keyboard, mouse, stylus, display, or touchscreen, facilitating user interaction and data analysis. For instance, the computer system 108 may facilitate receiving image data depicting scanned slides from the scanner 104 to memory, enabling subsequent analysis and interpretation.
[0040] The computer system 108 may receive one or more images from the scanner 104 and perform assessment via the QC tool 106, which may involve assessing focus quality, non-tissuearea, or other image parameters. The QC tool 106 may score different regions of the microscopic images e.g., as in-focus (InF), non-tissue area (NTA) or out-of-focus (OOF) generating a QC output 110. If the OOF region score is above a predefined threshold indicating regions with more focus problems, an alert may be triggered indicating focus status of the microscopic image. In response, the slide may be assigned back to the scanner 104 for the rescanning. Similarly, if the OOF score is below the predefined threshold, the image may be considered as validated in-focus image, which can be sent further for further analysis or processing.
[0041] The image data, which encompasses various imaging modalities such as brightfield, fluorescence, and others, may include data related to color channels or frequency channels. Biological specimen, for example, a tissue section typically undergoes staining assays involving comprising one or more different biomarkers associated with chromogenic stains for brightfield imaging (e.g., multiplex images) or fluorophores for fluorescence imaging. Staining assays can use chromogenic stains for brightfield imaging, organic fluorophores, quantum dots, or organic fluorophores together with quantum dots for fluorescence imaging, or any other combination of stains and viewing or imaging devices. In the analysis of biological specimens, for example, cancerous tissues, different stains may be specified to identify one or more types of biomarkers, including immune cells.
[0042] Scanned microscopic images may include an array, usually a rectangular matrix, of pixels. Each “pixel” is one picture element and is a digital quantity that represents some property of the image at a location in the array corresponding to a particular location in the image.Typically, in continuous tone black and white images the pixel values represent a gray scale value. Pixel values for a digital image typically conform to a specified range. For example, each array element may be one byte (i.e., eight bits) representing pixel values in the range of 0 to 255. In a gray scale image, a “255” may represent absolute white and zero (‘0’) an absolute black (or visa-versa). Color images may comprise of three-color planes, generally corresponding to red, green, and blue (RGB). For a particular pixel, there is one value for each of these color planes, (i.e., a red component, a green component, and a blue component that is a three-dimensional pixel also called voxel). By varying the intensity of these three components, all colors in the color spectrum are typically created.
[0043] The communications network 105 connecting the scanner 104 with the computer system 108 may include, internet, an intranet, a wired LAN (local area network), a wireless LAN(WiLAN), a WAN (wide area network), a MAN (metropolitan area network), a PSTN (public switched telephone network), and other types of communications networks. The communications network 105 may further include communication devices such as one or more gateways, routers, or bridges to facilitate the transmission between the scanner 104 and the computer system 108. Merely by way of example, the communications network 105 can have one or more servers and one or more web-sites accessible by users to send and receive information usable by the one or more computer systems 108. The communications network 105 may be any type of network familiar to those skilled in the art that can support data communications using any of a variety of available protocols, including without limitation TCP / IP (transmission control protocol / Intemet protocol), SNA (systems network architecture), IPX (internet packet exchange), AppleTalk®, and the like.
[0044] Moreover, in various configurations, the automatic QC tool 106 may be running on a computer system 108 either embedded within the scanner or connected externally through the communication network. In the former scenario, the QC tool 106 may assess different regions of the microscopic images locally and transfer the in-focus images to the external systems via the communication network 105. Alternatively, when the QC tool 106 operates in a separate computer system 108, images may be transmitted directly from the scanner 104 to the computer system 108 using interfaces such as universal serial bus (USB) or other direct data transfer mechanisms. This may facilitate subsequent analysis and processing tasks, including data storage, conversion, and image processing, as required.
[0045] The computer system 108 of the exemplary workflow 100 may include a processing system with one or more high-speed central processing unit(s) (CPU), processors and one or more memories. The computer system 108 may also include a memory for storing a plurality of processing modules or logical instructions that are executed by the one or more processors coupled. The computer memory that stores data may also be maintained on a computer-readable medium including magnetic disks, optical disks, organic memory, and any other volatile (e.g., random access memory (RAM)) or non-volatile (e.g., read-only memory (ROM), flash memory, etc.) mass storage system readable by the CPU. The computer-readable medium may include cooperating or interconnected computer-readable medium, which exist on the processing system or can be distributed among multiple interconnected processing systems that may be local or remote to the processing system.
[0046] The computer system 108 may further include one or more databases (not shown) for the processing and storing of data (e.g., microscopic images such as brightfield or darkfield images). Database may be integral to a memory system on the computer or in secondary storage such as a hard disk, floppy disk, optical disk, or other non-volatile mass storage devices. The computer system 108 and the databases may be further connected to one or more communications networks. The computer system 108 may include a client terminal in communication with one or more servers, or personal digital / data assistants (PDA), laptop computers, mobile computers, internet appliances, one or two-way pagers, mobile phones, or other similar desktop, mobile or hand-held electronic devices.
[0047] FIG. 2 shows an example illustration 200 of in-focus and out-of-focus images when a sample 202 is scanned with a microscope 204. An accurate analysis particularly in brightfield imaging may be performed for a microscopic image obtained with a suitable focus. A typical microscopic imaging set up, as shown in example illustration 200 includes an objective lens with a focal plane 206. The objective lens, responsible for visualizing and magnifying the image of a sample 202, focuses light onto the specific plane known as focal plane 206 where the image of the sample 202 may be found sharpest and more detailed. The focal plane 206 represents a point at which the light rays converge to form a clear image. The position of the focal plane 206 from the objective lens is determined by different parameters of lens e.g., magnifying power, numerical aperture, and lens design. In brightfield microscopy, obtaining a correct focal plane 208 involves placing the sample 202 at a working distance 208 within a depth of focus 210 to enable precise detailing, resolution, and contrast. The working distance 208, denoted as z, is the distance between the objective lens and the sample 202 dictating how the light rays travel through the sample 202 and converge at the focal plane 206. For an image to be in focus, the sample 202 is precisely placed at a correct working distance 208 within the depth of focus 210 - a range along the z-axis within which a sample 202 remains in focus.
[0048] The depth of focus 210 is a measure of tolerance for slight misalignment of the sample 202 from the precise focal plane 206. It is influenced by various factors e.g. numerical aperture of the objective lens, the wavelength of the light used for imaging, and the magnification. Higher magnification lenses with higher numerical aperture values may have a shallower depth of focus. The focus issues may be broadly categorized into negative and positive offsets that may arise when the sample is not positioned at the correct working distance (z) 208.When the sample is placed high along the z-axis (i.e., towards the objective lens (z < 0)), the light rays may not converge sufficiently by the time they reach the image plane where the sample is placed (e.g., sample 202a) causing a negative offset 212. Positive offset 214 may arise when the sample 202 is positioned below the focal plane 208 (i.e., away from the objective lens (z > 0)). In this case, the light rays converge early and diverge before reaching the image plane where the sample is placed (e.g., sample 202b). Both positive and negative offsets can significantly degrade the quality of the microscopic images resulting in blurriness and low- resolution, making it difficult to accurately analyze the sample 202. FIG. 2 shows an example depiction of brightfield images taken with different offsets e.g., the images in top row 216 shows effect of negative offsets and images in the bottom row 218 shows positive offsets.
[0049] FIG. 3 illustrates an exemplary architecture 300 of the automated quality control (QC) tool that performs classification of microscopic images 302. To perform quality control on microscopic images 302, a dataset 304 may be collected for a set of different working distances or z-offsets 208 e.g., -2.5pm to 2.5pm with a uniform spacing of 0.5. Apart from these images with various z-offsets, this dataset 304 may also include non-tissue area (NTA) images that refer to regions of the microscopic image 302 that include no biological tissue. Identifying these regions may be significant for accurate analysis and improvement of the classification performance of microscopic images 302. Non-tissue area in microscopic images may include regions showing glass surface of the slide, a void of cells, and / or empty spaces around the tissue specimen that may often appear as uniformly bright or dark area depending upon the imaging condition and staining.
[0050] To prepare dataset 304 for various machine-learning (ML) / deep learning (DL) models, one or more preprocessing techniques 306 may be applied to the microscopic images 302. For example, microscopic images may vary in intensities due to different acquisition conditions. For consistency, transforms such as intensity normalization or z-score normalization can standardize the pixel intensity values across images of varying intensities. This may be particularly useful for handling variations caused by different focus levels and z-offsets. Other transforms may include data augmentation, denoising for reducing artifacts and noise to improve signal -to-noise ratio making relevant features prominent, histogram equalization for adjusting contrast of images and / or Gaussian filter for enhancing edges and fine details making features more visible in InF and OOF images and easier for the ML / DL model to learn. The dataaugmentation may include affine transformation e.g., scaling, translation, random rotations, random flips, and other transformations such as image cropping, reshaping, image resampling and resizing. The transformations such as sampling and resizing may help obtain uniform size and resolution across the dataset, thus enabling consistency in spatial characteristics.
[0051] For preprocessing 306 of the microscopic images 302, the sampling or resampling techniques may be used to prepare and manipulate the image before further processing. This may involve changing the spatial resolution of an image which can be done by e.g., upsampling or downsampling depending upon the aim to either reduce or increase the resolution of the image to have a consistency across the dataset. For example, downsampling that reduces the size of the image thereby decreasing numbers of pixels or voxels, may be useful for reducing computation overhead or matching the resolution of different images within the dataset. Similarly, upsampling may involve increasing the number of pixels or voxels for matching the resolution or to obtain a high-resolution image. For sampling i.e., upsampling or downsampling different interpolation methods such as nearest neighbor, linear or cubic interpolation may be used.
[0052] By leveraging these transforms, the dataset 304 may undergo a rigorous preprocessing 306 phase enhancing the quality and consistency of the data being fed into the ML model i.e., a deep learning (DL) model 314 for further processing. Moreover, the incorporation of the preprocessing techniques 306 for the disclosed QC tool 106 can make the DL model 314 exposed to a wide range of conditions during training, thus enhancing robustness and generalization capabilities of the model. Subsequent or prior to preprocessing 306, segmentation 310 may be performed on the microscopic images 302 yielding patches 312 (e g., of similar or varying lengths). For extracting patches, the segmentation 310 may be performed by spatially sliding a fixed window or a variable size window (e.g., of size 64x 64, 256x 256 pixels) horizontally and vertically over the entire image with a certain stride (e g., a 50% overlap or no overlap). The size of the patch may be selected such that each patch includes relevant information and avoids unnecessary overlapping with previously extracted patches. For creating ground-truth images, the patches 312 with z-offset between a determined z-offset margin 308 may be labeled as in-focus (InF) images while the rest of the patches 312 may be labeled as out- of-focus (OOF). These labeled patches 312 may also include NTA patches that may be fed into the DL model 314 to classify each patch 312 as either InF, OOF, or NTA (i.e., 315).
[0053] Segmentation of microscopic images can be performed using various other methods, so as to detect one or more portions of a microscopic image to be subsequently individually analyzed. Because of the high resolution of a digital pathology image, the segmentation can facilitate more efficient and feasible computational processing. One basic approach is thresholding, where pixel values are segmented based on intensity thresholds. Global thresholding may apply a single threshold value across the entire image, while adaptive thresholding may calculate thresholds locally for different regions, adapting to varying lighting conditions and contrasts within the image. The thresholding approach can (for example) result in part or all of a background area (that does not depict a part of a sample) being excluded from subsequent processing. Other segmentation techniques may include graph-based segmentation (that models each pixel as node and edges represent similarity between neighboring pixels), edge detection such as Sobel, Canny, or Laplacian filters. These methods focus on detecting edges within an image, which can then be used to define boundaries of regions or objects. Regionbased segmentation methods, such as region growing and region splitting and merging, partition the image into regions that are similar based on criteria such as intensity or texture. These techniques may start with seed points and expand or divide regions based on homogeneity.
[0054] For each patch, the deep learning model 314 may generate a predicted probability indicating an extent to which a patch belongs to a specific class. The model may predict patches 312 for predefined categories by passing the patches 312 through its multiple layers for extracting and transforming features. The last layer or the output layer of such model 314 may generate logits (raw, unnormalized scores) that may be converted into probabilities by a function such as SoftMax or sigmoid thereby providing a probabilistic interpretation of the network predictions. After computing probabilities for each class, the class with highest probability is chosen as the predicted class. The training process involves optimizing these probabilities to closely match the true labels so that the model 314 may classify each patch accurately.
[0055] The classified patches corresponding to an image may be aggregated to make decisions at the image level if needed. For aggregation 316, various statistical techniques may be used that combine the independent predictions on patches within a single image, based on a metric. One statistical approach may involve calculating percentage as a metric by counting the number of patches that fall into each class label. For these counts, the percentage of patches belonging to each class may be estimated. For instance, if a microscopic image is segmented into100 patches and predicted by the DL model 314 with 60 as in-focus, 30 as out-of-focus, and 10 as non-tissue areas, the respective percentages would be 60%, 30%, and 10%. Alternatively, area-based calculation may be performed that is useful if each patch covers a known area, such as 10x10 pixels, the count of patches in each class may be multiplied by the area of a single patch to get the total area covered by each class. Then, the percentages of the total image area each class represents may be determined. This method may provide a more intuitive understanding when dealing with varying patch sizes or when the physical area is more relevant than the count. For situations where patches have different levels of importance or confidence, a weighted sum approach can be employed. Assigning weights to each patch based on criteria e g., the confidence score of the classifier may allow for a more nuanced aggregation. Summing the weighted counts for each class and then calculating the weighted percentages may provide a refined understanding, considering the varying significance of different patches.
[0056] Aggregation 316 or reconstruction of the classified patches into the original image layout may allow for more sophisticated analysis, such as identifying and measuring contiguous regions of each class. By reconstructing the microscopic image, each patch may be placed back into its original position, creating a new image where different classes may be represented by distinct values and / or labels. Using image processing techniques, these contiguous regions can be highlighted and labeled based on the metric e.g., percentage, area, or count and be shown in a graphical user interface (GUI) as illustrated in QC output image 110 of FIG. 3. To achieve the display of patches with different highlighted or marked regions in a GUI, image processing techniques such as color mapping or overlaying can be utilized. For instance, patches labeled as in-focus (InF) can be overlaid with a no color, while out-of-focus (OOF) patches can be overlaid with a semi-transparent red tint, and non-tissue areas (NTA) with a semi-transparent blue tint. Alternatively, border colors or patterns can be applied to the patches to indicate their labels. These visual enhancements can be implemented programmatically within the GUI framework, allowing for the dynamic display of patches with distinct labels, aiding in visually assessing and interpreting the dataset. This pictorial representation may be particularly useful for visualizing the spatial distribution and extent of each class within the image.
[0057] Moreover, this spatial analysis and visual representation via the GUI can be used to understand the distribution and continuity of patches within the image by aggregating or grouping adjacent patches of the same class into larger regions. This approach can revealpatterns and structures within the image that may not be apparent when considering patches individually, offering insights into the spatial relationships and clustering of different classes. Combining these methods may provide a robust framework for analyzing and interpreting the classification results of image patches. Whether through simple counting, area-based calculations, image reconstruction, weighted sums, or spatial analysis, each method may add up offering insights, enabling understanding of the image composition and the distribution of infocus, out-of-focus, and non-tissue regions.
[0058] The determination of z-offset margin 308 can be challenging in labeling the dataset 304. The determination of z-offset margin may be based on characteristics of the dataset such as types of tissue, staining protocols and expected biological variations. For example, tissue samples from different organs, such as lungs and kidneys may exhibit distinct cellular structures and e.g., if stained with certain dyes / chromogens may exhibit different staining patterns. Similarly, samples from a healthy subject versus those with diseases e.g., cancer or chronic kidney diseases may present different cell distributions, densities, or stain intensities. To address these variations within the dataset, context-aware thresholding may be adopted to determine a z- offset margin. This process may involve creating a representative dataset of microscopic images captured from one or more reference slides of samples from the target context to determine directly relevant variables (e.g., a cell count or stain intensities). These reference slides may be captured for specific regions by perturbing the focus points of the microscope for a range of z- offsets (e.g., approximately 10 to 20 different values such as -6pm to 6pm with uniform spacing of 0.5). This way a series of images at various focal planes can be captured, thus giving a view of the samples with different visibility levels of structures and details.
[0059] On this representative dataset, a quantitative analysis and / or qualitative analysis may be performed to determine z-offset margin. The quantitative analysis may include identifying relevant variables directly (e.g., cells, vessels, a specific biomarker) or indirectly by methods (e.g., ML-based such as autoencoder, Fourier transform, texture analysis or Laplacian filter). For these identified variables, a statistical analysis may be performed that may involve aggregating a quantitative measure (e.g., reconstruction error, average of power spectrum of each image, cell count or percentage) computed for the images associated with each z-offset. Based on qualitative or empirical analysis, a percentage drop (e.g., 5% to 20%) from the value where the quantitative measure is maximum may be estimated. The dataset may be labeled or annotated correspondingto each z-offset or focal depth within the dataset as InF (i.e., mapping to a z-offset of 0) or OOF based on the percentage drop that defines acceptable z-offset margin.
[0060] For example, quantitative analysis may include identifying directly one or more relevant (quantitative) variables from the representative dataset that influence the quality of focus or characteristics of the samples. These variables may include, but are not limited to, type of tissue, presence of a particular biomarker, expected structure such as cells, nuclei or vessels count or percentages. For example, slides stained with hematoxylin and eosin (H&E) will have different visual (e.g., color, contrast, and target) and quantitative characteristics (e g., cell densities, morphometric properties such as cell nucleus, size, shape) compared to those stained with immunohistochemistry (IHC) markers. Once the relevant variables are identified, context specific thresholds can be determined. This may involve analyzing the representative dataset to determine acceptable ranges of these variables by leveraging image processing techniques, statistical techniques, qualitative methods, machine-learning (ML) based methods, or a combination thereof.
[0061] If, for example, the relevant variable is based on morphological properties then boundary detection and structures of interest within the images can be focused followed by a statistical analysis of the data e.g., counting, calculating mean, medians, or standard deviation. Analyzing boundaries of the specimen within the image may involve, for example, accurately counting relevant structures or features. For analyzing boundaries or morphological structures such as cells or vessels, the images can be processed to segment individual structure (cell) to identify and delineate the boundaries of the structure within the images. Once the particular structure is segmented for the entire image, automated counting algorithms or manual counting methods can be used to determine the total number of structures in each image. Automated counting algorithms may rely on image processing techniques, such as thresholding, edge detection, or ML-based approaches, to detect and count structures automatically. Manual counting methods may involve visual inspection of the images by human operators to manually count the structures. This counting of morphological structures may be performed for all the images within the representative dataset that correspond to each z-offset. At each z-offset, the count from all the images may be combined to determine a statistical value e g., a mean or a median. Based on this statistical data, an acceptable z-offset margin, representing the range of z- offsets within which the counting accuracy meets a predetermined criteria or quality thresholds,can be identified. For example, for counting the cell structures if the mean value of counting accuracy remains above a certain threshold (e.g., 85-95%) for a given z-offset (e.g., ±0.5pm (microns) from the focus plane), then this offset can be included in the z-offset margin.
[0062] In some examples, an autoencoder may be used to automatically determine acceptable z-offset margin. For training, the representative dataset may be arranged in input-output pairs such that each image with a non-zero z-offset is paired with a corresponding perfect focus image i.e., with zero offset. The autoencoder comprises of an encoder that takes the input non-zero offset image (OOF), compresses it to a latent space, and a decoder that aims to reconstruct the corresponding zero offset (in-focus) image from this latent representation. The autoencoder may be trained to minimize the reconstruction loss between the input image and the reconstructed output (zero offset image). The reconstruction loss may be computed using metrics such as mean squared error (MSE) or binary cross entropy indicating how well the autoencoder can recover details from the input. After training, a reconstruction error (i.e., difference between input and output) may be compared with a specified threshold for each image as a quantitative measure suggesting how far an input image is from being perfectly focused. For example, if the reconstruction error is below the specified threshold (e.g., 0.2), the image may be considered as in-focus. Higher reconstruction losses may indicate greater deviations from the target perfect focus. This approach may be used to determine the acceptable z-offset margin by considering the z-offsets of the images for which the reconstruction error is found to be lower than the specified threshold.
[0063] The quantitative analysis including the Fourier analysis may involve transforming spatial components of the image data into frequency components that are indicative of the level of details and sharpness. Sharp, in-focus images typically exhibit higher frequency components concentrated in the center of the frequency spectrum, whereas blurred or out-of-focus images show diffuse frequency components. A quantitative measure e.g., sum or average of the power spectrum values within a certain radius from the center of frequency domain may be defined that can be averaged for the images associated with a specific z-offset. The z-offset margin may be determined for a percentage drop (e.g., 5-10%) from the value where this measure is maximum. The percentage drop can be further validated by the expert by visually inspecting the acceptable visibility corresponding to the specific z-offset that can be included in z-offset margin. Similarly, Laplacian filter is a second-order derivative filter that highlights regions of rapid intensitychange, making it useful for edge detection and assessing focus. When Laplacian filter is applied on the image it may generate a map of edges, with sharp, in-focus images producing more defined edges and higher variance in Laplacian response. The quantitative measure may be computed by averaging sum or variance of the Laplacian responses for the images associated with the specific z-offset. The threshold may be set based on the percentage drop (that may be validated by the expert) from the maximum value of the quantitative measure.
[0064] FIG. 4 illustrates an example architecture of a densely connected convolutional neural network (DenseNet) 400 performing classification of the microscopic images 302. Neural networks are known for efficiency and performance in handling complex imaging data for various classification tasks. In an example implementation, densely connected convolutional neural network also known as DenseNet model 400 fine-tuned for a three-class classification problem (i.e., to classify a patch as OOF, InF, or NTA) is leveraged. It may be understood that other neural networks capable of classification may also be used, which may include fine-tuned ResNet (residual network), VGG (visual geometry group), inception network (GoogLeNet), AlexNet and other similar neural networks. DenseNet 400 is a deep learning architecture that features dense connections between layers to enable maximum flow of information and gradient propagation, which may help in mitigating vanishing gradient problem often encountered in deep neural networks. DenseNet 400 may be characterized by its specific connectivity pattern where each layer may receive additional inputs from all previous layers and pass on its feature-maps to all subsequent layers. In this way, each layer may receive a collective knowledge from all the preceding layers, resulting in diversified features that tend to have richer patterns. This dense connectivity may significantly improve reuse of features and flow of gradient, making DenseNet 400 efficient for various deep learning tasks including classification. By reusing features from preceding layers, where each layer contributes different features, DenseNet 400 may reduce redundancy and achieve performance efficiency with fewer parameters. Moreover, in DenseNet 400, classifier may use features of all complexity levels that tend to give more smooth decision boundaries and good performance when training data is insufficient.
[0065] The network 400 may comprise of multiple layers such as initial layer 402, one or more dense blocks 406, one or more transition layers 408 and a final classification layer 410 for performing classification task. The initial layer 402 may further include one or more convolution layers to reduce spatial dimensions of the input patches 312 (i.e., by having a larger kernel ofsize e.g., 7X7 and a stride value), normalization layer for normalizing data e.g., batch normalization, activation layer to introduce non-linearity e.g., Rectified linear unit (ReLU) and an optional pooling layer for further reducing the dimensions of the input. The dense block 406 in DenseNet 400 may comprise of multiple dense layers (e.g., 404a, 404b, 404c and 404d) that may further include layers that are connected by the specific pattern of DenseNet i.e., each layer is connected to every other layer in a feed-forward manner. For example, within the dense block 406, the input from node 404a, 404b and 404c is fed to 404d where each node represents a dense layer as illustrated in FIG. 4. All these inputs from preceding layers may be concatenated that suggests the size of feature maps to be the same within a dense block 406. Typical dense layer 404 may include one or more layers of normalization, activation, one or more convolution layers (e.g., typically a convolution layer with smaller kernel size such as 1X1 followed by another convolution layer with larger kernel size such as 3X3).
[0066] The dense blocks 406 may be stacked together (e.g., such that with each dense block the size of feature maps may be reduced) in the DenseNet 400 where between these blocks may reside one or more transition layers 408 that may be used to downsample the feature maps, reducing the spatial dimensions and number of channels. A transitional layer 408 may comprise of normalization layers, convolution layer and pooling layer, which can be used as the transition layers between two contiguous dense blocks 406. If a dense block 406 outputs m feature maps, the transition layer 408 generates 6m output feature maps, where 0 < 6 < 1 is referred to as the compression factor. At the end of the last dense block, the final classification layer 410 resides that may include an optional average pooling layer and a fully connected layer that maps the pooled features to the output classes (i.e., 110). Additionally, the final classification layer 410 may include activation functions e.g., SoftMax that predicts the probabilities for multiple classes such as predicting the probability of a patch belonging to each class, where the class with the highest probability is chosen as the predicted class. DenseNet 400 may leverage multiple dense blocks and transition layers. For example, DenseNet-121 includes a set of layers comprising an initial layer, four dense blocks, three transition layers (after each dense block), and final layers comprising average pooling and fully connected layers. For training, the objective function (also known as loss function) of the DL model 314 may be selected as categorical cross entropy for the multi-class classification problem between the label and the prediction.
[0067] FIG. 5 illustrates an exemplary method 500 of determining a z-offset (focus depth) margin 308 for labeling the microscopic images as in-focus (InF) or out-of-focus (OOF). The microscopic images 302 captured with a z-offset of 0 value may be considered as precisely infocus but determining this margin may be significant because even with non-zero z-offset values, the result may be slightly unfocused images maintaining sufficient quality for downstream analysis. To determine the z-offset margin 308, the representative dataset may be collected from one or more reference slides at different focus depths or z-offset. This representative dataset comprising microscopic images may include singleplex images of CD8 (cluster of differentiation) in which only the CD8 biomarker is visualized. (It will be appreciated that biomarkers other than CD8 may instead by visualized.) The slides may be stained by counterstain hematoxylin for staining cell nuclei and providing contrast to CD8. On this representative dataset, quantitative analysis may be performed that by leveraging a computer vision tool or image processing tool 502 such as HALO (High-throughput Analysis and Learning Optimized) predict the cell count (which is considered as directly relevant variables) for various z-offsets as mentioned in the discussion above. The predicted cell counts from the images associated with each z-offset within the representative dataset may be combined or averaged and a corresponding percentage may be computed that represents the percentage of correctly predicted cells at each z-offset.
[0068] From the maximum percentage of cell count that may be associated with a z-offset value of 0, an acceptable percentage drop of e.g., -5% or -20% may be set as the cut-off threshold defining the threshold margin. This percentage drop may be validated by the expert that can visually inspect the images at the cut-off z-offset whether to be a valid threshold. Optionally, input of its variation such as CD8 images highlighting positive cells (i.e., 504a) and negative cells 504b may also be fed to the image processing tool 502 for respective cell counting. CD8 is a glycoprotein biomarker found on the surface of certain immune cells, primarily cytotoxic T cells. These cells can be involved in recognizing and destroying cells infected with viruses or other pathogens, as well as cancerous cells. When referring to CD8 positive cells, it means cells that express the CD8 protein on their surface that are typically cytotoxic T cells, while CD8 negative cells are cells that do not express the CD8 protein.
[0069] FIG. 6 illustrates one or more plots for determining z-offset margin 308 for the exemplary method 500 of the FIG. 5. The plot 602 illustrates a curve showing the percentagecount of accurately predicted CD8 positive cells (drawn at y-axis) at specific z-offsets values (drawn at x-axis) ranging from [-6pm, 6pm] for the given CD8 positive cell images 504a at different focal depths. This plot 602 represents how the accuracy of predicted cell counts changes with different z-offsets. It may be noticed from the plot 602 that increasing focal depths positively or negatively decreases the image resolution and visibility resulting in blurred images. The plot 602 has a peak (or maximum) at the z-offset that refer to a precise focus (in this case, approximately around 0pm). The selection of an acceptable percentage drop from this peak value may be determined empirically by observing the percentage of predicted cell counts and the associated z-offset. Based on this observation, a 5% change from the z-offset value of 0, marked by the red dashed line, is considered as in-focus, while those below the red dashed line are classified as out-of-focus. The red dashed line shows an approximately accurate cell count above a threshold e.g., 75%. Similarly, the plot 604 illustrates the same plot as 602 with y-axis labeled for total number of CD8 positive cells predicted by the image processing tool at various z-offsets for the given CD8 positive cell images 504a. Finally, the plot 606 illustrates a total cell count corresponding to the CD8 image 504c.Example Implementation:
[0070] In an example implementation, a dataset 304 comprising InF, OOF and NTA images with a range of focus depths is collected by varying the focus offsets of a microscope. The dataset 304 is comprised of digital slides with different stains, where each slide was scanned with magnification of 20x. For each whole slide, the microscopic images were captured with various z-offsets e.g., [-5 pm to 5pm] with a uniform spacing of 0.5. This way the same slide was captured with different focal planes some of which not aligned with the sample height, thus resulting in different levels of blurring (out of focus). The model was trained for different patch sizes e.g., 64x64, 128x128 and 186x 186 extracted from the microscopic images. For training and testing, the train-test split ratio was set to 75-25%.
[0071] These images in the dataset 304 were preprocessed and adjusted using different transformations e.g., an appropriate sampling, as stated in the above discussions. For implementing the ML / DL-based quality control (QC) tool 106, Pytorch is leveraged as a deep learning framework. The model 314 is trained on images from the hematoxylin channel of the Flash scanning hardware. This Flash multiplexing technology features distinct images, designedbased on DISCOVERY RUO (Universal Secondary Antibody, an advanced instrument used in immunohistochemistry (IHC) and in situ hybridization (ISH) research) chromogens generating unique colors. Each chromogen absorbs a particular range of light wavelengths, resulting in pronounced and unique colors in the images produced by this technology. Flash technology employs a diverse dye portfolio paired with precise LED illumination. The 635 nm channel was determined to represent the suitable (optimal) channel for the module to analyze because hematoxylin counterstain is ubiquitously used in IHC assays, and therefore this channel will have more applicability across different assays. In this example implementation, the QC tool is developed leveraging a pre-trained DenseNet model and fine-tuned for a three-class classification.
[0072] FIG. 7 illustrates one or more graphs showing performance of the disclosed technique in accordance with an example implementation. In FIG. 7, an example pie graph 702 is depicted showing distribution of the collected dataset with a total size of 148 GB (giga byte). In this dataset, in-focus (InF), out-of-focus (OOF) and non-tissue area (NT A) images contribute to approximately 32%, 40% and 28%, respectively. FIG. 7 also illustrates a training loss plot 704 in which loss is on y-axis and epoch on x-axis during the training process. At the start (epoch 0), the loss is higher because the model parameters are randomly initialized, and the model 314 has not learnt. As training progresses, the model 314 leams from the input dataset 304, demonstrating a decreasing trend (i.e., falling approximately below 0.075 after 175 epochs). This plot 704 shows an improvement in prediction that the model 314 has achieved by adjusting its parameters to minimize loss function.
[0073] The performance of the QC module 106 was tested on the collected IHC images, which were manually scored as either in focus or out of focus. Well-focused images were collected by providing several focus points falling within the z-offset margin 308 for each image that is closer to (z=0). All images collected in this manner were labeled as in-focus images. Subsequently, the z-offset of the objective lens was intentionally altered outside the z-offset margin 308 to create defocused images, which were then labeled as out-of-focus. For accuracy of the labeling, edge detection was performed on the images, and both brightfield and edge-detected images were manually inspected. A plot 706 was generated using 35 in-focus images and 41 out- of-focus images. The in-focus class is represented by the green squares, while the out-of-focus class is represented by the red circles. The y-axis denotes the percentage of the images that arein-focus, and the x-axis represents the percentage of the images that are out-of-focus. Each point in the plot 706 represents the performance of the QC module 106 on a specific set of images. Some values may not sum up to 100% due to the presence of three classes: in-focus, out-of- focus, and non-tissue area (not depicted in the plot 706).
[0074] In another experiment, the efficacy of the QC module 106 was evaluated using a set of 24 images. These images were manually categorized as either in-focus or out-of-focus. Following the methodology outlined in the previous example of 706, the dataset was created. For precise labeling, edge detection was applied to the images, which were then meticulously reviewed both in their original form and as edge-detected versions. Subsequently, a graphical representation 708 was generated utilizing 12 in-focus and 12 out-of-focus images. The focused category is denoted by the green squares, while the out-of-focus category is indicated by red circles. The y-axis illustrates the proportion of the images in-focus, while the x-axis represents the proportion of the images out-of-focus. It may be observed that the values may not add up to 100% due to the presence of three distinct classes: in-focus, out-of-focus, and non-tissue area (not depicted in the plot 708). By considering 50% as the threshold between in-focus and OOF, the module accurately classified 22 out of 24 images.
[0075] FIG. 8 illustrates an example 800 showing a quality control (QC) image of a first microscopic image 802, which is generated by the disclosed QC tool 106. The first microscopic image 802 800 represents a standard brightfield image from the collected dataset of IHC images. The example 800 further illustrates a corresponding edge-detected image 804, and the QC image 806 depicting results from the QC tool 106. In edge detected image 804, areas highlighted in green boundary signify OOF pixels. The QC tool 106 exhibits proficient OOF detection, clearly marked by the highlighted regions in a solid red color in the QC image 806. Additionally, NTA pixels, in QC image 806, are marked with a solid blue color, showcasing the accuracy of the model in handling NTA detection. The QC tool 106 provides a graphical illustration of the numerical values for all three classes, as evident in the QC image 806, for quantitative assessments.
[0076] FIG. 9 shows another example 900 illustrating the performance of the QC tool 106 for the second microscopic image 902 that represents a standard brightfield image. Meanwhile, the edge-detected image 904 highlights the edges or boundaries of the objects within the image that are OOF. The QC image 110 serves as the primary output of the QC tool 106, offeringinsights based on the analysis of both the standard brightfield image 902 and edge-detected image 904. Within this QC image 110, areas highlighted in red signify OOF pixels. The QC tool 106 adeptly detects these OOF pixels, distinctly marked by slightly transparent red highlights, indicating accurate identification of regions lacking focus. Furthermore, the QC image 110 marks pixels with a solid blue color to represent non-tissue area (NTA) pixels. Thus, the QC tool 106 demonstrates its capability in effectively handling NTA detection. Additionally, the QC tool 106 provides pictorial demonstration of numerical values for all three classes (brightfield InF, OOF, and NTA), facilitating quantitative assessment of the image quality.
[0077] FIG. 10 illustrates an example 1000 illustrating the performance of QC tool 106 for a third microscopic image 1002. The FIG. 10 illustrates a third example of a brightfield image 1002 taken from the collected IHC image dataset. The corresponding edge-detected image 1006 highlights the region that is OOF. Feeding this brightfield image to QC tool 106 generates the QC image 1004 where OOF region is highlighted by red and NTA area is highlighted by distinctive blue color. A different visualization of the QC image 1008 can be observed that highlights the boundaries of the patches with green, red, and blue representing InF, OOF, and NTA, respectively (as opposed to QC image 1004 that fills the patches).
[0078] FIG. 11 illustrates an example process flow 1100 for determining a focus quality of the microscopic images 302. The blocks in process flow 1100 are illustrated in a specific order, while the order can be modified, for example, some blocks may be performed before other, and some blocks may be performed simultaneously. The block can be performed by hardware, software, or a combination thereof. The process at block 1102 may include receiving a microscopic image 302 that depicts a slide including a slice of a sample. In various microscopy techniques, including brightfield, darkfield microscopy, this sample may be prepared by employing various techniques for the acquisition of high-quality images. In these techniques, samples may be preserved by different fixatives and embedded in a medium for providing structural support during sectioning. The sample may be then cut into thin slices and placed onto glass slides. Additionally, staining protocols may be applied for improving contrast and visualization. The prepared slides may be placed in a scanner such as microscopes, confocal scanners, brightfield or fluorescence scanner that are designed to capture high-resolution images of the samples or other specimens.
[0079] At block 1104, the captured microscopic image 302 e g., histopathology image, IHC image, fluorescence image may be segmented by spatially sliding horizontally and vertically a window over the entire microscopic image with a stride to extract a set of patches. The size of the window may be selected considering factors such as avoiding unnecessary or redundant information that may occur for overlapping windows and acquiring a suitable classification resolution for each patch. The value of stride may result in overlapping or non-overlapping window. For example, if the side length of a square window is kF, a stride of VF / 2 may result in 50% overlap, similarly a stride of VF will result in no overlap. In one example, segmentation results in non-overlapping and uniform sized patches. Alternatively, a variable window size may be used when different regions of the microscopic image may include structures of varying scales. At block 1106, the extracted set of patches may be fed into the deep learning model 314 that is configured to generate a predicted probability indicating an extent to which a patch corresponds to each predefined label. These predefined set of labels may include in-focus (InF), out-of-focus (OOF), and non-tissue area (NTA).
[0080] The deep learning model may be trained on a labeled dataset comprising a set of microscopic images 302, alternatively, a set of patches 312 extracted from these microscopic images 302, captured for a range of focus depths (e.g., -3 pm to 3 pm with non-uniform or uniform spacing of 1pm). These patches (or microscopic images) with z-offsets between a determined z-offset (or focus depth) margin 308 may be labeled as in-focus (InF) images while the rest of the patches (or microscopic) may be labeled as out-of-focus (OOF). These labeled patches including additional NTA patches may be fed into the ML / DL model for the prediction task. At block 1108, the predicted patches corresponding to the microscopic image 302 may be aggregated by various statistical techniques that, based on a metric, combine the independent predictions on patches corresponding to each label within a single image. For example, the statistical approach may involve computing the metric such as percentage of patches that fall into each class label (by counting total number of patches belonging to each class), area covered by patches (specifically useful for varying patch sizes), or weighted sum of number of patches belonging to each class where weights may be assigned by a confidence score of the classifier. Based on the computed metric, the statistical technique may assign a set of scores to each predicted label. At block 1110, if it is determined that the score assigned to OOF region is above a threshold indicating regions with more focus problems, an alert may be triggered at block1112. In response to this determination, the slide may be assigned back to the scanner for the rescanning, alternatively, it may be validated by a user (or an expert whether to be rescanned or accepted). Similarly, if the OOF score is below this threshold, the image may be considered as an in-focus image, which can be sent for further analysis or processing.
[0081] The disclosed QC tool may further include reconstructing the microscopic image by placing each patch back into its original position, creating a reconstructed image where different classes may be represented by the assigned scores. Using image processing techniques, these contiguous regions can be highlighted with different colors and labeled (e.g., by predicted labels such as InF, OOF or NTA, and / or the scores such as percentages, areas etc.) and shown in a graphical user interface (GUI). This method may be particularly useful for visualizing the spatial distribution and extent of each class within the image.
[0082] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods and / or part or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer-program product tangibly embodied in a non-transitory machine- readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods and / or part or all of one or more processes disclosed herein.
[0083] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.
[0084] The present description provides preferred exemplary embodiments only, and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the presentdescription of the preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It is understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.
[0085] Specific details are given in the present description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method comprising: receiving a microscopic image depicting a slide that includes a slice of a sample; segmenting the microscopic image by spatially sliding a window with a stride over the microscopic image to extract a set of patches; generating a set of predicted labels associated with the set of patches by leveraging a deep learning model that is configured to generate a predicted probability indicating an extent to which a patch corresponds to a label of a set of labels, wherein the deep learning model is trained on a labeled dataset comprising a plurality of patches extracted from a set of microscopic images captured for a range of focus depths; generating a set of scores for each predicted label of the set of predicted labels by a statistical technique that based on a metric aggregates one or more patches of the set of patches corresponding to each predicted label of the set of predicted labels; determining that a score of the set of scores is above a threshold; and triggering, based on the determination, an alert for a rescanning of the slide.
2. The computer-implemented method of claim 1, further including: outputting, via a graphical user interface (GUI), a reconstructed microscopic image by placing the set of patches into an original layout of the microscopic image; and applying an image processing technique for highlighting the one or more patches corresponding to each predicted label of the set of predicted labels within the reconstructed microscopic image.
3. The computer-implemented method of claim 1, further including: determining a margin of focus depth within which a patch is assigned a label of the set of labels as in-focus by averaging a cell count at a given focus depth of the range of focus depths for which the cell count is above a given threshold; and assigning the label of the set of labels to each patch of the plurality of patches within the labeled dataset based on the determined margin of focus depth.
4. The computer-implemented method of claim 1, wherein the deep learning model is a neural network including one or more dense blocks that comprise of one or more dense layers such that each dense layer receives feature maps from one or more preceding layers.
5. The computer-implemented method of claim 1, further including: preprocessing the labeled dataset by leveraging one or more transforms that include normalization, data augmentation, or sampling.
6. The computer-implemented method of claim 1, wherein the set of labels includes infocus (InF), out-of-focus (OOF), and non-tissue area (NTA).
7. The computer-implemented method of claim 1, wherein Pytorch is leveraged for implementing the deep learning model.
8. A system comprising: one or more data processors; and a non-transitory computer-readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform actions including: receiving a microscopic image depicting a slide that includes a slice of a sample; segmenting the microscopic image by spatially sliding a window with a stride over the microscopic image to extract a set of patches; generating a set of predicted labels associated with the set of patches by leveraging a deep learning model that is configured to generate a predicted probability indicating an extent to which a patch corresponds to a label of a set of labels, wherein the deep learning model is trained on a labeled dataset comprising a plurality of patches extracted from a set of microscopic images captured for a range of focus depths; generating a set of scores for each predicted label of the set of predicted labels by a statistical technique that based on a metric aggregates one or more patches of the set of patches corresponding to each predicted label of the set of predicted labels; determining that a score of the set of scores is above a threshold; andtriggering, based on the determination, an alert for a rescanning of the slide.
9. The system of claim 8, further including: outputting, via a graphical user interface (GUI), a reconstructed microscopic image by placing the set of patches into an original layout of the microscopic image; and applying an image processing technique for highlighting the one or more patches corresponding to each predicted label of the set of predicted labels within the reconstructed microscopic image.
10. The system of claim 8, further including: determining a margin of focus depth within which a patch is assigned a label of the set of labels as in-focus by averaging a cell count at a given focus depth of the range of focus depths for which the cell count is above a given threshold; and assigning the label of the set of labels to each patch of the plurality of patches within the labeled dataset based on the determined margin of focus depth.
11. The system of claim 8, wherein the deep learning model is a neural network including one or more dense blocks that comprise of one or more dense layers such that each dense layer receives feature maps from one or more preceding layers, wherein Pytorch is leveraged for implementing the deep learning model.
12. The system of claim 8, further including: preprocessing the labeled dataset by leveraging one or more transforms that include normalization, data augmentation, or sampling.
13. The system of claim 8, wherein the set of labels includes in-focus (InF), out-of-focus (OOF), and non-tissue area (NTA).
14. A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform action including:receiving a microscopic image depicting a slide that includes a slice of a sample; segmenting the microscopic image by spatially sliding a window with a stride over the microscopic image to extract a set of patches; generating a set of predicted labels associated with the set of patches by leveraging a deep learning model that is configured to generate a predicted probability indicating an extent to which a patch corresponds to a label of a set of labels, wherein the deep learning model is trained on a labeled dataset comprising a plurality of patches extracted from a set of microscopic images captured for a range of focus depths; generating a set of scores for each predicted label of the set of predicted labels by a statistical technique that based on a metric aggregates one or more patches of the set of patches corresponding to each predicted label of the set of predicted labels; determining that a score of the set of scores is above a threshold; and triggering, based on the determination, an alert for a rescanning of the slide.
15. The computer-program product of claim 14, further including: outputting, via a graphical user interface (GUI), a reconstructed microscopic image by placing the set of patches into an original layout of the microscopic image; and applying an image processing technique for highlighting the one or more patches corresponding to each predicted label of the set of predicted labels within the reconstructed microscopic image.
16. The computer-program product of claim 14, further including: determining a margin of focus depth within which a patch is assigned a label of the set of labels as in-focus by averaging a cell count at a given focus depth of the range of focus depths for which the cell count is above a given threshold; and assigning the label of the set of labels to each patch of the plurality of patches within the labeled dataset based on the determined margin of focus depth.
17. The computer-program product of claim 14, wherein the deep learning model is a neural network including one or more dense blocks that comprise of one or more dense layers such that each dense layer receives feature maps from one or more preceding layers.
18. The computer-program product of claim 14, further including: preprocessing the labeled dataset by leveraging one or more transforms that include normalization, data augmentation, or sampling.
19. The computer-program product of claim 14, wherein the set of labels includes in-focus (InF), out-of-focus (OOF), and non-tissue area (NTA).
20. The computer-program product of claim 14, wherein Pytorch is leveraged for implementing the deep learning model.
Citation Information
Patent Citations
Machine-learning techniques for detecting artifact pixels in images
WO2023064186A1
Cited By
Cell occlusion relation estimation method and system based on pixel-level focus evaluation curve
CN121904378A