Systems and methods for region of interest detection using slide thumbnail images
By classifying slide thumbnail images into five types and designing a custom detection algorithm, the problems of low accuracy and efficiency in slide thumbnail image detection are solved, achieving efficient and accurate region of interest detection and improving detection precision and recall.
Patent Information
- Application Number
- CN202111098397.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-01-31
- Filing Date
- 2016-01-29
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2036-01-29
AI Technical Summary
In existing technologies, the detection of regions of interest (AOI) in slide thumbnail images suffers from low accuracy and efficiency, especially when processing different types of slides, making it difficult to achieve efficient and accurate automatic AOI detection.
An AOI detection system and method based on tissue morphology and layout is provided. By classifying slide thumbnail images into five types and designing customized detection algorithms for each type, including ThinPrep (RTM), Tissue Microarray (TMA), control slides, smears and general slides, the accuracy of detection is ensured by using soft-weighted images and error correction modules.
It achieves efficient and accurate AOI detection on different types of wafers, improving detection accuracy and recall rate, which is superior to existing methods, and is verified by real ground data.
Smart Images

Figure CN113763377B_ABST
Abstract
Description
[0001] This application is a divisional application of application number 201680007777.8 filed on January 29, 2016, entitled "System and method for detecting region of interest using a slide thumbnail image". Technical Field
[0002] This disclosure relates to digital pathology. More specifically, it relates to tissue examination on slides containing tissue biopsies used to generate full-slide scans. Background Technology
[0003] In the field of digital pathology, biological samples such as tissue sections, blood, cell cultures, etc., may be stained using one or more staining agents and analyzed by viewing or imaging the stained samples. Various processes (including disease diagnosis, prognostic and / or predictive assessment in response to treatment), and the development of new drugs to combat diseases are achieved by observing the stained samples in conjunction with additional clinical information. As used herein, a target or object is a feature of the sample identified by the staining agent. A target or object may be a protein, protein fragment, nucleic acid, other objects of interest recognized by an antibody, a molecular probe, or a nonspecific staining agent. Those targets specifically identified may be referred to as biomarkers in the disclosure of this subject matter. Some staining agents do not explicitly target biomarkers (e.g., the commonly used counterstaining agent hematoxylin). Although hematoxylin has a fixed relationship with its targets, most biomarkers can be identified using the user's staining agent selection. That is, a wide variety of staining agents can be used to visualize specific biomarkers depending on the specific needs of the analyte. After staining, the analyte can be imaged for further analysis of the contents of the tissue sample. Imaging involves scanning slides containing stained tissue samples or biopsies. Staining is crucial because the stains required for cellular / microscopic structures are indicators of the tissue's pathological condition and can be used for medical diagnosis. In manual mode, trained pathologists read the slides under a microscope. In digital mode, the entire slide scan can be read on a computer monitor or automatically scored / analyzed throughout the scan using imaging algorithms. Therefore, high-quality scanning is absolutely essential for efficient use in digital pathology.
[0004] For high-quality scanning, proper detection of tissue on the glass slide is crucial. The tissue area on the slide is called the region of interest (AOI). This AOI detection can be done manually from thumbnail images, which are low-resolution captures of the slide. However, for rapid batch scanning, AOI extraction from slide thumbnail images is necessary. Unfortunately, the first step in existing methods for acquiring thumbnail images before any further analysis is limited to using low-resolution thumbnail cameras, and accurate and precise AOI detection is extremely challenging, especially considering the variability of data from different types of slides with varying staining intensities, morphology, tissue structures, etc. The use of individual AOI detection methods for such a diverse range of slides makes solving the AOI problem very difficult. Summary of the Invention
[0005] This disclosure addresses the problem of text recognition by providing a system and computer-implemented method for the accurate determination of AOI (Area of Interest) for various types of input slides. Slide thumbnail images (or thumbnails) can be designated as one of five different types, and a separation algorithm for AOI detection can be executed based on the slide type. The goal is to perform operations that can efficiently and accurately calculate the AOI region from the entire slide thumbnail image, as a single generic solution cannot account for all data variability. Each thumbnail can be assigned to one of five types based on tissue morphology, layout, tissue structure on the slide thumbnail, etc. Slide types include: ThinPrep (RTM) slides with a single disc structure (also known as "cell line slides") (or any other similar stained slides); tissue microarray (TMA) slides with a grid structure; control slides with blurred tissue (used to control HER2 tissue slides, with a specific number of circular cores, typically four); smear slides with tissue scattered throughout the slide; and default or generic slides. A control slide is a type of slide that includes both the tissue sample (which should actually be examined and analyzed) and control tissue samples (which are typically used to validate histological techniques and reagent reactions). It comprises multiple tissue regions (cores) of the tissue sample to be analyzed at defined locations along a straight line.
[0006] Any slide that cannot be visually categorized into any of the first four categories can be considered a generic slide. Slide types can be specified based on user input. Customized AOI detection operations are provided for each slide type. Furthermore, if the user inputs an incorrect slide type, the disclosed operations include detecting the incorrect input and performing appropriate methods. The result of each AOI detection operation provides a soft-weighted image as its output, with zero intensity values at pixels detected as not belonging to tissue and higher intensity values assigned to pixels detected as potentially belonging to tissue areas. The detected tissue areas are called regions of interest (AOIs). Optimizing AOI detection is a crucial step in the entire scanning process and enables subsequent steps (such as focus assignment based on soft-weighted values assigned to potential tissue areas in the output image). Higher-weighted AOI areas are more likely to be assigned focus points. The operation outperforms previous methods in terms of accuracy and recall scores, and this has been validated using scores calculated using ground-based real AOI data.
[0007] In one exemplary embodiment, this subject matter disclosure provides a system for detecting a region of interest (AOI) on a thumbnail image of a tissue slide, comprising a processor and a memory coupled to the processor for storing computer-executable instructions executed by the processor to perform operations including: receiving input including a thumbnail image and a thumbnail image type; and determining the region of interest (AOI) from the thumbnail image using one of a plurality of AOI detection methods according to the thumbnail image type, wherein when it is determined that the thumbnail image type is an incorrect input, the determination of the AOI uses another of the plurality of AOI detection methods.
[0008] In another exemplary embodiment, the subject matter disclosure provides a system for detecting a region of interest (AOI) on a thumbnail image of a tissue slide, comprising a processor and memory coupled to the processor for storing computer-executable instructions executed by the processor to perform operations including: determining the region of interest (AOI) from the thumbnail image using one of a variety of AOI detection methods according to a thumbnail image type; and outputting a soft-weighted image depicting the detected AOI, wherein the thumbnail image type represents one of a ThinPrep (RTM) slide, a tissue microarray slide, a control slide, a smear slide, or a general-purpose slide.
[0009] In yet another exemplary embodiment, this subject matter disclosure provides a system for detecting a region of interest (AOI) on a thumbnail image of a tissue slide, comprising a processor and a memory coupled to the processor for storing computer-executable instructions executed by the processor to perform operations including: receiving input comprising a thumbnail image and a thumbnail image type; determining the region of interest (AOI) from the thumbnail image using one of a variety of AOI detection methods according to the thumbnail image type, the variety of AOI detection methods including ThinPrep (RTM) method, tissue microarray method, control method, smear method, or general method; and using the general method when it is determined that the thumbnail image type is an incorrect input.
[0010] The terms AOI detection and tissue region detection will be used synonymously below.
[0011] In another aspect, the present invention relates to an image analysis system configured for detecting tissue regions in digital images of tissue slides. Tissue samples are mounted on slides. The image analysis system includes a processor and a storage medium. The storage medium includes multiple tissue detection routines specific to each slide type and general tissue detection routines. The image analysis system is configured to perform methods including:
[0012] - Select one of the tissue detection routines for a specific slide type;
[0013] - Before and / or simultaneously perform the tissue detection routine for the selected slide type, check whether the tissue detection routine for the selected slide type corresponds to the tissue slide type depicted in the digital image;
[0014] - If so, automatically execute the tissue detection routine for the selected specific slide type to detect tissue regions in the digital image;
[0015] - If not, the generic tissue detection routine is automatically executed to detect tissue regions in digital images. This generic tissue detection routine is also referred to herein as the "default" tissue detection routine, module, or algorithm. Attached Figure Description
[0016] Figure 1 An overview of a slide scanning process described in accordance with exemplary embodiments of the disclosure in this subject matter.
[0017] Figure 2A-2C A system and user interface for AOI detection, respectively, are depicted according to exemplary embodiments disclosed in this subject matter.
[0018] Figures 3A-3BA method for AOI detection on a ThinPrep (RTM) chip and the results of the method are described in accordance with exemplary embodiments of the disclosure of this subject matter.
[0019] Figures 4A-4B A method for AOI detection on a tissue microarray (TMA) chip and the results of the method are described in accordance with exemplary embodiments of the disclosure of this subject matter.
[0020] Figures 5A-5D A method for controlling AOI detection on a wafer, and the results of the method, are described in accordance with exemplary embodiments of the disclosure of this subject matter.
[0021] Figures 6A-6B A method for AOI detection on a smear slide and the results of the method are described in accordance with exemplary embodiments of the disclosure of this subject matter.
[0022] Figures 7A-7C A default method for AOI detection, based on exemplary embodiments of the disclosures in this subject matter, and the results of the method are described.
[0023] Figure 8 A histogram of a slide containing tissue samples is plotted, and this histogram is used to calculate multiple thresholds and corresponding intermediate masks. Detailed Implementation
[0024] This subject matter disclosure provides a system and computer-implemented method for the accurate determination of AOI (Automated Area of Interest) on various types of input slides. The AOI detection module described herein is crucial for optimized tissue detection on slides and is therefore an essential part of automated scanning systems. Once a digital scan is obtained from the entire slide, the user can read the scanned image to make a medical diagnosis (digital reading) or use image analysis algorithms to automatically analyze and score the image (image analysis reading). Therefore, the AOI module is an enabling module where accurate and timely scanning of slides is achieved, and thus it enables all subsequent digital pathology applications. Because scanning is performed at the beginning of the digital pathology workflow, stained tissue biopsy slides are subsequently acquired and scanned using the disclosed AOI detection. For each slide type (including general), a locally adaptive determination of one or more thresholds can be made for the tissue, and the results are incorporated into a "tissue probability" image. That is, the probability of a pixel being tissue is determined based on global and local constraints.
[0025] Generally, once a user selects a slide type, the disclosed method includes automatically performing internal validation to check if the user has entered an incorrect slide type (corresponding to an incorrectly selected tissue detection algorithm for that specific tissue slide type) and running the appropriate tissue detection algorithm. For slides labeled as control by the user, the operations disclosed herein determine whether the slide is an example of a control HER2 with four cores placed along a straight line, and if not, invoke a general tissue finder module for AOI extraction. For slides labeled as TMA by the user, the method (based on training TMA slides) determines whether the slide is an example of a single rectangular grid of cores, and if not, the disclosed method invokes a general tissue finder mode for AOI extraction. For ThinPrep(RTM) images, a single core (AOI) can be captured under the assumption that the radius of a single core is known and within a range of empirically determined variation; and a radially symmetric-based voting performed on the gradient magnitude of a downsampled version of the enhanced grayscale image is used to determine the center of the core and its radius. This method can be used to process other slides similar to ThinPrep(RTM) slide staining. For a description of radial symmetry voting typically used to detect blobular or circular objects, see Parvin, Bahram, et al., “Iterative voting for inference of structural saliency and characterization of subcellular events,” Image Processing, IEEE Transactions on 16.3(2007):615-623, the entire contents of which are incorporated herein by reference. For blurred control images, the operation automatically determines whether it is type “general” (blurred slice) or control HER2 (with 4 blurred cores) and can detect all 4 cores assuming their approximate size range is known and their centers are approximately located in a line. The control HER2 method uses radial symmetry followed by a difference of Gaussian (DoG) based filter to capture the possible locations of the core centers; then, empirically determined rules are used based on how the core center locations are aligned, the distance between the core centers, and the angle formed by the most likely line adapting to the core centers. For TMA images, where the slice cores are positioned in a rectangular grid, the method detects all relevant tissue regions assuming that the individual core sizes of the TMA cores and the distances from one core to its nearest core are very similar. To detect the core, constraints related to size and shape are used, and once the core is detected, empirically determined rules are used to determine between the true core and the anomalous core based on distance constraints.For smear slides, assuming the smear tissue can be distributed throughout the entire slide, all tissue regions across the slide can be reliably identified. Lower and upper thresholds have been calculated here based on the luminance and color images derived from the LUV color space representation of the thumbnail image, where the inverse of L is used as the luminance image and the square roots of the squared values of U and V are used as the color image. A hysteresis thresholding method is performed based on these thresholds and empirically determined area constraints. For the default / general image, a threshold range darker than the glass is used to appropriately detect tissue regions, and size and distance constraints are used to detect and discard certain smaller regions. Furthermore, once an AOI detection module is invoked and the user can see the generated AOI on the thumbnail image, the user can add additional focus points, in rare cases where a tissue has been missed by the algorithm. In rare scenarios where the calculated AOI is insufficient to capture all tissue regions, additional focus points can be placed using the GUI, if deemed necessary by the user, especially when the algorithm fails to do so for tissue regions.
[0026] The thumbnail image may be provided directly by an image scanner or may be calculated from the original image by an image analysis system. For example, the image scanner may include a low-resolution camera for acquiring the thumbnail image and a high-resolution camera for acquiring high-resolution images of the tissue slide. The thumbnail image may coarsely comprise 1000 x 3000 pixels, whereby one pixel in the thumbnail image may correspond to 25.4 μm of the tissue slide. The "image analysis system" may be, for example, a digital data processing device (e.g., a computer) including interfaces for receiving image data from a slide scanner, camera, network, and / or storage medium.
[0027] The embodiments described herein are merely exemplary, and while they disclose the best mode that enables those skilled in the art to reproduce the results depicted herein, readers of this patent application may be able to perform variations of the operations disclosed herein, and the claims should be construed as including all variations and equivalents of the disclosed operations and are not exclusively limited to the disclosed embodiments.
[0028] Figure 1This document provides an overview of a slide scanning process according to exemplary embodiments of the disclosure herein. Generally, one or more stained tissue biopsy slides may be placed in a scanning system for imaging operations. This system may include means 101 for capturing thumbnail images of the slides. The means may include a camera designed for the purpose. The image may be of any type, including RGB images. Processing operations 102 may be performed on the thumbnail image to generate a soft-weighted map including a depiction of a region of interest (AOI) for further analysis. The objectives of the embodiments described herein are to provide a system and computer-implemented method for optimally processing thumbnail images and accurately providing an AOI map indicating the probability of a specific or target tissue structure. The AOI map may be used to specify a focal point 103 based on this probability, and for each focal point, each tile, along with a set of z-layers corresponding to the tile location, is considered 104 around the focal point, followed by determining the optimal z-layer 105, 2D interpolation between the z-layers 106, and using the interpolated layers to create an image 107. The resulting image creation 107 enables other analytical steps, such as diagnosing prognosis, etc. The process may include additional or fewer steps depending on the imaging system used, and is intended to be interpreted merely as providing context, without any limitations intended to be imposed on any of the features described herein.
[0029] Figure 2AA system 200 for AOI detection is depicted according to exemplary embodiments of the disclosure of this subject matter. System 200 may include hardware and software for detecting AOI from one or more thumbnail images. For example, system 200 may include imaging components 201, such as a camera mounted on a slide level or slide tray, and / or a whole slide scanner with a microscope and camera. Imaging component 201 typically depends on the type of image generated. In this embodiment, imaging component 201 includes at least a camera for generating thumbnail images of the slide. The slide may include a stained sample that has been stained with the application of a staining analyte containing one or more different biomarkers associated with chromogenic staining for bright-field imaging. Tissue areas are stained, and AOI detection involves picking up the stained tissue areas. The stained region of interest may depict a desired tissue type (such as a tumor, a lymphatic region in an H&E slide, etc.) or a hotspot represented by a high biomarker in an IHC-stained slide (such as any tumor, immune or vascular marker tumor markers, immune markers, etc.). Images can be scanned at any zoom level, encoded in any format, or provided by imaging unit 201 to memory 210 for processing according to logic modules stored thereon. Memory 210 stores multiple processing modules or logic instructions executed by processor 220 coupled to computer 225. In addition to processor 220 and memory 210, computer 225 may also include user input and output devices (such as a keyboard, mouse, stylus, monitor / touchscreen), and networked elements. For example, execution of processing modules within memory 210 may be triggered by user input, as well as input provided via a network from a web server or database used for storage and later retrieved by computer 225. System 200 may include a whole-slide image viewer to facilitate viewing of the entire slide scan.
[0030] As described above, the module includes logic executed by processor 220. As used herein and throughout this disclosure, "logic" refers to any information having the form of instruction signals and / or data that can be applied to affect the operation of the processor. Software is one example of such logic. Examples of processors are computer processors (processing units), microprocessors, digital signal processors, controllers, and microcontrollers, etc. Logic may be formed by signals stored on a computer-readable medium, such as memory 210, which in an exemplary embodiment may be random access memory (RAM), read-only memory (ROM), erasable / electrically erasable programmable read-only memory (EPROM / EEPROM), flash memory, etc. Logic may also include digital and / or analog hardware circuitry, such as hardware circuitry including logical AND, OR, XOR, NAND, NOR, and other logical operations. Logic may be formed by a combination of software and hardware. On a network, logic may be programmed on a server or a complex of servers. A particular logical unit is not limited to a single logical location on the network. Furthermore, modules do not need to be executed in any particular order. Each module may invoke another module when needed.
[0031] The input detection module can receive thumbnail images and user input identifying slide types from the imaging unit 201. The thumbnail image (or "thumbnail") can be assigned a category based on the slide type input by the user, and separation algorithms 213-217 for AOI detection can be executed according to that category. Each thumbnail can be classified into one or more types based on tissue morphology, layout, and the structure of the tissue on the slide thumbnail. Slide types include, but are not limited to, ThinPrep (RTM) slides with a single disk structure; tissue microarray (TMA) slides with a grid structure; control slides with blurred tissue having a specific number (typically 4) of circular cores; smear slides with tissue scattered throughout the slide; and default or general slides (which include all slides not applicable to any of the previous categories). As an example, Figure 2B Describes the user interface for receiving slide types based on user input. Users can input the correct slide type based on selection list 219 when viewing slide thumbnails. For example, users can select from the following slide types: Grid (referring to TMA slides, where tissue cores exist within a single rectangular grid), Blurred (control slides are typically blurry, and due to control staining, they are never explicitly stained, making the thumbnail image appear blurry), Round (in ThinPrep (RTM) slides, this is a single circular core existing within a certain radius range in the slide thumbnail), Dispersed (for smears, tissue is typically scattered throughout the slide and not confined to a specific area like other slide types), and Default (also known as General).
[0032] Each of the modules 213-217 is customized for AOI detection for each of these slide types. For example, for ThinPrep (RTM) slides, tissue AOI is typically a single disc, and therefore a circular finder within a given radius range is an effective tool. Specific modules can be used to process other slides similar to ThinPrep (RTM) slides (such as those where the tissue sample is circular or near-circular in shape). Similarly, control HER2 slides typically have 4 cores with a known approximate radius range and a variation in intensity from darker to more blurred at the bottom to the top. In smear slides, the tissue is spread out and occupies a larger portion of the entire thumbnail image, and AOI is not concentrated in a smaller, more compact shape, as this is for other types of cases. In TMA or mesh slides, AOI is typically seen in an mxn grid, where grid elements can vary in size (individual TMA spots can be much smaller in size compared to spots seen in other non-TMA thumbnail images). Therefore, understanding the grid helps to constrain the expected size of the resulting spots, thus aiding in the removal of anomalous spots and shapes. The universal AOI detection module 217 provides omnidirectional operation to any slide that does not explicitly / visually belong to any of these four slide types. The tissue area detected by each detection module is called the region of interest (AOI).
[0033] The result of each AOI detection operation (which may be performed by one of the tissue detection routines for a general or tissue-specific slide type) is provided to a soft-weighted image output module 218 to generate a soft-weighted image output that has zero intensity values at pixels detected as not belonging to tissue and assigns higher intensity values to pixels detected as likely belonging to tissue regions. In the "soft-weighted image," each pixel has an assigned weight, also referred to herein as a "soft weight." As used herein, a "soft weight" is a weight whose value falls within a range of values that includes more than two distinct values (i.e., not just "zero" and "one"). For example, this range could cover all integers from 0 to 255 or could be a floating-point number between 0 and 1. This weight indicates the probability that a pixel is located in a tissue region rather than a non-tissue region (e.g., a glassy region). The larger the weight value, the higher the probability that the pixel is located in a tissue region.
[0034] The soft-weighted image output module 218 is invoked to compute a soft-weighted image for a given brightness and color image and a binary mask corresponding to the AOI (i.e., the unmasked area of the mask representing the tissue region). Generally, colored pixels are more likely to be tissue than glass; therefore, pixels with higher chromaticity (color) components are weighted more heavily. Darker pixels (based on brightness) are more likely to be tissue than blurrier pixels. This module takes an RGB image (e.g., an RGB thumbnail image) and a binary mask M as its input, wherein the binary mask M is provided by any of the AOI detection modules 213-217 (i.e., by one of the tissue detection routines for a general or tissue-specific slide type), and outputs a soft-weighted image SW with pixel values, for example, in [0, 255].
[0035] The image is first converted from the RGB color space to the L,UV color space (L = luminance, U and V are color channels). Let L' = max(L) – L, where max(L) is the maximum luminance value observed in the converted digital image, and L is the luminance value of the currently transformed pixel; (higher L' -> darker area).
[0036] The UV color space (also known as the chromaticity color space) or the UV color channel is calculated using UV = sqrt(U^2 + V^2).
[0037] For a luminance image, the image analysis method identifies all pixels where M > 0 (e.g., regions of the image will not be masked by the generated mask M), and calculates a lower threshold (L') in the L' domain for the identified pixels. low : 5% of the sorted L' values) and higher threshold (L' high (95% of the values).
[0038] Similarly, calculate the lower threshold (UV) in the UV domain. low : 5% of the sorted UV values) and higher thresholds (UV values) high (95% of the value). Instead of the stated 5%, values in the range of 2-7% can also be used. Instead of the stated 95%, values in the range of 90-98% can also be used.
[0039] Use L' low and L' high The weighted L image is calculated from the domain L'. The mapping from L' to the weighted L image during creation includes setting L'(x,y) <= L'. low ->WeightedL(x,y) = 0, and set L' low <L’(x,y)<L’ high ->WeightedL(x,y)=(L'(x,y)–L' low) / (L' high –L' low ), and set L'(x,y)>L' high ->WeightedL(x,y)=1. This means that if its original L' value is lower than L' low The threshold assigns a "0" value to pixels in the weighted L image if its original L' value is higher than L'. high The threshold assigns a "1" value to the pixels of the weighted L image, and in all other cases, a value (L'(x,y)–L'). low ) / (L' high –L' low The value (L'(x,y)–L') low ) / (L' high –L' low ) is the normalized inverse chromaticity value, which is normalized on the unmasked pixels.
[0040] Similarly, weighted UVs are calculated from UV images.
[0041] Image analysis systems use UV' low and UV' high The weighted UV image is calculated from the UV' domain. The mapping from UV' to weighted UV when creating the weighted UV image includes setting UV'(x,y) <= UV' low ->WeightedUV(x,y) = 0, and set UV' low <UV’(x,y)<UV’ high ->Weighted UV(x,y)=(UV'(x,y)–UV' low ) / (UV' high –UV' low ), and set UV'(x,y)>UV' high ->Weighted UV(x,y) = 1. This means that if its original UV' value is lower than UV'... low The threshold assigns a "0" value to pixels in the weighted UV image if its original UV' value is higher than UV''. high The threshold assigns a "1" value to pixels in the weighted UV image, and a value (UV'(x,y)–UV') in all other cases. low ) / (UV' high –UV' low Value (UV'(x,y)–UV') low ) / (UV' high –UV' low () is the normalized inverse luminance value, which is normalized on the unmasked pixels.
[0042] Then, a combined weighted image W can be created by merging the weighted L image and the weighted UV image. For example, the weight assigned to each pixel in the weighted image W will be calculated as the average of the weights of that pixel in the weighted L and weighted UV images.
[0043] The image analysis system then maps the weighted image W to a soft-weighted image (whose pixels have correspondingly assigned weights), also known as “soft weights” between 0 and 1.
[0044] The mapping from (weighted image W, mask image M) to SW includes the following steps based on s = (255–128) / (W). max –W min To calculate the scaling factor s, where W max The maximum observed weight in the weighted image W and W min The minimum observed weight in the weighted image W is used to selectively calculate the max and min values in W for pixels of the digital image, where M > 0, and M(x,y) = 0 -> SW(x,y) = 0, so M(x,y) > 0 -> value = (W(x,y) – Wmin) * s + 128; SW(x,y) = min(max(128,value), 255). Thus, 255 represents the maximum possible intensity value of the RGB image, and 128 represents half of that value. Pixels in the soft-weighted image SW have zero intensity in (non-organized) areas masked by the mask M and a "normalized" intensity value in the range of 128-255, thereby deriving the intensity value from the pixel's weight and indicating the probability that the pixel is an organized pixel.
[0045] Optimizing AOI detection is a crucial step in the entire scanning process and enables subsequent steps (such as soft-weighted focal point assignment based on potential tissue regions assigned to the output image). Higher-weighted AOI regions are more likely to contain focal points.
[0046] As used herein, a “focal point” is a point in a digital image that is automatically identified and / or selected by the user because it is predicted or assumed to indicate relevant biomedical information. In some embodiments, automatic or manual selection of a focal point can trigger a scanner to automatically retrieve additional high-resolution image data from the area surrounding the focal point.
[0047] However, if the obtained AOI is not accurate and the scanning technician or other operator of System 200 wishes to add more focal points to obtain a better scan, such options can be implemented via the user interface. For example, Figure 2C Describe an exemplary interface for manually adding additional focus points. Figure 2CThe interface can be presented on a computer coupled to, for example, a scanner used for scanning tissue slides. Furthermore, the operation outperforms previous methods in terms of accuracy and recall scores, and this has been validated using scores calculated using real-world AOI data. For example, AOI detection modules 213-217 were developed based on a training set of 510 thumbnail images. The AOI detection modules have used some empirically set parameters for features, such as minimum connecting part size, distance of the effective tissue area from the sliding edge, possible sliding position, and color threshold for distinguishing the effective tissue area from the dark pen color, wherein these parameters are set based on observed thumbnails, as further described herein.
[0048] Memory 210 also stores error correction module 212. Error correction module 212 can be executed when attempting to detect AOIs of the input slide type and encountering unexpected results. Error correction module 212 includes logic for determining that an incorrect slide type has been entered and selecting the appropriate AOI detection module regardless of the input. For example, TMA AOI detection module 216 may expect to see a single rectangular grid; however, it is possible that multiple grids exist in the same thumbnail. In such cases, TMA AOI detection may capture only a single grid pattern and discard the rest. In this case, general AOI detection module 217 may be more helpful in identifying all appropriate AOIs in the slide. General AOI detection module 217 can be executed if the user inputs a slide type as a “blurred” slide. For example, the user may instruct this via a graphical user interface (GUI) (e.g., by pressing a button or selecting a menu item indicating that the currently analyzed image is a “blurred” image). As used herein, a “blurred” image is an image that includes tissue areas that are not very dark compared to the brightness of the glass (i.e., have low contrast between the tissue and glass areas). In inverse grayscale images, the brightest parts are the densest tissue and the darkest parts are the least dense tissue, with completely dark (masked) parts corresponding to glass. However, the term "blurred" can also be applied to control images, i.e., control images with four cores. Therefore, when a "blurred" image type is received as input, the control AOI detection module 214 can be executed, and if none of the four cores are detected, the error correction module 212 can be invoked instead of the general AOI detection module 217. Similarly, for some TMA images, the TMA AOI detection module 216 may retrieve only a small portion of the expected AOI data, thereby triggering the determination that multiple grids may exist in the thumbnail, and thus the general AOI detection may be better than the TMA AOI detection. These and other alternative options for incorrect input slice types are described below with reference to Table 1.
[0049] Table 1: Operation of Error Correction Module 212 (Expected performance is indicated in parentheses).
[0050]
[0051] Each row in Table 1 indicates the slide type entered by the user, and each column indicates the actual slide type. Each cell indicates the expected performance level (poor, moderate, good, etc.) based on the operation. No default mode is provided for smears and ThinPrep (RTM) images; that is, there is no branch to a general mode for either of these cases. It should be noted that it is generally preferable to capture more false tissue areas rather than miss potential tissue areas, as missed detections are penalized more severely than false detections.
[0052] Figures 3A-3B This document describes a method for AOI detection on a ThinPrep (RTM) wafer according to exemplary embodiments of the disclosure of this subject matter, and the results of said method. (The last sentence appears to be incomplete and possibly refers to a different topic.) Figure 1 The method of Figure 3 can be performed using any combination of modules depicted in the subsystem or any other combination of subsystems and modules. The steps listed in the method do not need to be performed in the specific order shown. In ThinPrep (RTM) slices, the region of interest is typically circular and more blurred. Therefore, contrast stretching (S331) is performed as an image enhancement operation to make the individual core (region of interest) more prominent. The top few rows of pixels may be discarded to remove labels (S332). For example, empirical observation shows labels with text or blank spaces. In other words, the tissue is typically far below the top 300 rows. Next, to enable the calculation of both the possible center (S334) and radius of the circular core, the thumbnail image is downsampled (S333). For example, the thumbnail may be downsampled three times, each time by a factor of 2. The pyramid method may be used for faster downsampling. The downsampled image makes it easier / faster to calculate the possible location of the circular core (S334). Radial symmetry-based voting can be used to find the most likely core center C (S334). Radial symmetry operations can use a radius range of [42, 55] within the three downsampled images. This corresponds to the empirical observation that the radius of a single tissue core in the input image may lie between 336 and 440 pixels. The AOI region is then defined (S335) as a disk obtained at the center C (disk radius = the average radius of all points whose votes have accumulated at C). Finally, weights are assigned to the AOI pixels based on brightness and color characteristics (S336).
[0053] Therefore, the general assumption used by this method is that a single core exists in the image, the organized content in the first 300 pixels of the image is ignored or discarded, and the radius of the single core should be between 336 and 440 pixels (based on empirical / training observations and the actual size of the slide thumbnail image). Figure 3B Depicting Figure 3AThe result of the method is as follows: The image 301 of the ThinPrep (RTM) slide is delineated using rectangles 303 that mark the region of interest, including tissue spots or cores 304. A soft-weighted AOI image 302 (or mask) is then returned using this method.
[0054] Figures 4A-4B This document describes a method for AOI detection on a tissue microarray (TMA) substrate, based on exemplary embodiments of the disclosure of this subject matter, and the results of said method. (The last sentence appears to be incomplete and possibly refers to a different topic.) Figure 1 The method of Figure 4 can be performed by any combination of modules depicted in the subsystems, or any other combination of subsystems and modules. The steps listed in the method do not need to be performed in the specific order shown. Overall, the method is similar to... Figure 7A The “general” method described herein. For the various blobs present in the thumbnail, it is determined whether different blobs constitute a mesh. Size constraints are imposed to discard very small or very large blobs to prevent them from becoming potential TMA blobs, and shape constraints are used to discard highly non-circular shapes. There may be certain cases where the TMA method branches to the general method, such as when only a small portion of the possible AOI region, such as that obtained from a soft-weighted foreground image, is captured in the TMA mesh.
[0055] More specifically, this method begins by obtaining a grayscale image from an RGB thumbnail image (S441) and discards regions of the image that are known to be irrelevant based on training data (S442). For example, black support labels at the bottom of the thumbnail image can be discarded. Sliding edges can be detected, and regions outside the sliding edges can be discarded. Dark areas can be detected, i.e., those with gray pixels and an intensity <40 units (assuming an 8-bit thumbnail image), and these dark areas can be discarded to prevent them from becoming part of the AOI. The typical width of the sliding cover is about 24 mm, but significantly larger or smaller slides can also be used. However, embodiments of the invention assume that there is no more than one slide per slide.
[0056] In addition to these areas, edges, and pen marks, dark symbols can also be detected and discarded. Control windows are also detected if they exist on the slide. As used herein, a “control window” is an area on a tissue slide that includes additional tissue sections with known characteristics that have undergone the same staining and / or cleaning protocols when the tissue sample is detected and analyzed during image analysis. Tissue within the control window typically serves as a reference image for evaluating whether tissue slide cleaning, staining, and processing have been performed correctly. If a control window is found, the area corresponding to the control window boundary can be discarded. For cleaner AOI output, areas within a 100-pixel boundary line can also be discarded to help avoid parasitic AOI areas closer to the edges, based on the assumption that the TMA core will not be so close to the thumbnail boundary line.
[0057] Then, a histogram is calculated (S443) for all remaining valid image regions in the image. Based on the image histogram, a threshold range can be calculated (S444) to perform an adaptive thresholding method. This operation includes determining the pure glass region (i.e., where the image histogram peak is at the pixel value corresponding to glass) and calculating a threshold range R smaller than the glass intensity value. A Based on this image histogram, it is further determined whether the slice is sufficiently blurred (S445), and if so, a threshold range R is automatically created. B In this range R B Compared to range R A Closer to the glass strength value. Create a (S447) mask based on the threshold range.
[0058] According to some embodiments, the threshold range can be calculated as follows:
[0059] Set the pixel intensity value that causes the histogram peak to maxHistLocation (this can correspond to glass).
[0060] Sets the maximum intensity value in the grayscale image to max_grayscale_index. This value represents the "glass_right_cutoff" value.
[0061] According to some embodiments, the following operation is used to select a range of thresholds and store them in a vector with elements labeled, for example, different_cutoffs:
[0062] Set the right gap to max_grayscale_index – maxHistLocation.
[0063] Make glass_left_cutoff=maxHistLocation–rightgap.
[0064] Let interval_range_gap=max(1,round(rightgap / 5)).
[0065] Make min_rightgap=-rightgap.
[0066] Set max_rightgap = min(rightgap – 1, round(rightgap * 0.75). The value 0.75 is a predefined gap_cutoff_fraction value that is usually in the range of 0.5-0.95 (preferably in the range of 0.7-0.8).
[0067] Let different_cutoffs[0]=glass_left_cutoff–2*rightgap.
[0068] Let number_interval_terms=(max_rightgap-min_rightgap) / interval_range_gap.
[0069] The size of the vector different_cutoffs is (number_interval_terms+2).
[0070] For i = 0: number_interval_terms – 1:
[0071] different_cutoffs[i+1]=(glass_left_cutoff–rightgap)+interval_range_gap*idifferent_cutoffs[number_interval_terms+1]=glass_left_cutoff+max_rightgap.
[0072] Figure 8 Histograms (800) depict tissue samples including threshold ranges and corresponding intermediate masks applicable according to embodiments of the invention. Typically, the intensity peak mhl caused by the glass region is around 245 and usually falls within the range of 235-250, while mgs typically falls within the range 0-15 units higher than mhl but is limited to a maximum of 255. The basic idea is that in a typical image, pixels corresponding to glass form fairly symmetrical peaks, only 3-10 units wide (outside 255) on each side. The “right gap” (mgs-mhl) is the observed width of this peak on the right side. “Glass_left_cutoff” is an estimate of where the glass end is to the left of this peak. Substances with lower peak intensities are typically tissue. However, particularly for blurred tissue slides, it has been observed that a single threshold is not always sufficient to correctly distinguish tissue from glass. Therefore, embodiments of the invention use... Figure 8 The values identified in the histogram depicted are used to automatically extract a threshold range rts ranging from the pixel intensity most likely to be tissue to the pixel intensity most likely to be glass. In the depicted example, the threshold range ranges from dco[0] (i.e., from the values located at the two right gaps from glco to the left) to values close to mhl. For example, the value located close to mhl can be calculated as:
[0073] Dco[max-i] = different_cutoffs[number_interval_terms+1] = glass_left_cutoff + max_rightgap. It is also possible to select only mhl as the maximum threshold within the threshold range rts. According to an embodiment, the number of thresholds (number_interval_terms) included in the threshold range rts can be predefined or calculated as the derivative across the range of values in rts.
[0074] Each threshold in the RTS threshold range is applied to the digital image or its derivative (e.g., a grayscale image) to compute the corresponding mask. The generated masks are iteratively combined in pairs to generate intermediate masks and ultimately a single mask M. The size of the pixel spot in each mask is then compared to an empirically determined minimum spot area, and this size is examined to determine the presence of other spots in the neighboring region of that spot in the corresponding other masks.
[0075] For example, spots in one of the generated masks, the first mask1 and the second mask2, can be identified through connection component analysis. If an unmasked pixel exists in mask1 within, for example, 50 pixels of the i-th spot in mask2 (another distance threshold, typically less than 100 pixels, can also be used) and if its total size is >= (minimum_area / 2), then that spot in mask2 is considered a "real spot," and its unmasked pixels are preserved. If no unmasked pixel exists in mask1 within, for example, the aforementioned 50 pixels of the pixel in the i-th spot mask (i.e., the i-th spot in mask2 is relatively isolated), then the pixels of that spot in mask2 are only preserved as unmasked pixels if the spot size is greater than, for example, 1.5 * minimum_area. The intermediate mask generated by merging mask1 and mask2 includes the pixels of the blob in mask2 as unmasked pixels, provided that the blob is either large enough to exceed a first empirically derived threshold (>1.5*minimum_area) or at least exceeds a second threshold (>0.5*minimum_area) smaller than the first threshold, and has at least some unmasked pixels in mask1 within its neighborhood (within 50 pixels). The value of minimum_area has been empirically determined to be, for example, 150 pixels.
[0076] A soft-weighted image SW (with values in the range [0,1]) is created using a grayscale image and different_cutoffs. A binary mask (BWeff) is created (S447), which works at all pixels if the soft-weighted image is >0. For binary mask creation (S447), a binary mask image is created for all thresholds in different_cutoffs (the corresponding pixels in the mask image are the pixels with grayscale values < thresholds); a vector of the binary mask image is produced. This vector of the mask image is called the mask. Assuming there are N elements in the mask, another vector of the mask image (called FinalMask) is created, which has (N-1) elements. The function that combines masks[i] and masks[i+1] to generate FinalMask[i] can be called CombineTwoMasks, and it requires the minimum size of the blob as input. A binary OR operation of all binary images in FinalMask produces BWeff. The function that takes grayscale, different_cutoffs, and minimum_blob_size as input and returns SW and BWeff as output is called CreateWeightedImageMultipleCutoffs.
[0077] Functions that combine two masks (called CombineTwoMasks) can include the following operations:
[0078] Let FinalMask=CombineTwoMasks(mask1,mask2,minimum_area).
[0079] Initialization: FinalMask is made to the same size as mask1 and mask2 and all pixels in FinalMask are set to 0.
[0080] Perform the connection component (CC) on mask2; make the number of CCs M.
[0081] Calculate the distance transformation on mask1; make the distance transformation matrix Dist.
[0082] For i = 1:M (a loop over all M connected parts), consider the i-th CC in mask2 and see if there are any active pixels in mask1 within 50 pixels of the i-th CC in mask2 (doed using knowledge of Dist); if so, then only consider whether the area's total size is >= (minimum_area / 2), and set all pixels in the FinalMask corresponding to the i-th CC mask to active. If there are no active pixels in mask1 within 50 pixels of the pixels in the i-th CC mask (i.e., the i-th CC in mask2 is relatively isolated), then only consider the CC if it is large enough (i.e., the area of the mask based on the i-th CC > (1.5 * minimum_area)) and set all pixels in the FinalMask corresponding to the i-th CC mask to active. Therefore, in summary, the role of CombineTwoMasks is to consider only those spots in mask2 that are either large enough (>1.5*minimum_area) or somewhat large (>0.5*minimum_area) and have at least some functional pixels in mask1 that are sufficiently close to them (within 50 pixels). The value of minimum_area has been empirically determined to be 150 pixels.
[0083] Once the binary mask BWeff is calculated, soft weights are assigned to pixels based on grayscale intensity values. The fundamental concept is that darker pixels (lower grayscale intensity values) are more likely to be organized and therefore will receive higher weights. Soft weights also enable the subsequent determination of valid blobs in subsequent operations (S449), as blobs with soft weights exceeding a certain threshold may be retained and others discarded (S450). For example, blobs that are large enough or have a strong “darker” foreground with a sufficient number of pixels may be valid, and other blobs are discarded. To map the soft-weighted image, the input is a grayscale image, the valid binary mask image BWeff, a vector of thresholds called different_cutoffs, and the output is a soft-weighted foreground image SW. To determine the active pixels in the valid binary mask image BWeff (the total number of elements in the different_cutoffs vector is N):
[0084] Let norm_value=(grayscale value–different_cutoffs[0]) / (different_cutoffs[N-1]–different_cutoffs[0]).
[0085] If norm_value > 1, then norm_value = 1; otherwise if norm_value < 0, then norm_value = 0.
[0086] The corresponding floating-point value in SW = 1 – norm_value.
[0087] If the corresponding gray-level value < different_cutoffs[0], then the corresponding floating-point value in SW = 1 (if the pixel is dark enough, the assigned soft weight value = 1).
[0088] If the corresponding gray-level value > different_cutoffs[N – 1], then the corresponding floating-point value in SW = 0 (if the pixel is light enough, the assigned soft weight value = 0).
[0089] According to the embodiment, one or more of the general tissue region detection routines and the tissue region detection routines for specific tissue slide types are further configured to calculate a (soft) weighted image based on the digital image of the tissue sample, thereby obtaining a gray-level image, a mask image (BWeff), and a soft-weighted foreground image (SW) of the tissue sample. In SW, we use 21 non-uniformly spaced bins for the histogram to consider all pixels between 0.001 and 1. We determine whether the slide is sufficiently blurred based on the distribution of data in the bins. If the image is sufficiently blurred, the distribution of data in the higher-end histogram bins will be sufficiently low; similarly, the distribution of data in the lower-end histogram bins will be sufficiently high.
[0090] The above-mentioned steps of histogram generation, threshold extraction, mask generation, mask fusion, and / or (soft) weight calculation can be implemented by a subroutine of one of the general and / or specific tissue slide type tissue detection routines.
[0091] Once the binary mask (S447) is generated according to the above steps, the connected component detection (S448) can be performed to see whether each connected component is a possible TMA core. The detected connected components are subject to the following rules regarding effective size, shape, and distance (S449):
[0092] The width of the connected component (CC) < (cols / 6): (the width of the thumbnail image with the label below = cols)
[0093] The height of the CC < (rows / 8): (the height of the thumbnail image with the label below = rows)
[0094] The area of the CC > 55 pixels (small waste spots may be discarded)
[0095] The eccentricity of the CC mask is >0.62 (the TMA core is assumed to be circular).
[0096] After identifying potential TMA cores, determine the size distribution of CCs to decide which TMA core is sufficiently appropriate: make MedianArea = the median of all CCs retained after the previous steps of identifying valid CCs.
[0097] Retain areas with a size greater than or equal to MedianArea / 2 and areas less than or equal to 2 * MedianArea (CCs).
[0098] Then, examine the inter-spot distances for all valid CCs and use the following rules to determine whether a spot is retained or discarded. The terminology used in the rules may include:
[0099] NN_dist_values[i] = the distance between the center of the i-th spot and the center of the nearest TMA spot.
[0100] radius_current_blob_plus_NN_blob[i] = radius of the i-th TMA blob + radius of the blob closest to the i-th TMA blob.
[0101] The median of DistMedian = (NN_dist_values).
[0102] MADvalue = (NN_dist_values) is the mean absolute deviation.
[0103] Sdeff = max(50, 2*MAD).
[0104] The rule governing the spots can be expressed as: NN_dist_values[i] <= 1.3 * (DistMedian + max(Sdeff, radius_current_blob_plus_NN_blob[i]) => for the i-th spot to be a valid TMA spot. When considering the TMA core, the distance between any two closest cores is generally similar in the TMA, so all these assumptions are used in the distance-based rule.
[0105] therefore, Figure 4AThe TMA algorithm in [0] is based on the following assumptions: the effective TMA cores are small enough, the effective TMA cores are larger than a certain threshold so as to discard spots based on small waste / dirt / dust, each effective TMA core is sufficiently circular and thus the eccentricity should be closer to 1, and all effective TMA cores should be close enough and the distance between cores and the closest core should be similar enough so that when this distance constraint is violated, these cores are discarded from the AOI. It can be further assumed that there is no tissue area outside the detected slide. For the module detecting dark areas in the detected image, it is assumed that pen marks and symbols are darker than the tissue area. Thus, a gray-level value of 40 can be used to discard dark areas. Figure 4B shows Figure 4A the result of the method in [4]. The image 401 of the TMA slide is demarcated by a rectangle 403 marking the region of interest including the tissue sample array 404. The soft-weighted AOI map 402 (or mask) is returned by this method.
[0106] Figures 5A-5B Depicts a method for controlling AOI detection on a slide and the result of the method according to an exemplary embodiment of the present disclosure. It can be performed by Figure 1 any combination of the modules depicted in the subsystem of
[10] or any other combination of the subsystem and modules. The steps listed in the method do not need to be performed in the specific order shown. This method is generally used to detect 4 tissue section cores (of the tested tissue, rather than the control tissue in the control window of the slide) on the control slide using a radially symmetric operation on an enhanced version of the gray-level image, where the excitation is that votes accumulate at the centers of the 4 cores. A multi-level difference of Gaussians (DOG) operation is used on the voting image to detect the core centers. Radially symmetric voting helps the accumulation of votes, and when a multi-level DoG is used on the voting image, it helps to better determine the location of the core centers. Then, segmentation techniques can be applied to the area near the discovered core centers because the spots are darker than their closest background and thus thresholding helps to segment out the cores. Although this method is for a "control" slide, it can work on a general fuzzy slide and also work on both control HER2 (4-core) images. If the thumbnail does not represent a 4-core control HER2 image, this method can automatically branch out from the general method.
[0107] Specifically, the method obtains a grayscale image of the RGB thumbnail (S551), and performs a radially symmetric vote (S552) and DOG (S553) on this image to obtain the cores in the image (S554). If the image appears to be a general tissue slide, without a specific arrangement, and is non-brown in color (a control slide is a slide where the control color is presented and the control stained tissue is not clearly stained), then it is determined that it is not a 4-core control slide, and the method branches to general (S555). This determination is completed by calculating the median and maximum votes (S552) based on the radially symmetric vote (S552) and DOG (S553) for all detected cores, and this determination is based on the assumption that the 4-core slide has higher values for the (median, maximum) of (votes, DOG). The grayscale image can be downsampled twice to accelerate the voting process. For voting, a radius range of [13, 22] pixels can be used for the twice-downsampled image. The resulting voting matrix is subjected to DOG (S553), and the top 10 DOG peaks greater than the threshold are used as the expected core center positions. For example, based on all detected core centers, the mean and maximum of the DOG for all cores and the mean and maximum of the votes for all cores can be calculated, where mean_DOG_all_cells_cutoff is set to 1.71, mean_DOG_all_cells_cutoff is set to 2, mean_votes_all_cells_cutoff is set to 1.6, mean_votes_all_cells_cutoff is set to 1.85, and then it is determined whether the mean / maximum of DOG / votes exceeds these cutoffs. If, outside of 4 conditions, this condition satisfies 2 or more cases, then it may be decided that the image represents a control slide with 4 cores.
[0108] Given the expected core centers, a segmentation operation (S556) based on the Otsu threshold method is performed to segment out the foreground spots, which is based on the fact that generally the cores are darker than their closest background, which makes it possible to extract the cores based on intensity. Once the segmentation (S556) is performed, spots that do not match the radius range [13, 22] (S557) can be discarded (S558). The expected spots are circular, so spots with a calculated circularity feature < min_circularity (0.21) are discarded, where circularity is defined as (4 * PI * area / (perimeter * perimeter)). For non-circular elongated shapes, the circularity is low and for approximately circular shapes, the circularity is close to 1.
[0109] Regarding all remaining cores within the effective radius range [13,22], any four spots placed along a straight line, along with the nearest core distance assumed to be within a range of [38,68] pixels (i.e., the range between the two closest core centers), are detected (S559). To determine whether some spots are along the same straight line (S559), a spot can be considered as part of a line whose distance from the spot center is less than a certain distance constraint (23 pixels). If fewer than four cores are found after using these size and distance constraints (S560), the detection of additional cores is attempted (S561) to ultimately find a total of four cores, as per the previous method. Figures 5C-5D Further description.
[0110] This method typically assumes that if a non-4-core image does not have 5 or more smaller cores, then 4-core and non-4-core control images can be appropriately separated. Some 5-core images can be analyzed as generic images. The radius of each core is assumed to be within the range of [52, 88] pixels in the thumbnail image. To segment each core, it is assumed to be darker than its nearest background. This makes it possible to perform intensity-based thresholding on grayscale versions of the thumbnail image. The distance between the two nearest cores (for 4-core control HER2 thumbnail images) is assumed to be within the range of [152, 272] pixels. Figure 5B The results of this method are shown. A thumbnail image 501 of the control slide is defined using a rectangle 503 that marks the tissue slide area including the core 504. A soft-weighted AOI image 502 (or mask) is then returned using this method. Figures 5C-5D The apparatus for detecting more cores (S561) is described. For example, for each detected core 504, the expected core is assumed to be 50 pixels away from arrow 506. A better configuration is one with a larger number of effective spots (i.e., a configuration that matches the size and roundness constraints), and if both configurations have the same number of effective spots, the averaged roundness features on all effective cores are compared with the settling coefficient, and the best of the three options is selected. Figure 5C Only two spots were found, so the method selects from any of the three alternative configurations. Figure 5D Only 3 spots were found, so the method selects the best of the two alternative configurations.
[0111] Figures 6A-6B This invention describes an exemplary method for AOI detection on smear slides (e.g., slides containing blood samples) according to exemplary embodiments disclosed herein. It can be provided by... Figure 1The method of Figure 6 can be performed by any combination of modules depicted in the subsystems or any other combination of subsystems and modules. The steps listed in the method do not need to be performed in the specific order shown. Generally, the smear method considers two image channels calculated from an RGB image, an L and a sqrt(U^2+V^2) LUV image, where L is the luminance channel and U and V are the chrominance channels. The median and standard deviation are calculated for both the L and sqrt(U^2+V^2) images. Based on lower and higher thresholds calculated with respect to the median and standard deviation, a hysteresis thresholding operation is performed to return separate AOI masks for the L and sqrt(U^2+V^2) domains, and these masks are then combined. Post-processing operations can be performed to remove parasitic areas, especially near the bottom of the slide thumbnail.
[0112] Specifically, a grayscale image (S661) can be obtained from an RGB thumbnail image. A dark mask corresponding to pixels with grayscale values <40 is generated (S662). The RGB image is then converted (S663) into L and UV sub-channels (L = luminance, U, V: chrominance). The L channel is also called the luminance image. The U and V channels are the color channels. A threshold is calculated (S664) for both sub-channels (i.e., for the luminance image and for the chrominance image).
[0113] For the L channel, the inverse channel L' is calculated using L' = max(L) – L, where max(L) is related to the maximum luminance value observed in the L channel image and L is the luminance value of the pixel. Then, the median (MedianL) and mean absolute deviation (MADL) are calculated for the L' channel.
[0114] The UV of the chromaticity image is calculated using UV = sqrt(U^2 + V^2).
[0115] For channel L', according to thL low =MedianL + MADL to calculate the lower threshold thL low And according to thL high =MedianL + 2 * MADL to calculate the higher threshold thL high .
[0116] A similar method was used to calculate the threshold for the UV channel.
[0117] The UV of the chromaticity image is calculated using UV = sqrt(U^2 + V^2).
[0118] The inverse channel UV' is calculated using UV' = max(UV) – UV, where max(UV) is related to the maximum luminance value observed in the UV image and UV is the chromaticity value of the pixel in the image. Then, the median (MedianUV) and mean absolute deviation (MADUV) are calculated for the UV' channel.
[0119] For the chroma UV channel, according to thUV low =MedianUV + MADUV to calculate the lower threshold thUV low And according to thUV high =MedianUV + 2*MADUV to calculate the higher threshold thUV high .
[0120] Then, the hysteresis thresholding method (S665) is performed on both channels (L', UV'). On the L' channel, thL is used. low and thL high Furthermore, a hysteresis thresholding method with an area constraint of, for example, 150 pixels was used to obtain the binary mask (M) in the L' domain. L’ ).
[0121] Hysteresis thresholding on the UV channel uses, for example, an area constraint of 150 pixels to obtain a binary mask (M) in the UV domain. U,V ).
[0122] In hysteresis thresholding, "using area constraints" means that any pixel spot detected by hysteresis thresholding (as a group of neighboring pixels with values satisfying both the lower and upper thresholds) is retained as an "unmasked" pixel spot only in the generated mask if the spot is at least as large as the area constraint. The size of the area constraint is typically in the range of 130-170 pixels, for example, 150 pixels.
[0123] Through M L’ and M UV The binary OR combination (S666) of (M(x,y)=1 if ML'(x,y)=1 or MUV(x,y)=1) yields the final valid binary mask M.
[0124] Performing hysteresis thresholding on two different masks generated from luminance and chromaticity images, respectively, and then recombining the threshold masks, can increase the accuracy of tissue and glass region detection, particularly for smear slides. This is because tissue regions on a slide typically consist of cells scattered throughout the slide. These cells usually have low contrast relative to the slide. However, the color components represented in the chromaticity image can serve as a clear indicator of whether a particular slide area is covered by tissue.
[0125] The resulting combination can be soft-weighted as described herein to generate the final mask. Furthermore, analysis (S667) is performed on the regions between the support labels at the bottom of the thumbnail image to determine whether any tissue regions between support labels adjacent to the detected tissue AOI regions need to be included in the mask.
[0126] This method typically assumes that any sliding cover detection would result in the discarding of smear tissue, so tissue areas outside the sliding cover are automatically discarded. The benefit here is that if the tissue does not extend to the slide / edge, the sliding cover edge is ignored. Furthermore, because smear tissue can be located between support labels, it is assumed to be genuine tissue, especially if it is connected to tissue already detected in the area outside the support labels. For this method, after AOI detection, smaller connected components are not discarded based on the assumption that the smear tissue may not consist of consolidated spots but may exist as tissue clumps near protruding smear portions (which can be small enough in the region). Because losing tissue areas is riskier than capturing parasitic areas, a conservative approach involves retaining smaller connected components in the AOI. Figure 6B The results of this method are shown. A thumbnail image 601 of the smear slide is delineated using a rectangle 603 that marks the region of interest, including the tissue smear 604. This method is then used to return a soft-weighted AOI image 602 (or mask).
[0127] Figures 7A-7C This describes AOI detection of a wafer for general or error identification, based on exemplary embodiments of the disclosure of this subject matter. It can be provided by... Figure 1 The method of Figure 7 can be performed using any combination of modules depicted in the subsystems or any other combination of subsystems and modules. The steps listed in the method do not need to be performed in the specific order shown. Typically, the method uses a variety of thresholds to propose soft-weighted images with a higher probability of finding tissue regions, and discards blurred image regions from blurry artifacts based on the assumption that blurred regions closer to darker tissue regions are more likely to be tissue regions than blurred regions farther away from true tissue regions.
[0128] The method begins by obtaining a grayscale image from an RGB thumbnail image (S771) and discarding areas known to be of no interest (S772). For example, areas such as the black support label at the bottom of the thumbnail image and areas outside the detection edge of the slider are discarded. Dark areas with grayscale values <40 and any areas outside the edges or boundaries of the control window (if any area is detected) are discarded to prevent them from becoming part of the AOI. An image histogram is calculated (S773) for all image areas that were not previously discarded. Based on this image histogram, a peak value for glass is obtained and used to calculate the threshold range for which an adaptive thresholding method is applied to detect the expected tissue area. These operations S773-S776 are similar to those described above. Figure 4AThe TMA method described is performed as follows. However, in this case, if the image is blurred, the most prominent blob is retained (S777), i.e., the sum of the soft weights of all pixels belonging to that blob is at its maximum. Further, smaller blobs closer to the most prominent blob are retained when the soft weights are sufficiently high and sufficiently close. Connecting parts are detected (S778), and size, distance, and soft weight constraints are applied (S779-S780) to retain and discard blobs. The result is an AOI mask. Further, when an appropriate tissue area exists within the tissue area inside the touchable zeroing section, pixels in a 100-pixel wide boundary area are considered part of the AOI mask. Small blobs are discarded (S780) based on size, distance constraints, proximity to the dark support label at the bottom, soft weight constraints, etc.
[0129] This default or general method assumes that there are no tissue areas outside the detected slider, because it risks zeroing out all areas outside the slider if the slider is detected incorrectly (e.g., when a line-like pattern exists within a tissue area). Furthermore, for the module detecting dark areas in the image, the intuition is that pen marks and symbols are darker than tissue areas. This method further assumes that dark areas are discarded at a gray level of 40. A blob is considered small if it is <150 pixels, and is discarded if it is not within a certain distance (<150 pixels) from any other obviously non-small blob. Pixels with a soft weight >0.05 and within 75 pixels of the most salient blob are considered part of the AOI.
[0130] As described above, these AOI detection algorithms have been trained using ground-based data estimated from 510 training slices contained at least 30 slices from each of all five categories. The algorithms were further tested on a test set of 297 slice thumbnails, different from the training set. Based on ground-based AOI data, precision and recall scores were calculated, and it was observed that for all five slice types, the precision-recall scores obtained using these methods significantly outperformed prior art methods. These improvements may contribute to several features disclosed herein, including but not limited to adaptive thresholding (which uses a series of thresholds in relation to a single statistically found threshold typically done in the prior art), or lower and higher thresholds as in the case of hysteresis thresholding, performing connective component detection on binary masks to discard smaller blobs, and using a series of thresholds to remove dependence on a single selected threshold, and considering the distance between smaller (still significant) blobs and larger blobs (which are significant because small changes in intensity can separate larger blobs and smaller, still important blobs can be discarded by prior art methods). The operations disclosed in this paper can be ported to a hardware graphics processing unit (GPU) to enable multi-threaded parallel implementation.
[0131] In another aspect, the present invention relates to an image analysis system configured for detecting tissue regions in digital images of tissue slides. Tissue samples are mounted on slides. The image analysis system includes a processor and a storage medium. The storage medium includes multiple tissue detection routines specific to each slide type and general tissue detection routines. The image analysis system is configured to perform a method comprising:
[0132] - Select one of the tissue detection routines for a specific slide type;
[0133] - Before and / or simultaneously perform the tissue detection routine for the selected slide type, check whether the tissue detection routine for the selected slide type corresponds to the tissue slide type depicted in the digital image;
[0134] - If so, automatically execute the tissue detection routine for the selected specific slide type to detect tissue regions in the digital image;
[0135] - If not, automatically execute the general tissue detection routine to detect tissue regions in digital images.
[0136] The aforementioned features may be advantageous because they provide a fully automated image analysis system capable of identifying tissues in multiple different slide types. In the event that an incorrect slide type and corresponding tissue analysis method are selected by the user or automatically, the system can automatically determine the incorrect selection and execute a default tissue detection routine that has been shown to be effective for all or almost all types of tissue slides. The embodiments described below can be freely combined with any of the embodiments described above.
[0137] According to an embodiment, the digital image is a thumbnail image. This can be advantageous because thumbnails have lower resolution and can therefore be processed more efficiently compared to "full-size" tissue slide images.
[0138] According to an embodiment, the image analysis system includes a user interface. The selection of a tissue detection routine for a specific slide type includes receiving a user selection of a tissue slide type via the user interface; and selecting one of the tissue detection routines for the specific slide type specified by the user-selected tissue slide type.
[0139] According to an embodiment, one or more of the general tissue detection routine and the tissue detection routine (TMA routine) for a specific slide type include subroutines for generating a binary mask M that selectively masks non-tissue regions. The subroutines include:
[0140] - Calculate the histogram of the grayscale version of the digital image;
[0141] - Extract multiple intensity thresholds from the histogram;
[0142] - Apply multiple intensity thresholds to the digital image to generate multiple intermediate masks from the digital image;
[0143] - Generate a binary mask M by combining all intermediate masks.
[0144] According to an embodiment, extracting multiple intensity thresholds from a histogram includes:
[0145] - Identify the max-grayscale-index(mgs), which is the maximum grayscale intensity value observed in the histogram;
[0146] - Identify the max-histogram-location(mhl), which is the gray level intensity value with the highest frequency of occurrence in the histogram;
[0147] - The right gap is calculated by subtracting the maximum-histogram-location (mhl) from the max-grayscale-index (mgs);
[0148] - Calculate glass-left-cutoff based on glass-left-cutoff = max-histogram-location(mhl)-rightgap;
[0149] - Perform extraction of multiple thresholds such that the lowest of the thresholds (dco[0]) is equal to or greater than the intensity value of glass_left_cutoff–rightgap and the highest of the thresholds is equal to or less than glass_left_cutoff+rightgap. For example, the lowest threshold could be equal to or greater than the intensity value of glass_left_cutoff–rightgap+1 and the highest of the thresholds could be equal to or less than glass_left_cutoff+rightgap-1.
[0150] The aforementioned feature can be advantageous because the threshold is histogram-based and therefore dynamic adaptation can depend on the blurring of tissue and non-tissue areas of the slide due to many different factors, such as light source, staining intensity, tissue distribution, camera settings, etc. Thus, better comparability of images from different tissue slides and better overall accuracy in tissue detection can be achieved.
[0151] According to an embodiment, extracting multiple intensity thresholds from a histogram includes:
[0152] - Calculate interval-range-gap(irg) based on irg = max(1, round(rightgap / constant)), where constant is a predefined value between 2 and 20 (preferably between 3 and 8);
[0153] - The value of max_rightgap is calculated based on max_rightgap = min(rightgap–1, round(rightgap*gap_cutoff_fraction), where gap_cutoff_fraction is a predefined value in the range of 0.5-0.95 (preferably 0.7-0.8);
[0154] - Calculate number_interval_terms based on number_interval_terms = ((max_rightgap + rightgap) / interval_range_gap); and
[0155] - Create multiple thresholds such that their number is the same as the calculated number_interval_terms.
[0156] This feature may be advantageous because the number of generated thresholds is histogram-based and therefore dynamically adaptable to the blurring and intensity distribution of tissue and non-tissue regions on the slide, influenced by many different factors such as light source, staining intensity, tissue distribution, camera settings, etc. It has been observed that when the glass-related intensity peaks are very narrow, a smaller number of thresholds may be sufficient to accurately distinguish tissue regions from glass. When the glass-related intensity peaks are very wide and the corresponding right gap is large, it may be preferable to calculate a larger number of thresholds and the corresponding mask.
[0157] According to an embodiment, applying multiple intensity thresholds to a digital image to generate multiple intermediate masks includes:
[0158] a) Create a first intermediate mask by masking all pixels in a digital image whose grayscale value is lower than a first threshold among a plurality of thresholds;
[0159] b) A second intermediate mask is created by masking all pixels in a digital image whose grayscale value is lower than a second threshold among a plurality of thresholds; the second threshold may be, for example, the next highest threshold among a plurality of thresholds.
[0160] c) Perform connectivity analysis on the unmasked pixels of the first mask to identify the first spot;
[0161] d) Perform connectivity analysis on the unmasked pixels of the second mask to identify the second spot;
[0162] e) Merging the first and second masks into a merged mask, the merging including:
[0163] - Mask all first spots whose size is less than the absolute minimum area and mask all first spots whose size is greater than the absolute minimum area but less than the conditional minimum area, such that there are no first spots in a predefined neighboring region around the first spots;
[0164] - After masking one or more of the first spots has been performed, the unmasked areas of the first and second masks are combined to generate the unmasked area of the merged mask, and all other areas of the merged mask are masked areas;
[0165] - Use the merged mask as a new first intermediate mask, select a third threshold from the thresholds to calculate a new second intermediate mask based on the third threshold, repeat steps c)-e) until each of the multiple thresholds is selected, and output the final generated merged mask as a merged binary mask M.
[0166] Multi-threshold-based multi-mask merging, as described above, can be advantageous because different thresholds can capture points or spots of tissue with very different intensity levels. If a spot is missed by one mask, it may be captured by a mask based on a less stringent threshold. Furthermore, the way the masks are combined prevents them from being contaminated by noise, as the contextual information of the two different masks is evaluated. If an unmasked pixel spot on one mask is very large or spatially close to a spot detected in another mask, that spot is retained only during mask merging. This is typically made if tissue areas are being analyzed, rather than if artifacts in glass areas are being analyzed. The merging of luminance-based and chrominance-based masks is performed in the same manner as described above, whereby the luminance mask may be used, for example, as the first mask, the chrominance mask may be used as the second mask, and connectivity analysis is performed to detect first and second spots in the unmasked areas of the luminance and chrominance channels, respectively.
[0167] According to an embodiment, the absolute-minimum-area is in the range of 100%-200% of the expected size of the spot representing tissue cells, and the conditional-minimum-area is in the range of 30%-70% of the expected size of the spot representing tissue cells. A predefined neighboring region defined around the spot can be formed by pixels with a width of about 1-3 diameters (e.g., 50 pixels) around the spot, containing typical tissue cells.
[0168] According to an embodiment, a tissue detection routine for a specific slide type includes a TMA slide routine. This TMA slide routine is configured to detect tissue regions in a tissue microarray (TMA) slide. The TMA slide routine is configured to:
[0169] - Generate a binary mask M by executing a subroutine according to the above embodiments;
[0170] - A binary mask (M) is applied to the digital image and the unmasked areas in the digital image are selectively analyzed for the purpose of detecting the mesh of tissue regions. As used herein, a “core” is a single tissue sample or a portion thereof contained on a slide, which is separated from other tissue regions (other “cores”) on the slide by a predefined minimum distance, such as at least 2 mm or more.
[0171] Embodiments of the present invention utilize multiple thresholds for intermediate mask generation and then combine these masks to generate a final mask. The image analysis system uses this final mask to identify tissue regions. Applying multiple different thresholds to mask generation and then combining the masks via an observed intelligent masking algorithm provides higher accuracy in tissue and glass detection. For different tissue slide types, the manner in which multiple thresholds are generated and the corresponding masks are fused may differ slightly to further increase the accuracy of the tissue detection algorithm.
[0172] According to an embodiment, the tissue detection routine for the selected specific slide type is a TMA slide routine, and the image analysis system is further configured to:
[0173] - If no mesh of tissue region is detected in the unmasked area of the digital image, it is determined that the selected TMA slide tissue detection routine does not correspond to the tissue slide type depicted in the digital image; and
[0174] - Terminate the execution of the TMA slide tissue detection routine and begin the execution of the general tissue detection routine.
[0175] According to an embodiment, the tissue detection routine for this specific slide type includes a control slide routine. This control slide routine is configured to detect tissue regions in a control slide. A control slide is a slide that includes control tissue in addition to the tissue to be analyzed. The control tissue and the tissue to be analyzed are stained using a control stain (i.e., a non-specific biomarker stain) instead of a specific biomarker stain. Due to the control stain, the control tissue sections, and the actual tissue sections to be analyzed, are never clearly stained and are typically more blurred than experimental test slides. The control slide routine is configured to:
[0176] - Generate a binary mask M by executing any of the subroutines in the above embodiments for binary mask generation;
[0177] - Apply a binary mask M to the digital image; and
[0178] - Selectively analyze unmasked areas in digital images to detect numerous tissue areas located in straight lines.
[0179] According to an embodiment, the tissue detection routine for the selected specific slide type is a control slide routine. If no tissue regions in a straight line are detected in the unmasked area of the digital image, the image analysis system determines that the selected control slide tissue detection routine does not correspond to the tissue slide type depicted in the digital image and terminates the execution of the control slide tissue detection routine and begins the execution of the general tissue detection routine.
[0180] According to an embodiment, a tissue detection routine for a specific slide type includes a cell line slide routine. This cell line slide routine is configured to detect tissue regions within a cell line slide. A cell line slide is a tissue slide comprising a single disc-shaped tissue sample, and this cell line slide routine is configured to:
[0181] - The grayscale image is calculated as a derivative of the digital image; according to some embodiments, the grayscale image is "enhanced" by means of contrast stretching. According to an embodiment, the (enhanced) grayscale image is downsampled (stored at a reduced resolution) multiple times (e.g., using a factor of 2, cubic) before the gradient magnitude image is calculated;
[0182] - Calculate the derivative of the gradient magnitude image with a grayscale image;
[0183] - Calculate the histogram of a digital image;
[0184] - Extract multiple intensity thresholds from the histogram;
[0185] - Apply multiple intensity thresholds to the digital image to generate multiple intermediate masks from the digital image;
[0186] - Generate a binary mask (M) by combining all intermediate masks;
[0187] - Apply a binary mask (M) to digital images and gradient magnitude images; and
[0188] - Selectively perform radially symmetric voting on unmasked areas in gradient magnitude images to detect the center and radius of individual tissue regions.
[0189] According to an embodiment, the tissue detection routine for the selected specific slide type is a cell line slide routine, and the image analysis system is further configured to:
[0190] - If no single tissue region is detected, or if the detected radius or center position deviates from the radius or center position expected for the cell line slide type, then the selected cell line slide tissue detection routine is determined not to correspond to the tissue slide type depicted in the digital image; and
[0191] - Terminate the execution of the cell line slide tissue testing routine and begin the execution of the general tissue testing routine.
[0192] According to an embodiment, the tissue detection routine for this particular slide type includes a smear slide routine. This smear slide routine is configured to detect tissue regions in a smear slide. A smear slide is a type of slide on which tissue regions diffuse throughout the slide. The smear slide routine is configured to:
[0193] - Calculate the luminance image as a derivative of the digital image, which is the L-channel image of the digital image represented in the LUV color space;
[0194] - Calculate the chroma image as a derivative from the digital image, which is the derivative of the U and V channels of the digital image in the LUV color space representation;
[0195] - Perform a hysteresis thresholding method on the inverse version of the brightness image to generate a brightness-based mask;
[0196] - Perform a hysteresis thresholding method on the inverse version of the chroma image to generate a chroma-based mask;
[0197] - Combine luminance-based masks and chrominance-based masks to obtain a binary mask (M);
[0198] - Applying binary masks to digital images; and
[0199] - Selectively analyze unmasked areas in digital images to detect spreading tissue areas.
[0200] According to an embodiment, performing the hysteresis threshold method includes:
[0201] - For each set of neighboring pixels whose pixel values satisfy the lower and upper thresholds applied during the hysteresis thresholding method, determine whether the set of pixels at least covers an area determined empirically; for example, the area determined empirically may correspond to a typical number of pixels of tissue cells depicted in a digital image at a given resolution.
[0202] - If so (if the set of pixels at least covers the area), then generate a binary mask that makes the set of pixels unmasked; treating the pixels as unmasked pixels means that the pixels are treated as organized regions while masked pixels are treated as unorganized regions (e.g., glass regions);
[0203] - If not, then generate a binary mask that masks the set of pixels.
[0204] Comparing neighboring pixel sets may have the advantage of filtering out small speckles and noise that are smaller than those in typical cells by the generated mask.
[0205] According to an embodiment, the tissue slide is a glass slide, and the tissue detection routine is configured to generate a mask that selectively masks the glass region, wherein the tissue region is the unmasked area in the mask.
[0206] According to an embodiment, the subroutine further includes using size and distance constraints to detect and discard smaller regions from the intermediate mask when combining intermediate masks to generate a binary mask.
[0207] According to embodiments, one or more tissue detection routines, both general and specific tissue slide types, are further configured to generate a weighted image from a digital image, in which each pixel has been assigned a weight within a range of values including more than two distinct values. This weight indicates the likelihood that the pixel belongs to a tissue region compared to belonging to glass. This can be advantageous because the mask merely generates a binary image, but the probability of a pixel being a tissue region rather than a glass region can be between zero and 1. For some pixels, it may be nearly impossible to definitively determine whether they are tissue pixels or not. By assigning weights that can have any of a plurality of possible values, it is possible to show the user that some expected tissue regions have a high probability of being true tissue pixels and some expected tissue regions have a lower probability of being tissue regions (but are, however, unmasked and therefore considered tissue regions).
[0208] According to an embodiment, the generation of the weighted image uses a digital image and a binary mask M as input. The soft-weighted image has zero values at all pixels (except for pixels that are non-zero inside the binary mask), and these non-zero pixels are assigned higher or lower values depending on whether they belong to more or fewer tissues compared to belonging to glass.
[0209] According to an embodiment, the generation of the weighted image includes:
[0210] - Calculate the luminance image as a derivative of the digital image, which is an L-channel image of the digital image in the LUV color space representation, or access a luminance image that has already been calculated for use in generating a binary mask (M).
[0211] - Calculate the chroma image as a derivative from the digital image, which is the derivative of the U and V channels of the digital image's LUV color space representation, or access the already calculated chroma image to generate a binary mask M.
[0212] - Use a binary mask (M) to mask luminance and / or chrominance, which selectively masks unorganized areas;
[0213] - Generate a weighted image SW as a derivative of the luminance and / or chrominance images, wherein each pixel in the weighted image has been assigned a value that is positively correlated with the chrominance of the corresponding pixel in the luminance image and / or positively correlated with the darkness of the corresponding pixel in the luminance image.
[0214] In general, colored pixels (whose color information is contained in the chromaticity image) and darker pixels (luminance image) are more likely to be tissue than glass; therefore, pixels with high chromaticity (color) components and / or darker pixels (according to luminance) have higher weights.
[0215] As described with respect to embodiments of the present invention, the calculation of the chromaticity UV and luminance L images is performed by an image analysis system.
[0216] According to an embodiment, the image analysis system includes a graphical user interface configured to display selectable GUI elements. These GUI elements allow the user to select one of a specific slide type or a general tissue detection routine. The specific slide type tissue detection routine is one of a TMA slide routine, a control slide routine, a cell line slide routine, or a smear slide routine. Alternatively or additionally, the image analysis system displays the detected tissue region, wherein the color and / or brightness of the pixels of the displayed tissue region depends on a weight assigned to each of the pixels in each tissue region [soft weighting]. Alternatively or additionally, the GUI allows the user to add additional focal points to each area of the displayed (soft-weighted) tissue region, the addition of which triggers the scanner to scan the area indicated by the focal point in greater detail.
[0217] In another aspect, the present invention relates to an image analysis method for detecting tissue regions in digital images of tissue slides. Tissue samples are mounted on slides. The method is implemented in an image analysis system including a processor and a storage medium. The method includes:
[0218] - Select one of the tissue detection routines for a specific slide type;
[0219] - Before and / or simultaneously perform the tissue detection routine for the selected slide type, check whether the tissue detection routine for the selected slide type corresponds to the tissue slide type depicted in the digital image;
[0220] - If so, automatically execute the tissue detection routine for the selected specific slide type to detect tissue regions in the digital image;
[0221] - If not, automatically execute the general tissue detection routine to detect tissue regions in digital images.
[0222] In another aspect, the present invention relates to a storage medium comprising instructions interpretable by a processor of an image analysis system, which, when executed by the processor, cause the image analysis system to perform a method according to an image processing method implemented in any of the embodiments described herein.
[0223] Computers typically include known components such as processors, operating systems, system memory, memory storage devices, input / output controllers, input / output devices, and display devices. Those skilled in the art will also understand that many possible configurations and components exist for computers, and that they may also include cache memory, data backup units, and many other devices. Examples of input devices include keyboards, cursor control devices (e.g., mice), microphones, scanners, and so on. Examples of output devices include display devices (e.g., monitors or projectors), speakers, printers, network cards, and so on. Display devices may include display devices that provide visual information, which may typically be logically and / or physically organized as an array of pixels. Interface controllers may also be included, which may include any of a wide variety of known or future software programs for providing input and output interfaces. For example, an interface may include an interface that provides one or more graphical representations to a user, often referred to as a “graphical user interface” (GUI). Interfaces typically enable the acceptance of user input using selection or input devices known to those skilled in the art. The interface may also be a touchscreen device; in some or alternative embodiments, applications on the computer may employ an interface, including interfaces referred to as “command line interfaces” (also known as CLI). CLIs typically provide text-based interaction between applications and users. Typically, a command-line interface presents output via a display device and receives input as lines of text. For example, some implementations may include those referred to as "shells," such as the Unix Shell known to those skilled in the art, or the Microsoft Windows PowerShell, which employs an object-oriented programming architecture (such as the Microsoft .NET Framework).
[0224] Those skilled in the art will recognize that the interface may include one or more GUIs, CLIs, or combinations thereof. The processor may include commercially available processors, such as Celeron, Core, or Pentium processors manufactured by Intel Corporation, SPARC processors manufactured by Sun Microsystems, Athlon, Sempron, Phenom, or Opteron processors manufactured by AMD Corporation, or it may be one of other processors available or to become available. Some embodiments of the processor may include those referred to as multi-core processors and / or those enabling parallel processing techniques in single-core or multi-core configurations. For example, multi-core architectures typically include two or more processor “execution cores.” In this example, each execution core may execute as a separate processor implementing the parallel execution of multiple threads. Furthermore, those skilled in the art will recognize that the processor may be configured in those configurations commonly referred to as 32-bit or 64-bit architectures, or other architectural configurations now known or potentially to be developed in the future.
[0225] Processors typically run an operating system, which can be, for example, a Windows-type operating system from Microsoft; a Mac OS X operating system from Apple Computer; a Unix or Linux-type operating system available from many vendors or known as open source; other or future operating systems; or a combination thereof. The operating system interacts with firmware and hardware in a known manner and facilitates the processor's coordination and execution of various computer programs written in a wide variety of programming languages. The operating system, which typically works with the processor, coordinates and executes the functions of other components of the computer. The operating system also provides scheduling, input / output control, file and data management, memory management, and communication control, as well as related services, all based on known technologies.
[0226] System memory may include any of a wide variety of known or future memory storage devices that can be used to store desired information and are accessible to a computer. Computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Examples include any generally available random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), digital versatile disc (DVD), magnetic media (such as resident hard disks or tapes), optical media (such as read-and-write discs), or other memory storage devices. Memory storage devices may include any of a wide variety of known or future devices, including disc drives, tape drives, removable hard disk drives, USB or flash drives, or floppy disk drives. These types of memory storage devices typically read from and / or write to program storage media, such as discs, tapes, removable hard disks, USB or flash drives, or floppy disks. Any of these program storage media, or other program storage media currently in use or subsequently developed, can be considered a computer program product. As will be appreciated, these program storage media typically store computer software programs and / or data. Computer software programs (also referred to as computer control logic) are typically stored in system memory and / or program storage devices used in conjunction with memory storage devices. In some embodiments, a computer program product is described that includes a computer-usable medium in which control logic (computer software programs, including program code) is stored. When executed by a processor, this control logic causes the processor to perform the functions described herein. In other embodiments, for example, a hardware state machine is used to implement certain functions primarily in hardware. The implementation of a hardware state machine to perform the functions described herein will be apparent to those skilled in the art. Input / output controllers may include any of a wide variety of known devices for accepting and processing information from a user, whether human or machine, local or remote. Such devices include, for example, modem cards, wireless cards, network interface cards, sound cards, or other types of controllers for any of a wide variety of known input devices. The output controller may include a controller for presenting information to a wide variety of known display devices, whether human or machine, local or remote. In the previously described embodiments, the functional elements of the computer communicate with each other via a system bus. Some embodiments of the computer may use a network or other types of remote communication to communicate with certain functional elements.As will be apparent to those skilled in the art, if the instrument control and / or data processing application is implemented in software, it can be loaded into and executed from system memory and / or memory storage devices. All or part of the instrument control and / or data processing application may also reside in a read-only memory or similar device that does not require initial loading of the application via an input / output controller. Those skilled in the art will understand that, for the benefit of execution, the instrument control and / or data processing application, or parts thereof, can be loaded by a processor into system memory, cache memory, or both in a known manner. Furthermore, the computer may include one or more library files, experimental data files, and an internet client stored in system memory. For example, experimental data may include data relating to one or more experiments or analytes (such as detected signal values), or other values associated with one or more sequencing-by-synthesis (SBS) experiments or processes. Additionally, the internet client may include applications that enable access to remote services on another computer using a network and may include, for example, those generally referred to as "web browsers." In this example, some commonly used web browsers include Microsoft Internet Explorer (available from Microsoft Corporation), Mozilla Firefox (available from Mozilla Corporation), Safari (from Apple Computer Corporation), Google Chrome (from Google Corporation), or other types of web browsers currently known or to be developed in the art. Furthermore, in some or other embodiments, the internet client may include a dedicated software application or may be a component of a dedicated software application that enables access to remote information via a network (such as a data processing application for a biological application).
[0227] A network may include one or more of many different types of networks known to those skilled in the art. For example, a network may include a local area network (LAN) or a wide area network (WAN), which may employ those commonly referred to as the TCP / IP protocol suite for communication. A network may include a network of a global system containing an interconnected computer network commonly referred to as the Internet, or may also include various intranet architectures. Those skilled in the art will also recognize that some users in a networked environment may prefer to use those commonly referred to as “firewalls” (sometimes also called packet filters or perimeter protection devices) to control information traffic to and from hardware and / or software systems. For example, a firewall may include hardware or software components or a combination thereof, and is typically designed to implement security policies placed appropriately by users (such as network administrators, for example).
[0228] The foregoing disclosure, which provides exemplary embodiments of the subject matter, is intended for illustrative and descriptive purposes. It is not intended to be exhaustive or to limit the subject matter disclosure to the precise forms disclosed. Many variations and modifications of the embodiments described herein will be apparent to those skilled in the art based on the foregoing disclosure. The scope of this subject matter disclosure will be defined only by the claims appended herein and their equivalents.
[0229] Furthermore, in describing representative embodiments of the disclosed subject matter, the specification may have presented the methods and / or processes of the disclosed subject matter as a specific sequence of steps. However, the method or process should not be limited to the specific order of the steps set forth herein, as the method or process is independent of this specific order. Other orders of steps are possible, as will be readily recognized by those skilled in the art. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation of the claims. Moreover, the claims to the methods and / or processes of the disclosed subject matter should not be limited to their steps in the written order, and those skilled in the art will readily recognize that the order can be changed while still remaining within the spirit and scope of the disclosed subject matter.
Claims
1. An image analysis system configured for detecting tissue regions in a digital image of a tissue slide, wherein a tissue sample is mounted on a tissue slide, the image analysis system including a processor (220) and a storage medium including a plurality of tissue detection routines for specific slide types and general tissue detection routines, the image analysis system being configured to perform methods including the following: - Select one of the tissue detection routines for a specific slide type; - Before and / or simultaneously perform the tissue detection routine for the selected slide type, check whether the tissue detection routine for the selected slide type corresponds to the tissue slide type depicted in the digital image; - If so, automatically execute the tissue detection routine for the selected specific slide type to detect tissue regions in the digital image; - If not, automatically execute the general tissue detection routine to detect tissue regions in digital images.
2. The image analysis system according to claim 1, wherein the digital image is a thumbnail image.
3. The image analysis system according to claim 1 or 2, wherein the image analysis system includes a user interface, wherein the selection of a tissue detection routine for a specific slide type includes receiving a user selection of a tissue slide type via the user interface; and selecting one of the tissue detection routines for a specific slide type specified by the tissue slide type selected by the user.
4. The image analysis system of claim 1, wherein one or more of the general tissue detection routine and the tissue detection routine for a specific slide type include a subroutine for generating a binary mask that selectively masks non-tissue regions, the subroutine comprising: - Calculate the histogram of the grayscale version of the digital image; - Extract multiple intensity thresholds from the histogram; - Apply multiple intensity thresholds to the digital image to generate multiple intermediate masks from the digital image; - Generate a binary mask by combining all intermediate masks.
5. The image analysis system according to claim 4, wherein extracting multiple intensity thresholds from the histogram includes: - Identify max_grayscale_index, which is the maximum grayscale intensity value observed in the histogram; - Identify the max_histogram_location, which is the gray level intensity value with the highest frequency of occurrence in the histogram; - The rightgap is calculated by subtracting the max_histogram_location from the max_grayscale_index, where the rightgap is the width of the peak of the observed histogram. - The glass_left_cutoff is calculated based on glass_left_cutoff = max_histogram_location - rightgap, where glass_left_cutoff is an estimate of where the glass end of the slide is to the left of the peak; - Perform multiple threshold extractions such that the lowest of the thresholds is equal to or greater than the intensity value of glass_left_cutoff–rightgap and the highest of the thresholds is equal to or less than glass_left_cutoff+rightgap.
6. The image analysis system according to claim 5, wherein extracting multiple intensity thresholds from the histogram includes: - The interval_range_gap is calculated based on interval_range_gap = max(1, round(rightgap / constant)), where constant is a predefined value between 2 and 20; - The value of max_rightgap is calculated based on max_rightgap = min(rightgap–1, round(rightgap*gap_cutoff_fraction)), where gap_cutoff_fraction is a predefined value in the range of 0.5-0.95; - Calculate number_interval_terms based on number_interval_terms = ((max_rightgap + rightgap) / interval_range_gap); and - Create multiple thresholds such that their number is the same as the calculated number_interval_terms.
7. The image analysis system according to claim 4, wherein applying multiple intensity thresholds to a digital image to generate multiple intermediate masks comprises: a) Create a first intermediate mask by masking all pixels in a digital image whose grayscale value is lower than a first threshold among a plurality of thresholds; b) Create a second intermediate mask by masking all pixels in a digital image whose grayscale value is lower than a second threshold among multiple thresholds; c) Perform connectivity analysis on the unmasked pixels of the first mask to identify the first spot; d) Perform connectivity analysis on the unmasked pixels of the second mask to identify the second spot; e) Merging the first and second masks into a merged mask, the merging including: - Mask all first spots whose size is less than the absolute minimum area and mask all first spots whose size is greater than the absolute minimum area but less than the conditional minimum area, such that there are no first spots in a predefined neighboring region around the first spots; - After masking one or more of the first spots has been performed, the unmasked areas of the first and second masks are combined to generate the unmasked area of the merged mask, and all other areas of the merged mask are masked areas; f) Use the merged mask as a new first intermediate mask, select a third threshold from the thresholds to calculate a new second intermediate mask based on the third threshold, repeat steps c)-e) until each of the multiple thresholds has been selected, and output the final generated merged mask as a merged binary mask.
8. The image analysis system according to claim 7, wherein the absolute minimum area is in the range of 100%-200% of the expected size of the spot representing the tissue cells, and the conditional minimum area is in the range of 30%-70% of the expected size of the spot representing the tissue cells.
9. The image analysis system according to any one of claims 4-8, wherein the tissue detection routine for the particular slide type includes a tissue microarray slide routine configured to detect tissue regions in a tissue microarray slide, the tissue microarray slide routine being configured to: - Generate a binary mask by executing the subroutine; - Apply a binary mask to a digital image and selectively analyze unmasked areas in the digital image for mesh detection of tissue regions.
10. The image analysis system of claim 9, wherein the tissue detection routine of the selected specific slide type is a tissue microarray slide routine and the image analysis system is further configured to: - If no mesh of the tissue region is detected in the unmasked area of the digital image, it is determined that the selected tissue microarray slide tissue detection routine does not correspond to the tissue slide type depicted in the digital image; and - Terminate the execution of the tissue microarray slide tissue detection routine and begin the execution of the general tissue detection routine.
11. The image analysis system according to any one of claims 4-8, wherein the tissue detection routine for the particular slide type includes a control slide routine configured to detect tissue regions in a control slide, the control slide routine being configured to: - Generate a binary mask by executing the subroutine; - Apply a binary mask to a digital image; as well as - Selectively analyze unmasked areas in digital images to detect numerous tissue areas located in straight lines.
12. The image analysis system of claim 11, wherein the tissue detection routine for the selected specific slide type is a slide control routine, and the image analysis system is further configured to: - If no tissue regions in a straight line are detected in the unmasked area of the digital image, it is determined that the selected control slide tissue detection routine does not correspond to the tissue slide type depicted in the digital image; and - Terminate the execution of the control slide tissue detection routine and begin the execution of the general tissue detection routine.
13. The image analysis system of claim 1 or 2, wherein the tissue detection routine for the particular slide type includes a cell line slide routine configured to detect tissue regions in a cell line slide, the cell line slide being a tissue slide comprising a single disc-shaped tissue sample, the cell line slide routine being configured to: - Calculate the derivative of a grayscale image with respect to a digital image; - Calculate the derivative of the gradient magnitude image with a grayscale image; - Calculate the histogram of a digital image; - Extract multiple intensity thresholds from the histogram; - Apply multiple intensity thresholds to the digital image to generate multiple intermediate masks from the digital image; - Generate a binary mask by combining all intermediate masks; - Apply binary masks to digital images and gradient magnitude images; as well as - Selectively perform radially symmetric voting on unmasked areas in gradient magnitude images to detect the center and radius of individual tissue regions.
14. The image analysis system of claim 13, wherein the tissue detection routine for the selected specific slide type is a cell line slide routine and the image analysis system is further configured to: - If no single tissue region is detected, or if the detected radius or center position deviates from the radius or center position expected for the cell line slide type, then the selected cell line slide tissue detection routine is determined not to correspond to the tissue slide type depicted in the digital image; and - Terminate the execution of the cell line slide tissue testing routine and begin the execution of the general tissue testing routine.
15. The image analysis system of claim 1 or 2, wherein the tissue detection routine for the particular slide type includes a smear slide routine configured to detect tissue regions in a smear slide, the smear slide being a slide type in which tissue regions diffuse throughout the slide, the smear slide routine being configured to: - Calculate the luminance image as a derivative of the digital image, which is the L-channel image of the digital image represented in the LUV color space; - Calculate the chroma image as a derivative from the digital image, which is the derivative of the U and V channels of the digital image in the LUV color space representation; - Perform a hysteresis thresholding method on the inverse version of the brightness image to generate a brightness-based mask; - Perform a hysteresis thresholding method on the inverse version of the chroma image to generate a chroma-based mask; - Combine luminance-based masks and chrominance-based masks to obtain a binary mask; - Apply a binary mask to a digital image; as well as - Selectively analyze unmasked areas in digital images to detect spreading tissue areas.
16. The image analysis system of claim 15, wherein performing the hysteresis thresholding method comprises: - For each set of neighboring pixels whose pixel values satisfy the lower and upper thresholds applied during the hysteresis thresholding method, determine whether the set of pixels at least covers an area determined empirically. -If so, generate a binary mask that makes the set of pixels unmasked; - If not, then generate a binary mask that masks the set of pixels.
17. The image analysis system of claim 1 or 2, wherein the tissue slide is a glass slide, and the tissue detection routine is configured to generate a mask that selectively masks the glass region, wherein the tissue region is an unmasked area in the mask.
18. The image analysis system of claim 4, further comprising using size and distance constraints to detect and discard smaller regions from the intermediate masks when combining intermediate masks to generate binary masks.
19. The image analysis system according to claim 1 or 2, wherein one or more of the general and tissue slide-specific tissue detection routines are further configured to: - Generate a weighted image from a digital image, in which each pixel has been assigned a weight in a range of values including more than two distinct values, the weight indicating the likelihood that the pixel belongs to a tissue region compared to belonging to glass.
20. The image analysis system according to claim 19, wherein the generation of the weighted image comprises: - Calculate the luminance image as a derivative of the digital image, which is an L-channel image of the digital image in the LUV color space representation, or access an already calculated luminance image for generating a binary mask. - Calculate the chroma image as a derivative from the digital image, which is the derivative of the U and V channels of the digital image's LUV color space representation, or access an already calculated chroma image for generating a binary mask. -Use a binary mask to mask luminance and / or chrominance, which selectively masks unorganized areas; - Generate a weighted image (SW) as a derivative of the luminance and / or chrominance images, wherein each pixel in the weighted image has been assigned a value that is positively correlated with the chrominance of the corresponding pixel in the luminance image and / or positively correlated with the darkness of the corresponding pixel in the luminance image.
21. The image analysis system according to claim 1 or 2, wherein the image analysis system includes a graphical user interface configured for: - Displays selectable GUI elements; these GUI elements allow the user to select one of a specific slide type or a general tissue detection routine; and / or - Display the detected tissue regions, wherein the color and / or brightness of the pixels in the displayed tissue regions depend on the weights assigned to each pixel in each tissue region; and / or - This allows users to add additional focus points to each area of the displayed tissue area. The addition of additional focus points triggers the scanner to scan the area indicated by the focus point in more detail.
22. An image analysis method for detecting tissue regions in a digital image of a tissue slide, wherein a tissue sample is mounted on a slide, and the method is implemented in an image analysis system including a processor and a storage medium, the method comprising: - Select one of the tissue detection routines for a specific slide type; - Before and / or simultaneously perform the tissue detection routine for the selected slide type, check whether the tissue detection routine for the selected slide type corresponds to the tissue slide type depicted in the digital image; - If so, automatically execute the tissue detection routine for the selected specific slide type to detect tissue regions in the digital image; - If not, automatically execute the general tissue detection routine to detect tissue regions in digital images.
23. A storage medium comprising instructions interpretable by a processor (220) of an image analysis system, wherein, when executed by the processor, the instructions cause the image analysis system to perform the method according to claim 22.
Citation Information
Patent Citations
Regression testing case selection method based on hierarchical slicing
CN101859276A
Systems and methods for sample display and review
CN103930762A