Optimized data processing for medical image analysis
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- VENTANA MEDICAL SYSTEMS INC
- Filing Date
- 2024-10-17
- Publication Date
- 2026-06-02
Smart Images

Figure 00000027_0000 
Figure 00000027_0001 
Figure 00000027_0002
Abstract
Description
[Technical field]
[0001] Field The present disclosure relates to digital pathology, and in particular to techniques for optimized data processing for medical image analysis. [Background technology]
[0002] background Digital pathology involves the scanning of a specimen slide (e.g., a histopathology or cytopathology glass slide) into a digital image. The tissues and / or cells in the digital image may then be examined by digital pathology (DP) image analysis and / or interpreted by a pathologist for a variety of reasons, including diagnosing disease, evaluating response to treatment, and developing drugs to combat disease. For example, evaluation of tissue changes caused by disease may be performed by examining thin tissue sections. A tissue sample may be sliced to obtain a series of sections (each section having a thickness of, for example, 4-5 microns), and each tissue section may be stained with a different stain or marker to reveal different features of the tissue. Each section may be mounted on a slide and scanned to create a digital image for examination by a pathologist. The pathologist may review and manually annotate the digital image of the slide (e.g., tumor areas, necrosis, etc.) to enable the extraction of meaningful quantitative measures using image analysis algorithms. Because tissues and / or cells are substantially transparent, preparation of pathology slides typically involves using various staining assays (e.g., immunostaining) that selectively bind tissue and / or cellular components to facilitate examination (e.g., by increasing the contrast between relevant features).
[0003] One of the most common examples of staining assays is the hematoxylin-eosin (H&E) staining assay, which includes two dyes that help identify tissue anatomical information. Hematoxylin primarily stains cell nuclei with a generally blue color, while eosin acts primarily as a generally pink stain for the cytoplasm, with other structures exhibiting different shades, hues, and combinations of these colors. H&E staining assays can be used to identify target substances in tissues based on the chemical, biological, or pathological properties of the target substances. Another example of a staining assay is the immunohistochemistry (IHC) staining assay, which involves the process of selectively identifying antigens (proteins) in cells of tissue sections by utilizing the principle of antibodies and other compounds (or substances) that specifically bind to antigens in biological tissues. In some assays, the target antigens in the specimen for the dye can be called biomarkers. Digital pathology image analysis can then be performed on the digital images of the stained tissues and / or cells to identify and quantify staining for antigens (e.g., biomarkers indicative of tumor cells) in the biological tissue.
[0004] In multiplexed slides of tissue specimens, different nuclei and tissue structures are stained simultaneously with certain biomarker-specific stains, which can be either chromogenic or fluorescent dyes. Each stain has a distinct spectral signature, in terms of spectral shape and spread. The spectral signatures of different biomarkers can be either broad or narrow spectral banded and may overlap spectrally. Slides containing specimens (e.g., tumor specimens) stained with several combinations of dyes are imaged using a multispectral imaging system. Each channel of the resulting image corresponds to a spectral band. Thus, the multispectral image stacks generated by the imaging system provide a rich set of images of the underlying component biomarkers, which in some instances may be colocalized. More recently, quantum dots have been widely used for immunofluorescence staining of biomarkers of interest due to their strong and stable fluorescence. Summary of the Invention
[0005] overview An apparatus and method for optimized data processing for medical image analysis is provided.
[0006] According to various aspects, a method is provided for analyzing an image of a tissue section, the method comprising: acquiring a plurality of image locations, each corresponding to a separate one of a plurality of biological structures; acquiring a plurality of locations of a first biomarker in the image; and calculating a distance transform array of at least a portion of the image including a plurality of seed locations. The method may include: for each of the plurality of seed locations and based on information from the first distance transform array, detecting whether the first biomarker is expressed at the seed location; and storing an indication of whether expression of the first biomarker was detected at the seed location in a data structure associated with the seed location. The method may include detecting colocalization of at least two phenotypes in at least a portion of the tissue section based on the stored indication.
[0007] According to various aspects, another method is provided for analyzing an image of a tissue section, comprising acquiring a plurality of image locations each corresponding to a separate one of a plurality of biological structures, and acquiring a first sparse binary segmentation mask including a first tissue region of the tissue section and excluding a second tissue region of the tissue section. The first sparse binary segmentation mask may include a plurality of pixel membership values and a plurality of microtile membership values and may indicate, for each of the plurality of pixels, a corresponding state of the first binary membership value. Each of the plurality of pixel membership values may correspond to a respective pixel of the plurality of pixels and may indicate a state of the first binary membership value for the pixel, and each of the plurality of microtile membership values may correspond to a respective microtile of the plurality of microtiles of the first binary mask and may indicate a state of the first binary membership value for all pixels in a block of the image corresponding to the microtile. The method may include determining, for each of a plurality of seed locations and based on information from the first sparse binary segmentation mask, whether a first binary membership value state for a pixel among the plurality of pixels corresponding to the seed location is a first state or a second state, and storing the first binary membership value state of the pixel in a data structure associated with the seed location. The method may include providing an analysis result including a calculated distance or distribution between biomarkers in cells of the first tissue region based on the stored states.
[0008] According to various aspects, further methods are provided for analyzing an image of a tissue section including a plurality of pixels and indicative of a plurality of biological structures. In some aspects, the method may include acquiring a plurality of image locations, each of the image locations corresponding to a separate one of the plurality of biological structures and indicative of a location of a description of the biological structures within the image, and acquiring a first binary mask for the image, the first binary mask indicative of a corresponding state of a first binary membership value for each of the plurality of pixels of the image. The first binary mask may include a plurality of pixel membership values and a plurality of microtile membership values, each of the plurality of pixel membership values corresponding to a separate one of the plurality of pixels and indicative of a state of the first binary membership value for the pixel, and each of the plurality of microtile membership values corresponding to a separate one of the plurality of microtiles and indicative of a state of the first binary membership value for all pixels within a block of the image corresponding to the microtile. In some aspects, the method may include, for each of the plurality of image locations and based on information from the first binary mask, acquiring a first binary mask for the image, the first binary mask indicative of a corresponding state of a first binary membership value for each of the plurality of pixels of the image. Based on the first binary membership value state of the pixel corresponding to the image location, the method further includes storing the first binary membership value state of the pixel corresponding to the image location in a data structure associated with the image location.
[0009] In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a first binary marker value indicative of a positive status of a first biomarker in the corresponding biological structure.
[0010] The method may further include obtaining a second binary mask for the image, the second binary mask indicating, for each of a plurality of pixels, a corresponding state of the second binary membership value. The second binary mask may include a plurality of pixel membership values and a plurality of microtile membership values, each of the plurality of pixel membership values corresponding to a separate one of the plurality of pixels and indicating a state of the second binary membership value for the pixel, and each of the plurality of microtile membership values corresponding to a separate one of the plurality of microtiles and indicating a state of the second binary membership value for all pixels in a block of the image corresponding to the microtile. In some aspects, the method further includes, for each of the plurality of image locations and based on information from the second binary mask, storing in a data structure associated with the image location a state of the second binary membership value of the pixel corresponding to the image location.
[0011] The method may further include calculating a distance transform array for at least a portion of the image, each value of the distance transform array corresponding to a different respective pixel of the image and indicating a distance between the pixel and a closest image location among the plurality of image locations at which the first binary marker value indicates a first positive state.
[0012] In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a second binary marker value indicative of a positive status of a second biomarker in the corresponding biological structure.
[0013] The method may further include storing values of the distance transform array corresponding to image locations where the second binary marker value indicates the first positive condition. The method may further include sorting the stored values in order of magnitude.
[0014] In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a first binary marker value indicative of a positive state of the first biomarker in the corresponding biological structure. In such a case, the method may further include calculating, for each of the plurality of overlapping tiles of the image, a distance transform array including, for each pixel of the tile, a corresponding value indicative of a distance between the pixel and a closest one of the plurality of image locations at which the first binary marker value indicates the first positive state.
[0015] In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a second binary marker value indicative of a positive state of a second biomarker in the corresponding biological structure, and each of the plurality of tiles may include an inner region that does not overlap with an inner region of any other tile of the plurality of tiles. In such a case, the method may further include storing, for each of the inner regions, a value of the distance transform array of the corresponding tile that corresponds to the image location where the second binary marker value indicates the first positive state.
[0016] In any of the various aspects of the method, e.g., as described above, each of the plurality of biological structures can be a cell nucleus. Additionally or alternatively, the image can be a multiplexed immunofluorescence image having multiple channels.
[0017] According to various aspects, a non-transitory computer-readable medium is provided. In some aspects, the non-transitory computer-readable medium may include instructions for causing one or more processors to perform operations for analyzing an image of a tissue section including a plurality of pixels and depicting a plurality of biological structures, the operations including acquiring a plurality of image locations, each of the image locations corresponding to a separate one of the plurality of biological structures and indicating a location of a depiction of the biological structures within the image, and acquiring a first binary mask for the image, the first binary mask indicating a corresponding state of a first binary membership value for each of a plurality of pixels of the image. The first binary mask may include a plurality of pixel membership values and a plurality of microtile membership values, each of the plurality of pixel membership values corresponding to a separate one of the plurality of pixels and indicating a state of the first binary membership value for the pixel, and each of the plurality of microtile membership values corresponding to a separate one of the plurality of microtiles and indicating a state of the first binary membership value for all pixels in a block of the image corresponding to the microtile. In some aspects, the method further includes, for each of the plurality of image locations and based on information from the first binary mask, storing a state of a first binary membership value of a pixel corresponding to the image location in a data structure associated with the image location.
[0018] In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a first binary marker value indicative of a positive status of a first biomarker in the corresponding biological structure.
[0019] The non-transitory computer-readable medium may further include instructions for causing the one or more processors to perform operations including obtaining a second binary mask for the image, the second binary mask indicating, for each of a plurality of pixels, a corresponding state of a second binary membership value. The second binary mask may include a plurality of pixel membership values and a plurality of microtile membership values, each of the plurality of pixel membership values corresponding to a separate one of the plurality of pixels and indicating a state of the second binary membership value for the pixel, and each of the plurality of microtile membership values corresponding to a separate one of the plurality of microtiles and indicating a state of the second binary membership value for all pixels in a block of the image corresponding to the microtile. In some aspects, the operations further include, for each of the plurality of image locations and based on information from the second binary mask, storing in a data structure associated with the image location a state of the second binary membership value of the pixel corresponding to the image location.
[0020] The non-transitory computer-readable medium may further include instructions to cause the one or more processors to perform operations including calculating a distance transform array for at least a portion of the image, where each value in the distance transform array corresponds to a different respective pixel of the image and indicates a distance between the pixel and a closest one of a plurality of image locations where the first binary marker value indicates a first positive state.
[0021] In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a second binary marker value indicative of a positive status of a second biomarker in the corresponding biological structure.
[0022] The non-transitory computer readable medium may be configured to cause the one or more processors to store values of the distance transform array corresponding to image locations where the second binary marker value indicates a first positive condition. The method may further include instructions for performing operations including: sorting the stored values by magnitude.
[0023] In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a first binary marker value indicative of a positive state of the first biomarker in the corresponding biological structure. In such a case, the operations may further include calculating, for each of the plurality of overlapping tiles of the image, a distance transform array including, for each pixel of the tile, a corresponding value indicative of a distance between the pixel and a closest one of the plurality of image locations at which the first binary marker value indicates the first positive state.
[0024] In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a second binary marker value indicative of a positive state of a second biomarker in the corresponding biological structure, and each of the plurality of tiles may include an inner region that does not overlap with an inner region of any other tile of the plurality of tiles. In such a case, the operations may further include storing, for each of the inner regions, a value of the distance transform array of the corresponding tile that corresponds to the image location where the second binary marker value indicates the first positive state.
[0025] In any of the various aspects of the non-transitory computer readable medium, such as those described above, each of the plurality of biological structures may be a cell nucleus. Additionally or alternatively, the image may be a multiplexed immunofluorescence image having multiple channels.
[0026] Numerous benefits are achieved through various embodiments over conventional techniques. For example, various embodiments provide methods and systems that can be used to effectively acquire data from large DP images (e.g., MPX images) and rapidly process statistical analysis and calculations in and between such images. In some embodiments, a sparse segmentation mask allows for rapid and accurate association of image locations with corresponding tissue types. In some embodiments, a two-level tile architecture supports multi-threaded processing. In some embodiments, a bitmap data structure supports compression of associated data for rapid (e.g., interactive) statistical and / or spatial analysis. In some embodiments, distance transform calculations over overlapping image regions support efficient calculation of distances and distributions. These and other embodiments, along with many of the advantages and features thereof, are described in more detail in conjunction with the following text and accompanying figures.
[0027] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0028] BRIEF DESCRIPTION OF THE DRAWINGS Aspects and features of various embodiments will become more apparent from the following detailed description of the invention, taken in conjunction with the accompanying drawings, in which: FIG. [Brief description of the drawings]
[0029] [Figure 1] 1 shows examples of multiplexed immunofluorescence (MPX) images. [Figure 2A] 1 shows another example of an MPX image. [Figure 2B] A portion of the MPX image of FIG. 2A is shown. [Figure 3A] 1 illustrates an example of a result tile divided into computation tiles in accordance with some aspects of the present disclosure. [Figure 3B] 3B shows the example of FIG. 3A where the tiles and the interior areas of the tiles are shown hatched. [Figure 4] 1 illustrates a grid for dividing a portion of an image into result tiles in accordance with some aspects of the present disclosure. [Diagram 5] 1 illustrates an example of a computational tile in accordance with some aspects of the present disclosure. [Figure 6A] 1 illustrates another example of a computational tile of an MPX image in accordance with some aspects of the present disclosure. [Figure 6B] 6B shows a computational tile of an epithelial tumor binary mask corresponding to the computational tile of FIG. 6A. [Figure 6C] 6B shows a computational tile of a stroma binary mask corresponding to the computational tile of FIG. 6A. [Figure 7A] 1 illustrates an example of an epithelial tumor binary mask integrated onto a result tile, according to some aspects of the present disclosure. [Figure 7B] 13 illustrates an example of a stromal binary mask integrated onto a result tile in accordance with some aspects of the present disclosure. [Figure 8] 1 illustrates a division of a binary mask into cells and microtiles in accordance with some aspects of the present disclosure. [Figure 9] 9 illustrates a classification of the microtiles of FIG. 8 into three classes in accordance with some embodiments of the present disclosure. [Figure 10A] 1 illustrates a bitmap data structure in accordance with some aspects of the present disclosure. [Figure 10B] 1 illustrates indexing of multiple phenotypes according to some embodiments of the present disclosure. [Figure 11A] 1 shows the distance between occurrences of different phenotypes in a portion of the MPX image. [Figure 11B] 1 is a flowchart illustrating an example of a method for calculating distances between locations of medical images in accordance with some aspects of the present disclosure. [Figure 12] 1 illustrates an application of a method for calculating distances between locations in medical images in accordance with some aspects of the present disclosure. [Figure 13]1 illustrates an application of a method for calculating distances between locations in medical images in accordance with some aspects of the present disclosure. [Figure 14] 1 is a flow chart illustrating an example of a method for image analysis in accordance with some aspects of the present disclosure. [Figure 15A] 1 is a flow chart illustrating another example of a method for image analysis in accordance with some aspects of the present disclosure. [Figure 15B] 1 is a flowchart illustrating a further example of a method for image analysis in accordance with some aspects of the present disclosure. [Figure 16] 1 is a flowchart illustrating a further example of a method for image analysis in accordance with some aspects of the present disclosure. [Figure 17] 1 is a flowchart illustrating a further example of a method for image analysis in accordance with some aspects of the present disclosure. [Figure 18] FIG. 1 is a block diagram of an exemplary computing environment having an exemplary computing device suitable for use in some exemplary implementations. [Figure 19] FIG. 1 shows an example of an annotated MPX image. [Figure 20] 13 shows examples of corresponding analysis results that may be obtained from such images in accordance with some aspects of the present disclosure. [Figure 21] 13 shows examples of corresponding analysis results that may be obtained from such images in accordance with some aspects of the present disclosure. [Figure 22] 13 shows another example of annotated MPX imagery. [Diagram 23] 13 shows examples of corresponding analysis results that may be obtained from such images in accordance with some aspects of the present disclosure. [Figure 24] 13 shows examples of corresponding analysis results that may be obtained from such images in accordance with some aspects of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0030] Detailed Description Although specific embodiments have been described, these embodiments are presented by way of example only and are not intended to limit the scope of protection. The devices, methods, and systems described herein may be embodied in various other forms. Furthermore, various omissions, substitutions, and modifications of the form of the exemplary methods and systems described herein may be made without departing from the scope of protection.
[0031] I. Overview The ability to characterize multiple biomarkers within a tissue (e.g., within a tumor tissue) and measure the heterogeneity of the presence and levels of such biomarkers within and between tissues can provide important information for understanding and characterizing various disease states and / or for appropriately selecting available targeted therapeutics for a patient's disease state. The ability to identify and measure regions within a tissue with differential distribution of important biomarkers can provide important information to inform the development of targeted and combination therapies. The development and selection of appropriate combination therapies can also be an important factor in preventing recurrence.
[0032] Multiplexed immunofluorescence (MPX) images of tissue sections can be obtained by staining the sections with two or more fluorophores that emit different respective wavelengths of light upon excitation (e.g., by ultraviolet light). Each channel of the resulting image can be acquired by controlling the excitation light (e.g., selecting an excitation laser of the appropriate wavelength) to generate the desired emission profile of the target fluorophores and filtering the emission to block unwanted spectral components. MPX staining of tissue sections allows for the simultaneous detection of multiple biomarkers and their co-expression at the individual cell level. Figure 1 shows an example of a pseudocolor version of a multiplexed immunofluorescence (MPX) image in which a corresponding pseudocolor (e.g., pink, blue, green) has been assigned to each of the different channels of the MPX image. This image also includes three manual annotations of the regions of interest (as indicated by red, white, and yellow closed curves (color) or by three overlapping closed curves (grayscale) at the left, center, and upper right of the image). A pathologist may manually annotate the MPX images to, for example, identify portions of the tissue to be analyzed using image analysis (eg, tumor regions, necrotic regions, etc.) and / or regions to be excluded from image analysis.
[0033] Primary analysis of MPX images may include biomarker and phenotype detection (e.g., co-expression of specific combinations of biomarkers), image segmentation into different tissue classes (e.g., epithelial tumor (i.e., tumor epithelium), stroma (e.g., connective tissue, supporting tissue, or other non-functional tissue of an organ), etc.), and / or extraction of potentially relevant features (e.g., location of cells (e.g., cell nuclei), etc.). Such analysis may be performed manually, but is more typically performed using automated processes such as computer vision, machine learning, and / or deep learning. Figure 2A shows another example of an MPX image. Figures 2B and 2C show a portion of the MPX image of Figure 2A, where the calculated epithelial tumor inclusion and exclusion regions are indicated by a set of polygons (marked with a blue outline (color) or indicated by a bright central area (grayscale) in Figure 2B), and the detected PanCK-positive (PanCK+) cells are indicated by red dots (color) or gray dots (grayscale) (e.g., in the central area of Figure 2C). Such a computed polygon segmentation is typically stored as a matrix array in a storage device, such as a primary or secondary memory.
[0034] Secondary analysis of MPX images may use the results from the primary analysis to obtain the next level of information, e.g., density distribution of one or more biomarkers; spatial relationships between biomarkers; co-localization of multiple phenotypes in the tumor, epithelium, and / or stroma; and / or other statistics and / or metrics. Such "readout analyses" may be used by pharmaceutical companies to obtain other It may be important to correlate with genomic sequencing findings and / or molecular features of tumors, and may help determine patient treatment response and / or prognosis for drug development. Comprehensive automated readout statistical analysis may include one or more (possibly all) of the following: density of different cell phenotypes in a region of interest (ROI) (e.g., "tumor"); distance between different phenotypes in the ROI; distance from various cell phenotypes to various biomarker positive regions in the ROI; descriptive statistics / metrics of biomarker positive regions such as blood vessels; descriptive statistics / metrics of different cell phenotypes within a certain distance from the ROI (e.g., immune cells, CD8); descriptive statistics / metrics of different biomarker positive regions (e.g., fibroblast activation protein positive (FAP+) regions) within a certain distance from the ROI; descriptive statistics of intensity-based (e.g., cell-based and / or region-based) metrics of different biomarkers; calculation and representation of calculated ROIs such as epithelial and stromal portions of tumors, represented by the presence or absence of tumor markers.
[0035] MPX images typically have about six channels, and may even have 32 or 64 or more channels. Furthermore, the pixel values in each channel of an MPX image may have 16 bits or more of resolution (compared to, for example, 8 bits of resolution for each of the three channels of a typical RGB image). Because the size of an MPX whole slide image (WSI) may be on the order of 100,000 pixels wide by 100,000 pixels high, the total storage size of an MPX image may be 5 or 10 gigabytes or more.
[0036] When processing such large images, it is not feasible to keep even a mask of the image in working memory. For MPX images such as those shown in Figure 1, it may be impractical to even store a pixel-level mask showing all epithelial tumor and stroma segmentations detected in the entire slide (e.g., as shown in Figure 2B). Instead, it is common to store only the polygons describing the segmentation (e.g., for use in future statistical analyses). During such analyses, it is common that only a subset of the polygon segmentation (e.g., the subset corresponding to the tile or other region of the image currently being analyzed) resides in memory at any one time. Factors such as the lack of easy access to segmentation information at the pixel level can cause automated read-out analyses of large tissue sections or multiple tissue blocks to take hours (even days). Thus, the time currently required to generate a statistical analysis report on large MPX images may not meet project requirements, especially for large-scale clinical trials. Moreover, such fragmentation of the polygon segmentation makes it difficult to keep track of different layers of the segmentation to determine which regions are excluded and included, which may affect the accuracy of statistical calculations of biomarkers in these layers.
[0037] The techniques disclosed herein may be used, for example, to design an effective readout analysis system to effectively fetch MPX data, efficiently process the analysis of whole slide images, and / or perform corresponding statistical analysis. Such a system may effectively retrieve data from large MPX images and rapidly process statistical analysis and calculations within and between such images, especially for large-scale clinical trials and / or to meet the needs of pharmaceutical customers. For example, such a system may include an efficient data structure and architecture design optimized to meet the computational complexity and big data processing requirements in MPX images. Although the problems described herein may be exacerbated by darkfield images (e.g., MPX images), which may have a much larger number of channels (e.g., up to 32, even 64), these techniques may also be applied to the processing and analysis of general DP images (e.g., whole slide images), including brightfield images (e.g., optical microscope images).
[0038] II. Definition As used herein, when an action is "based on" something, this means that the action is based on at least some "Based at least in part on" means "based at least in part on
[0039] As used herein, the terms "substantially," "approximately," and "about" are defined as being largely, but not necessarily completely, specified (and including that which is completely specified), as would be understood by one of ordinary skill in the art. In any disclosed embodiment, the terms "substantially," "approximately," or "about" may be replaced with "within [a percentage]" of that which is specified, where the percentage includes 0.1, 1, 5, and 10%.
[0040] As used herein, the term "biological material or structure" refers to a naturally occurring material or structure that comprises all or part of a living structure (e.g., a cell nucleus, a cell membrane, a cytoplasm, a chromosome, DNA, a cell, a cluster of cells, etc.).
[0041] III. Techniques for Optimized Storage and Processing for Medical Image Analysis Due to the large size of typical DP images, it may be desirable to perform image analysis by dividing the image to be analyzed into smaller (typically square) parts of equal size, called "tiles", which are then processed separately. In MPX analysis, even better performance may be obtained by using different tile sizes for different stages of the analysis. In one such example, a two-level tile architecture includes large tiles (also called "result tiles") and small tiles (also called "computational tiles" or "computation tiles"). In such a design, division of the entire slide image (or desired regions thereof) into large tiles may be applied to efficiently fetch image data into working memory (e.g., from disk). Each large tile may then be divided into smaller tiles (e.g., by computer vision / deep learning / machine learning algorithms) for use during primary analysis, which may include operations such as phenotype / biomarker detection and / or feature extraction. Finally, the large tiles may be used to consolidate results calculated using the smaller "computation" tiles (e.g., region masks, phenotype locations).
[0042] Potential benefits of such a two-level design may include more efficient data fetching due to fewer transactions with the image server (e.g., because larger tiles are being read). Additionally, smaller tiles are typically more suitable as input to deep learning / machine learning / computer vision algorithms that perform image analysis for segmentation and detection. Using smaller tiles during computation may also facilitate improved use of processor (e.g., CPU and / or GPU) resources such as caches.
[0043] 3A and 3B show an example of a two-level tiling architecture in which an image is divided into result tiles and further divided into computation tiles. FIG. 3A shows an example of a result tile 304 (indicated by the outer thin square) divided into an overlapping 3×3 array of equal-sized computation tiles. The size of each computation tile may be, for example, but not limited to, 128×128 pixels, 256×256 pixels, or 512×512 pixels, with each result tile being roughly three times larger in each dimension (e.g., depending on the size of the overlap). For computation tiles of size 512×512 pixels, the size of each result tile is approximately 2K×2K pixels.
[0044] As indicated by the thin lines in Figures 3A and 3B, each result tile overlaps its neighboring result tiles in the image, and each computation tile also overlaps its neighboring result tiles within the result tile. Such overlap allows for efficient handling of boundary conditions between different tiles (e.g., as described herein with reference to distance calculations). It may be desirable to implement the overlap to be equal to the calculated maximum distance (eg, 100 microns, 500 microns).
[0045] Each result tile includes an inner region that does not overlap with any other result tile in the image, and each computational tile includes an inner region that does not overlap with any other computational tile in the result tile. In FIG. 3A, the inner region of the result tile 304 is shown with an outer bold square. FIG. 3B shows an example of FIG. 3A where the computational tile 308 and the inner region 312 of another computational tile are shaded. FIG. 4 shows an example of a grid that divides a portion of an MPX image into result tiles (result tile 404 is shown in red), and FIG. 5 shows an example of computational tiles 508 (shaded in blue) of a portion of an MPX image.
[0046] A two-level tile architecture (e.g., as shown in Figures 3A and 3B) may be implemented such that each result tile is read into memory by a corresponding CPU thread, which then processes the individual computational tiles within the result tile. During processing, results associated with the computation may be stored within each individual result tile and then accumulated in a backend. Such an architecture also supports completely independent processing of each computational tile, without requiring synchronization between different parts of the image or different stages of the process. Such independence of computation enables multi-threaded implementations in which multiple or multiple processors run in parallel, each processor decoding and processing a corresponding portion of the image.
[0047] As discussed above, describing image segmentation in the form of polygons may be an imperfect solution that results in processing inefficiencies and / or inaccuracies. Instead, it may be desirable to configure the primary analysis to generate binary masks (also called "region masks") corresponding to each layer of segmentation (e.g., tumor mask, epithelial tumor mask, stromal mask). Such an approach may be well suited to a tile-based image analysis process. The use of pixel-level segmentation masks (e.g., each pixel of the mask indicates the membership state of the corresponding pixel in the image) may also avoid the need to convert from polygons to pixels at runtime and / or resolve tile boundary inaccuracies. FIG. 6A shows an example of a computational tile 608 of an MPX image, and FIG. 6B and FIG. 6C show corresponding computational tiles 612 and 616, respectively, of epithelial tumor binary mask and stromal binary mask generated by segmentation analysis of the image computational tile 608 of FIG. 6A. FIG. 7A shows an example where a computational tile of an epithelial tumor binary mask has been integrated onto a result tile 704, and FIG. 7B shows an example where a computational tile of a stromal binary mask has been integrated onto a result tile 708 (e.g., for use in future readout analysis).
[0048] Unfortunately, using binary segmentation masks instead of polygon segmentation can significantly increase storage requirements and / or incur multiple disk access operations for each image tile. Because binary masks are typically stored as simple images (i.e., one byte per pixel), the amount of disk storage required can increase substantially, and the amount of storage required to maintain such masks in working memory can be prohibitive for processing full slide images. A tile-based analysis process may allow mask tiles to be swapped into working memory as needed, but such an approach also increases the number of disk accesses required to perform the analysis.
[0049] The use of sparse binary masks as described herein to represent binary computation region masks (e.g., for epithelial, stromal, and vascular regions) allows the operations to be implemented efficiently. For example, it has been shown that memory requirements can be significantly reduced (i.e., only a fraction of the bits per pixel).
[0050] As demonstrated with reference to FIG. 8 and FIG. 9, converting a binary mask (e.g., each pixel of the mask indicates the membership state of a corresponding pixel of the image) into a sparse binary mask may include partitioning the mask image into an array of non-overlapping cells (so as not to be confused with biological cells in a tissue section). FIG. 8 illustrates a portion of a binary computation domain mask partitioned into an array of non-overlapping cells (as indicated by a grid of white lines). Each cell (e.g., cell 804) is completely independent of one another in the mask image, supporting the use of multiple threads to process the WSI in parallel. Such partitioning of the WSI mask into cells may be performed on the entire mask image at once or piecewise (e.g., tile-wise), such as on a combined result tile of the mask (e.g., an inner region of a result tile). In one example, the size of each cell is 512×512 pixels, although other cell sizes (e.g., 256×256 pixels, 1024×1024 pixels) are also possible.
[0051] As shown in FIG. 8, each cell in the mask image is further divided into an array of non-overlapping microtiles (also called "mittels"). In one example, each 512×512 pixel cell is divided into an array of 16×16 microtiles, each having a size of 32×32 pixels, corresponding to a block of the image that is the same size as the microtile and contains pixels that correspond to the pixels in the microtile. Other microtile sizes (e.g., 16×16 pixels, 64×64 pixels) are also possible. Due to the relatively small size of these microtiles and the spatial coherency of the binary mask, many of the microtiles in a cell contain only "white" pixels (e.g., binary value 1) or only "black" pixels (e.g., binary value 0), but some microtiles contain both black and white pixels.
[0052] FIG. 9 illustrates the classification of the microtiles of FIG. 8 into three classes. In the first class, all pixels of the microtile are black (e.g., binary value 0), and in the second class, all pixels of the microtile are white (e.g., binary value 1). Thus, the values of all pixels of the microtile in either of these classes (i.e., the membership values of all pixels of the image corresponding to pixels in the microtile) may be represented as a single binary state. In one implementation of the sparse binary mask representation, no values are stored for microtiles of the first class, and a null pointer (or similar Flye weight value) is stored for each microtile of the second class. In the third class, the microtile contains pixels of both binary values (shown in FIG. 9 as orange (color) or gray (grayscale)). In this case, the value of each pixel of the microtile is stored (e.g., as an ordered string of 1024 bits for a microtile of size 32×32 pixels). In one implementation of the sparse binary mask representation, a pool allocator is used to allocate 32x32 bits (128 bytes) of memory, and the bitmask of the micro-tiles is stored in the allocated memory.
[0053] The application of the sparse binary mask implementation described herein allows for a significant reduction in memory requirements. For example, typical memory requirements of approximately 0.2 bits per pixel have been measured in practice for real MPX WSIs (e.g., as opposed to 8 bits per pixel in "naive" implementations). Such a reduction allows for in-memory storage of multiple WSI binary masks. Furthermore, the technique also allows for efficient implementation of various important binary mask operations, such as, for example, efficient downsampling and upsampling of binary masks to create image resolution pyramids. The use of one or more sparse binary segmentation masks may also enable statistical processing that is not feasible or practical with polygon segmentation (e.g., distribution of distances between biomarkers or other features, colocalization analysis).
[0054] The results of primary analysis of MPX images may include the localization of multiple biomarkers from such analysis, which may be used to detect colocalization of biomarkers and phenotypes. Effective detection of all the different combinations of colocalization between large MPX datasets may be a key condition to enable efficient readout analysis.
[0055] Primary analysis of MPX images may also include identification of relevant image locations, such as cell nuclei. It may be desirable to use such image locations as reference points for localization of detected biomarkers (e.g., for phenotype identification). Figure 10A shows an example of a bitmap data structure that uses logical bit calculations to record all of the different combinations of phenotype / biomarker colocalization for a particular image location (also called a "seed location"). Using such a data structure may significantly reduce memory usage during calculations and detection.
[0056] The example of FIG. 10A includes indices of five different biomarkers (denoted as marker 1 through marker 5), with each biomarker associated with a unique corresponding location in the bitmap. In this example, a binary value of "1" indicates that expression of the marker has been detected at the location (e.g., within a predefined neighborhood of the location, within the boundaries of a cell or other biological structure associated with the location), and a binary value of "0" indicates that expression of the marker has not been detected at the location. Thus, as shown in FIG. 10B, the particular example shown in FIG. 10A supports the identification of up to 32 different phenotypes (e.g., unique combinations of expression of each of the five biomarkers). These 32 phenotypes are indexed in FIG. 10B as phenotype 0 through phenotype 31 according to the decimal value of the string of bits assigned to the markers, marker 1 through marker 5, with the particular string as shown in FIG. 10A corresponding to phenotype 10. It will be appreciated that this data structure may be extended to include such indices of expression for each of any number of different biomarkers at a location.
[0057] Additionally or alternatively, such a bitmap data structure may include binary indicators, each corresponding to a separate tissue region (e.g., as indicated by the corresponding segmentation described herein). The example of FIG. 10A includes three additional binary indicators, each corresponding to a separate segmentation mask (e.g., a sparse binary mask as described herein). In this example, without limitation, the three masks are a stromal mask, an epithelial tumor mask, and a (e.g., vascular) mask, where a binary value of "1" indicates that the mask identifies the location as being included in the corresponding tissue region (e.g., based on the mask value corresponding to the location, or based on a majority of the mask values corresponding to pixel locations within a predefined neighborhood of the location, or within the boundaries of a cell or other biological structure associated with the location), and a binary value of "0" indicates that the mask identifies the location as being excluded from the corresponding tissue region. It will be appreciated that this data structure may be extended to include such indicators of region membership for the location for each of any number of different regions (e.g., stroma, epithelial tumor, tumor, "other region").
[0058] In the particular example of FIG. 10A, all information from the MPX data relevant to this location for a particular readout analysis is stored in a single byte. By encoding all important information as a combined bit pattern into one (or potentially multiple) bytes and storing that information only for the relevant location of the image (e.g., the center of each cell), a data representation that can be queried very efficiently can be obtained. In some implementations, once such information has been computed and recorded in such a bitmap data structure by a corresponding processor for each identified relevant location within the inner region of the MPX image result tile, the image tile and any corresponding segmentation mask tiles are discarded from memory. It can be done.
[0059] It may be desirable to calculate spatial relationships between different biomarkers in a particular region of interest (e.g., in tumor and active stroma regions). Knowledge of such complex spatial relationships may enable a better understanding of the relationships between different biomarkers / phenotypes and regional areas (e.g., blood vessels, active stroma, tumor, etc.). One example of such calculations may include the following operations: (1) For each occurrence of phenotype A in the MPX image (or a selected portion thereof), find and record the distance (e.g., Euclidean distance) to the nearest occurrence of phenotype B. (2) Optionally, calculate a histogram of the recorded distances (e.g., count the number of distances less than or equal to 10 microns, the number of distances greater than 10 microns and less than or equal to 20 microns, ..., the number of distances greater than 90 microns and less than or equal to 100 microns, etc.). (3) Optionally, calculate other statistical data such as the average (e.g., mean) of the collected distances, the standard deviation of the collected distances, etc. It will be appreciated that such calculations may use one or more other distance measures (e.g., city-block or L1 distance) instead of or in addition to Euclidean distance, may use distance bins of different sizes (possibly unequal sizes), and / or may use one or more other average measures (e.g., median, mode) instead of or in addition to the mean.
[0060] To support the calculation of spatial relationships between different biomarkers in a particular region of interest, it may be desirable to provide a method that can be used to efficiently calculate the distance for each occurrence of one selected phenotype to the nearest occurrence of a different selected phenotype. Such a method may be desirable to support such calculations for any pair of local phenotypes, or even for any combination of local phenotypes. Additionally or alternatively, such a method may be desirable to support one or more further restrictions of phenotype selection according to other factors (e.g., occurrence within a particular tissue region, e.g., epithelial tumor or stroma). It is noted that a bitmap data structure is described herein with reference to Figures 10A and 10B and can be used to assist in the efficient selection of image locations consistent with a desired phenotype selection.
[0061] 11A shows an example of an image portion with one occurrence A1 of phenotype A and four occurrences B1, B2, B3, B4 of phenotype B. Of the distances A1-B1, A1-B2, A1-B3, A1-B4 between the occurrences of these two phenotypes, the shortest distance is the distance A1-B3.
[0062] FIG. 11B is a flow chart showing an example of a method 1100 for calculating a distance between locations of a medical image (e.g., an MPX image) that meet a first and second criteria, according to some aspects of the present disclosure. Referring to FIG. 11B, in block 1104, pixels in a blank tile that correspond to image locations that meet the first criterion are marked to generate a marked tile. In one non-limiting example, all pixels of the marked tile have a binary value of "0", except for the marked pixel that has a binary value of "1". The first criterion may be, for example, a first selected phenotype (or a first combination of selected phenotypes), and may be further limited to a particular tissue region in some cases. In one example, the blank tile is a blank result tile that corresponds to a result tile of the MPX image (i.e., each pixel of the blank tile corresponds to a pixel at the same position in the result tile of the MPX image).
[0063] At block 1108, a distance transform array is calculated for the marked tile. For each pixel in the marked tile, the distance transform array has a corresponding value indicating the distance from that pixel to the nearest marked pixel in the marked tile. At block 1112, distances corresponding to image locations that satisfy a second criterion are calculated. Values of the transformation array are selected and stored. The second criterion may be, for example, a second selected phenotype (or a second combination of selected phenotypes), possibly further limited to a particular tissue region. In this manner, an instance of method 1100 may be performed (e.g., in parallel) for each tile of interest (e.g., each result tile of an image, or each result tile within an annotation of an image), and the selected values for each tile are stored together (e.g., in a common hash table) for future processing (e.g., sorting in order of magnitude, histogram calculation, statistical analysis, etc., as described herein). Such processing of the sorted table may include, for example, using different custom input variables to calculate corresponding histograms and report associated spatial relationships. Such a process (e.g., including a tile-level instance of method 1100) may be repeated (e.g., in parallel) for different annotations of an image, and / or for different selected first and second criteria for the same or different annotations of an image, as often as necessary.
[0064] Values in the distance transform array near the edges of the array may be unreliable. For that reason, it may be desirable to ignore values that are within a predetermined number of elements from any edge of the array in block 1112. If the blank tile is a blank result tile that corresponds to a result tile of the MPX image, it may be desirable, for example, to limit the selection to values from within the portion of the distance transform array that corresponds to an interior region of the result tile (i.e., to limit the selection to values of the distance transform array that correspond to pixels within the interior region of the result tile), and it may also be desirable to configure the overlap between result tiles to be at least as large as the maximum closest distance recorded.
[0065] FIG. 12 illustrates the application of method 1100 to calculating, recording, and classifying distances between a first biomarker and a second biomarker. In this example, the first criterion is the expression of the CD8 biomarker and the second criterion is the expression of the PanCK biomarker. The array on the left side of FIG. 12 illustrates a portion of a marked result tile (e.g., as generated in block 1104) in which the position corresponding to the CD8 marker is indicated with a "1" (in this example, the array contains only one such position). The array on the right side of FIG. 11 illustrates a corresponding portion of a resulting distance transform array (e.g., as generated in block 1108) indicating the distance from each corresponding pixel of the image to the position of the CD8 marker. In this example, the position of the PanCK marker is also indicated in these two arrays (i.e., by three Xs). Values of the distance transform array corresponding to these three positions are selected and stored (e.g., in block 1112) to obtain, for each occurrence of the PanCK marker in this portion of the image, the distance to the closest occurrence of the CD8 marker. In this example, these distances are 4.2426, 6.0828, and 7.0711 pixels. Before these values are stored (e.g., in a hash table), it may be desirable to convert these values to actual distances (e.g., in microns) according to a known correspondence between image pixel size and physical dimensions. Figure 13 illustrates a similar application of method 1100 for calculating, recording, and classifying distances between a first biomarker and a second biomarker where the CD8 marker occurs at multiple locations.
[0066] Further levels of analysis ("tertiary analysis") may include statistical analysis of any regions of interest within the analyzed MPX slide. For example, it may be desirable to provide a readout of desired statistical results (e.g., as obtained by the example of method 1100) regarding cells of tissues contained in complex user annotations of MPX images, where the annotations typically consist of a combination of hand-drawn inclusion and exclusion regions of any shape and size. It may even be desirable to provide such results interactively (e.g., in real time).
[0067] This operation of retrieving information related to a given region of interest is called a "spatial query." To enable rapid collection of such information (e.g., to enable interaction), support for spatial queries may be implemented using hierarchical data structures such as quadtrees. Such an approach also enables the use of algorithms for efficient quadtree traversal and polygon clipping. Additionally or alternatively, it may be desirable to store image data and / or analysis results using an internal data order that enables efficient compression (e.g., using a Hilbert curve that maps a 2D image space into a 1D storage space for efficient querying and retrieval).
[0068] 14 is a flow chart illustrating an example of a method 1400 for analyzing an image of a tissue section including a plurality of pixels and showing a plurality of biological structures, according to some aspects of the present disclosure. Referring to FIG. 14, in block 1404, a plurality of image locations are obtained (e.g., via analysis of the image and / or from storage). In some cases, each of the image locations corresponds to a separate one of the plurality of biological structures and indicates a location of a description of the biological structure within the image. The image may be a WSI (e.g., an MPX WSI) or a portion (e.g., a tile) of such an image.
[0069] At block 1408, a first binary mask is obtained for the image (e.g., via analysis of the image and / or from storage). The first binary mask indicates a corresponding state of a first binary membership value for each of a plurality of pixels of the image. The first binary mask may include a plurality of pixel membership values and a plurality of microtile membership values, each of the plurality of pixel membership values corresponding to a separate one of the plurality of pixels and indicating a state of the first binary membership value for the pixel, and each of the plurality of microtile membership values corresponding to a separate one of the plurality of microtiles and indicating a state of the first binary membership value for all pixels in a block of the image corresponding to the microtile.
[0070] At block 1412, for each of the plurality of image locations and based on information from the first binary mask, a first binary membership value state of the pixel corresponding to the image location is stored in a data structure associated with the image location. In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a first binary marker value indicative of a positive state of a first biomarker in the corresponding biological structure.
[0071] In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a first binary marker value indicative of a positive state of the first biomarker in the corresponding biological structure. In such a case, the method 1400 may further include, for each of the plurality of overlapping tiles of the image, calculating, for each pixel of the tile, a distance transform array including a corresponding value indicative of a distance between the pixel and a closest one of the plurality of image locations for which the first binary marker value indicates the first positive state.
[0072] In some aspects, for each of the plurality of image locations, the data structure associated with the image location may include a second binary marker value indicative of a positive state of a second biomarker in the corresponding biological structure, and each of the plurality of tiles may include an inner region that does not overlap with an inner region of any other tile of the plurality of tiles. In such a case, the method 1400 may further include, for each of the inner regions, storing a value of the distance transform array of the corresponding tile that corresponds to the image location where the second binary marker value indicates the first positive state.
[0073] 15A is a flow chart illustrating an example of an implementation 1500 of a method 1400 for analyzing an image of a tissue section including a plurality of pixels and showing a plurality of biological structures, according to some embodiments of the present disclosure. 2 may be implementations of blocks 1404, 1408, and 1412, respectively, as described herein. At block 1516, a distance transform array is calculated for at least a portion of the image, with each value of the distance transform array corresponding to a different respective pixel of the image and indicating a distance between the pixel and a closest one of the plurality of image locations at which a first binary marker value indicates a first positive state. In some embodiments, for each of the plurality of image locations, the data structure associated with the image location may include a second binary marker value indicating a positive state of a second biomarker in the corresponding biological structure.
[0074] 15B is a flow chart illustrating an example implementation 1502 of a method 1500 for analyzing an image of a tissue section including a plurality of pixels and showing a plurality of biological structures according to some aspects of the present disclosure. With reference to FIG. 15, blocks 1504, 1508, 1512, and 1516 may be as described herein. At block 1520, values of the distance transform array corresponding to image locations where the second binary marker value indicates a first positive condition may be stored. In some aspects, the method may further include sorting the stored values in order of magnitude.
[0075] It should be appreciated that the particular blocks depicted in Figures 14, 15A, and 15B provide a particular method for analyzing an image of a tissue section including a plurality of pixels and showing a plurality of biological structures according to embodiments disclosed herein. Other sequences of such operations may be performed according to alternative embodiments. For example, alternative embodiments of such methods may perform the operations outlined above in a different order. Additionally, the individual blocks depicted in Figures 14, 15A, and 15B may include multiple sub-operations that may be performed in various orders as appropriate for the individual blocks. Furthermore, additional operations may be added or removed depending on the particular application. Those skilled in the art will recognize many variations, modifications, and alternatives. In any of the various aspects or implementations of these methods, such as those described above, each of the plurality of biological structures may be a cell nucleus. Additionally or alternatively, the image may be a multiplexed immunofluorescence (MPX) image having multiple channels (e.g., 3, 4, 5, 6, 7 or more, 32, or 64).
[0076] Any of methods 1400, 1500, and 1502 may further include obtaining a second binary mask of the image, the second binary mask indicating, for each of the plurality of pixels, a corresponding state of the second binary membership value. The second binary mask may include a plurality of pixel membership values and a plurality of microtile membership values, each of the plurality of pixel membership values corresponding to a separate one of the plurality of pixels and indicating a state of the second binary membership value for the pixel, and each of the plurality of microtile membership values corresponding to a separate one of the plurality of microtiles and indicating a state of the second binary membership value for all pixels in the block of the image corresponding to the microtile. In some aspects, such methods further include, for each of the plurality of image locations and based on information from the second binary mask, storing in a data structure associated with the image location a state of the second binary membership value of the pixel corresponding to the image location.
[0077] 16 is a flow chart illustrating an example of a method 1600 for analyzing an image of a tissue section including a plurality of pixels and showing a plurality of biological structures, according to some aspects of the present disclosure. Referring to FIG. 16, in block 1604, a plurality of seed locations in the image are obtained (e.g., via analysis of the image and / or from storage). In some cases, each of the image locations corresponds to a separate one of the plurality of biological structures and indicates the location of a description of the biological structure within the image. The image may be a WSI (e.g., MPX WSI) or a portion (e.g., a tile) of such an image.
[0078] A plurality of locations of a first biomarker in an image is obtained at block 1608. In some cases, the plurality of seed locations are obtained from a first channel of the image and the plurality of locations of the first biomarker are obtained from a second channel of the image.
[0079] At block 1612, a first distance transform array is calculated for at least a portion of the image including a plurality of seed locations, each value of the first distance transform array corresponding to a respective pixel of the plurality of pixels and indicating a distance from the pixel to a nearest one of the plurality of locations of the first biomarker.
[0080] At block 1616, for each of the plurality of seed locations and based on the information from the first distance transform sequence, it is detected whether a first biomarker is expressed at the seed location. At block 1620, for each of the plurality of seed locations, an indication of whether expression of the first biomarker was detected at the seed location is stored in a data structure associated with the seed location.
[0081] At block 1624, analysis results are provided that include detecting co-localization of the at least two phenotypes in at least a portion of the tissue section based on the stored indices. In some cases, detecting co-localization of the at least two phenotypes includes detecting that a first phenotype of the at least two phenotypes occurs within a predetermined vicinity of a second phenotype of the at least two phenotypes.
[0082] 17 is a flow chart illustrating an example of a method 1700 for analyzing an image of a tissue section including a plurality of pixels and showing a plurality of biological structures, according to some aspects of the present disclosure. Referring to FIG. 17, in block 1704, a plurality of seed locations in the image are obtained (e.g., via analysis of the image and / or from storage). In some cases, each of the image locations corresponds to a separate one of the plurality of biological structures and indicates the location of a description of the biological structure within the image. The image may be a WSI (e.g., MPX WSI) or a portion (e.g., a tile) of such an image.
[0083] At block 1708, a first sparse binary segmentation mask is obtained that includes a first tissue region of the tissue section and excludes a second tissue region of the tissue section. The first sparse binary segmentation mask includes a plurality of pixel membership values and a plurality of microtile membership values, and indicates a corresponding state of the first binary membership value for each of the plurality of pixels. Each of the plurality of pixel membership values corresponds to a respective pixel of the plurality of pixels and indicates a state of the first binary membership value for the pixel. Each of the plurality of microtile membership values corresponds to a respective microtile of the plurality of microtiles of the first binary mask and indicates a state of the first binary membership value for all pixels in a block of the image corresponding to the microtile.
[0084] At block 1712, for each of the plurality of seed locations and based on information from the first sparse binary segmentation mask, it is determined whether a state of the first binary membership value for a pixel of the plurality of pixels corresponding to the seed location is a first state or a second state. In some cases, determining whether a state of the first binary membership value for the corresponding pixel is a first state or a second state includes detecting that the first sparse binary segmentation mask does not include a pixel membership value for the pixel. At block 1716, for each of the plurality of seed locations, the state of the first binary membership value of the pixel is stored in a data structure associated with the seed location.
[0085] At block 1720, a cellular population of the first tissue region is determined based on the stored state. In some cases, the analysis results include a distribution density of at least one phenotype in the first tissue region. In some cases, the analysis results include a distribution of distances between the locations of the biomarkers in the first tissue region.
[0086] Methods 1100, 1400, 1500, 1502, 1600, and 1700 may each be embodied on a non-transitory computer-readable medium, such as, but not limited to, a memory or other non-transitory computer-readable medium known to one of skill in the art, that stores a program including computer-executable instructions for causing a processor, computer, or other programmable device to perform the operations of the method.
[0087] IV. Exemplary Systems for Automated Image Analysis 18 is a block diagram of an exemplary computing environment having an exemplary computing device suitable for use in some exemplary implementations, e.g., executing methods 1100, 1400, 1500, 1502, 1600 and / or 1700. The computing device 1805 in the computing environment 1800 may include one or more processing units, cores, or processors 1810, memory 1815 (e.g., RAM, ROM, etc.), internal storage 1820 (e.g., magnetic, optical, solid state storage and / or organic), and / or I / O interfaces 1825, any of which may be coupled over a communication mechanism or bus 1830 for communicating information or incorporated into the computing device 1805.
[0088] The computing device 1805 may be communicatively coupled to an input / user interface 1835 and an output device / interface 1840. Either or both of the input / user interface 1835 and the output device / interface 1840 may be wired or wireless interfaces and may be removable. The input / user interface 1835 may include any device, component, sensor, or interface, physical or virtual, that may be used to provide input (e.g., buttons, touch screen interfaces, keyboards, pointing / cursor control, microphones, cameras, Braille, motion sensors, optical readers, etc.). The output device / interface 1840 may include displays, televisions, monitors, printers, speakers, Braille, etc. In some exemplary implementations, the input / user interface 1835 and the output device / interface 1840 may be incorporated or physically coupled with the computing device 1805. In other exemplary implementations, other computing devices may function as or provide the functionality of the input / user interface 1835 and the output device / interface 1840 for the computing device 1805.
[0089] The computing device 1805 may be communicatively coupled (e.g., via I / O interface 1825) to an external storage device 1845 and to a network 1850 for communicating with any number of networked components, devices, and systems, including one or more computing devices of the same or different configurations. The computing device 1805 or connected computing devices may function as, provide services for, or be referred to as a server, client, thin server, general purpose machine, special purpose machine, or another label.
[0090] The I / O interface 1825 may be implemented using, but is not limited to, any communication or I / O protocol or standard (e.g., I / O protocol, I / O interface, I / O protocol, I / O interface card, I / O card, I / O card reader ... Network 1850 may include wired and / or wireless interfaces using protocols such as Ethernet, 802.11x, Universal System Bus, WiMax, modem, cellular network protocols, etc. Network 1850 may be any network or combination of networks (e.g., the Internet, a local area network, a wide area network, a telephone network, a cellular network, a satellite network, etc.).
[0091] The computing device 1805 can use and / or communicate using computer usable or computer readable media, including transitory and non-transitory media. Transitory media include transmission media (e.g., metallic cables, optical fibers), signals, carrier waves, and the like. Non-transitory media include magnetic media (e.g., disks and tapes), optical media (e.g., CD ROM, digital video disks, Blu-ray® disks), solid-state media (e.g., RAM, ROM, flash memory, solid-state storage), and other non-volatile storage or memory.
[0092] The computing device 1805 may be used to implement techniques, methods, applications, processes, or computer-executable instructions in some exemplary computing environments. The computer-executable instructions may be retrieved from a transitory medium, stored on a non-transitory medium and retrieved from the non-transitory medium. The executable instructions may be from one or more of any programming language, scripting language, and machine language (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).
[0093] The processor 1810 may execute under any operating system (OS) (not shown) in a native or virtual environment. One or more applications may be deployed, including a logic unit 1860, an application programming interface (API) unit 1865, an input unit 1870, an output unit 1875, a boundary mapping unit 1880, a control point determination unit 1885, a transformation calculation and application unit 1890, and an inter-unit communication mechanism 1895 for different units to communicate with each other, with the OS, and with other applications (not shown). For example, the binary mask processing unit 1880, the image position processing unit 1885, and the data structure processing unit 1890 may implement one or more processes described and / or shown in Figure 14, Figure 15A, and / or Figure 15B. The described units and elements may vary in their design, function, configuration, or implementation, and are not limited to the provided description.
[0094] In some example implementations, once information or execution instructions are received by the API unit 1865, it may be communicated to one or more other units (e.g., logic unit 1860, input unit 1870, output unit 1875, binary mask processing unit 1880, image location processing unit 1885, and data structure processing unit 1890). For example, after the input unit 1870 detects a user input, it may use the API unit 1865 to communicate the user input to the binary mask processing unit 1880 to obtain a first binary mask. The binary mask processing unit 1880 may interact with the image location processing unit 1885 via the API unit 1865 to determine a state of a first binary membership value of a pixel corresponding to the image location. Using the API unit 1865, the image location processing unit 1885 may interact with the data structure processing unit 1890 to store a state of a first binary membership value of a pixel corresponding to the image location in a data structure associated with the image location. Further example implementations of applications that may be deployed may include a distance transform array calculation unit for calculating the distance transform arrays described herein (eg, with reference to FIG. 11B).
[0095] In some cases, logic unit 1860 may be configured to control information flow between units and direct services provided by API unit 1865, input unit 1870, output unit 1875, binary mask processing unit 1880, image position processing unit 1885, and data structure processing unit 1890 in some example implementations described above. For example, the flow of one or more processes or implementations may be controlled solely by logic unit 1860 or in cooperation with API unit 1865.
[0096] For one or more embodiments, at least one of the components depicted in one or more of the preceding figures may be configured to perform one or more of the operations, techniques, processes, or methods described in the following examples and claims presented below. EXAMPLES
[0097] V. Working Examples In the following sections, further exemplary embodiments are provided.
[0098] Example 1 is a method for analyzing an image of a tissue section including a plurality of pixels and indicative of a plurality of biological structures, comprising acquiring a plurality of image locations, each of the image locations corresponding to a separate one of the plurality of biological structures and indicative of a location of a description of the biological structures within the image, and acquiring a first binary mask for the image, the first binary mask indicating, for each of a plurality of pixels of the image, a corresponding state of a first binary membership value, the first binary mask including a plurality of pixel membership values and a plurality of microtile membership values, each of the plurality of pixel membership values being indicative of a corresponding state of a first binary membership value within the image. and obtaining a first binary mask for the image, each of a plurality of microtile membership values corresponding to a distinct one of the plurality of microtiles and indicating a state of the first binary membership value for the pixel, wherein each of the plurality of microtile membership values corresponds to a distinct one of the plurality of microtiles and indicates a state of the first binary membership value for all pixels in the block of the image corresponding to the microtile, the method further including, for each of the plurality of image locations and based on information from the first binary mask, storing in a data structure associated with the image location a state of the first binary membership value of the pixel corresponding to the image location.
[0099] Example 2 includes a method as described in Example 1 or any other example herein, wherein, for each of a plurality of image locations, a data structure associated with the image location includes a first binary marker value indicative of a positive status of a first biomarker in the corresponding biological structure.
[0100] Example 3 includes a method as described in Example 2 or any other example herein, where the method further includes calculating a distance transform array for at least a portion of the image, where each value in the distance transform array corresponds to a different respective pixel of the image and indicates a distance between the pixel and a nearest image location among a plurality of image locations where the first binary marker value indicates a first positive state.
[0101] Example 4 includes a method as described in Example 3 or any other example herein, wherein, for each of the plurality of image locations, the data structure associated with the image location includes a second binary marker value indicating a positive status of a second biomarker in the corresponding biological structure.
[0102] Example 5 includes a method as described in example 4 or any other example herein, wherein the method further includes storing values of the distance transform array corresponding to image locations where the second binary marker value indicates the first positive state.
[0103] Example 6 includes the method of example 5 or any other example herein, where the method further includes sorting the stored values in order of magnitude.
[0104] Example 7 includes a method as described in Example 1 or any other example herein, where, for each of the plurality of image locations, the data structure associated with the image location includes a first binary marker value indicating a positive state of a first biomarker in the corresponding biological structure, and the method further includes, for each of a plurality of overlapping tiles of the image, calculating, for each pixel of the tile, a distance transform array including a corresponding value indicating the distance between the pixel and a closest one of the plurality of image locations for which the first binary marker value indicates the first positive state.
[0105] Example 8 includes a method as described in Example 7 or any other example herein, wherein, for each of the plurality of image locations, the data structure associated with the image location includes a second binary marker value indicating a positive state of a second biomarker in the corresponding biological structure, and each of the plurality of tiles includes an inner region that does not overlap with an inner region of any other tile of the plurality of tiles, and the method further includes, for each inner region, storing a value of the distance transform array of the corresponding tile that corresponds to the image location where the second binary marker value indicates the first positive state.
[0106] Example 9 includes the method of any of Examples 1-8 or any other example herein, wherein each of the plurality of biological structures is a cell nucleus.
[0107] Example 10 includes the method of any of Examples 1-8 or any other example herein, where the image is a multiplexed immunofluorescence image having multiple channels.
[0108] Example 11 includes a system comprising one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations described in any of Examples 1 to 10 or other examples herein.
[0109] Example 12 is a non-transitory computer-readable medium having stored thereon instructions for causing one or more processors to execute a method for analyzing an image of a tissue section including a plurality of pixels and showing a plurality of biological structures, the processor-executable instructions including: acquiring a plurality of image locations, each of the image locations corresponding to a separate one of the plurality of biological structures and indicating a location of a description of the biological structures within the image; and acquiring a first binary mask for the image, the first binary mask indicating, for each of a plurality of pixels of the image, a corresponding state of a first binary membership value, the first binary mask indicating a corresponding state of a first binary membership value for each of the plurality of pixels of the image, the first binary mask indicating a corresponding state of a first binary membership value for each of the plurality of pixels of the image. and obtaining a first binary mask for the image, the first binary mask including a plurality of pixel membership values and a plurality of microtile membership values, each of the plurality of pixel membership values corresponding to a distinct one of the plurality of pixels and indicating a state of a first binary membership value for the pixel, each of the plurality of microtile membership values corresponding to a distinct one of the plurality of microtiles and indicating a state of a first binary membership value for all pixels in a block of the image corresponding to the microtile, the operations including obtaining, for each of the plurality of image locations and based on information from the first binary mask, a first binary mask for the image including a plurality of microtile membership values corresponding to the image location. The method further comprises storing a state of the first binary membership value of the pixel in a data structure associated with the image location.
[0110] Example 13 includes a non-transitory computer-readable medium as described in Example 12 or any other example herein, wherein, for each of a plurality of image locations, the data structure associated with the image location includes a first binary marker value indicative of a positive status of a first biomarker in the corresponding biological structure.
[0111] Example 14 includes the non-transitory computer-readable medium of Example 13 or any other example herein, further including instructions for performing an operation including calculating a distance transform array for at least a portion of the image, where each value in the distance transform array corresponds to a different respective pixel of the image and indicates a distance between the pixel and a closest one of a plurality of image locations where a first binary marker value indicates a first positive state.
[0112] Example 15 includes a non-transitory computer-readable medium as described in Example 14 or any other example herein, wherein, for each of the plurality of image locations, the data structure associated with the image location includes a second binary marker value indicative of a positive status of a second biomarker in the corresponding biological structure.
[0113] Example 16 includes the non-transitory computer-readable medium of Example 15 or any other example herein, further including instructions for performing an operation including storing values of the distance transform array corresponding to image locations where the second binary marker value indicates a first positive state.
[0114] Example 17 includes the non-transitory computer-readable medium of Example 16 or any other example herein, further including instructions for performing operations including sorting the stored values in order of magnitude.
[0115] Example 18 includes the non-transitory computer-readable medium of Example 12 or any other example herein, wherein for each of the plurality of image locations, the data structure associated with the image location includes a first binary marker value indicating a positive state of a first biomarker in a corresponding biological structure, and the medium further includes instructions for performing an operation including: for each of a plurality of overlapping tiles of the image, calculating, for each pixel of the tile, a distance transform array including a corresponding value indicating the distance between the pixel and a closest one of the plurality of image locations where the first binary marker value indicates a first positive state.
[0116] Example 19 includes a non-transitory computer-readable medium as described in Example 18 or any other example herein, wherein for each of the plurality of image locations, the data structure associated with the image location includes a second binary marker value indicating a positive state of a second biomarker in the corresponding biological structure, and each of the plurality of tiles includes an inner region that does not overlap with an inner region of any other tile of the plurality of tiles, and the medium further includes instructions for performing an operation including storing, for each inner region, a value of a distance transform array of the corresponding tile that corresponds to the image location where the second binary marker value indicates a first positive state.
[0117] Example 20 includes the non-transitory computer-readable medium of any of Examples 12-19 or any other example herein, wherein each of the plurality of biological structures is a cell nucleus.
[0118] VI. Further Considerations FIG. 19 shows an example of annotated MPX image of a small multi-tissue biopsy, and FIG. 20 and FIG. 21 show example area, density distribution, and distance distribution results that may be obtained from such an image using implementations of methods 1100, 1300, 1400, and / or 1402 described herein. Using the existing framework, it took 2.24 seconds to report the density distribution and spatial relationships of 132,766 4',6-diamidino-2-phenylindole (DAPI) stained nuclear cells in tumor, epithelial tumor, and stroma, as shown in FIG. 20 and FIG. 21 (area and density distribution of DAPI cells; spatial relationships between DAPI nuclei and tumor, stroma, and epithelial tumor regions; respectively).
[0119] FIG. 22 shows an example of annotated MPX image of a large tissue section, and FIG. 23 and FIG. 24 show example area, density distribution, and distance distribution results that may be obtained from such an image using implementations of methods 1100, 1300, 1400, and / or 1402 described herein. Using the techniques disclosed herein, it took only 1.13 seconds to report the density distribution and spatial characteristics of 630,276 DAPI nucleated cells, as shown in FIG. 23 and FIG. 24 (area and density distribution of DAPI cells; spatial relationship between DAPI nuclei and tumor, stromal, and epithelial tumor regions; respectively), while the existing framework took 231.49 seconds (e.g., the techniques disclosed herein provided results 205 times faster than the existing framework).
[0120] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the methods and / or some or all of the processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause the one or more data processors to perform some or all of the methods and / or some or all of the processes disclosed herein.
[0121] The terms and expressions used are used as terms of description and not of limitation, and in the use of such terms and expressions there is no intention to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the invention as claimed has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.
[0122] The above description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the above description of preferred exemplary embodiments provides those skilled in the art with an effective description for implementing various embodiments. It will be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.
[0123] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, and other components may be shown in block diagram form in order to avoid obscuring the embodiments. Processes, algorithms, structures, and techniques may be shown without unnecessary detail.
Claims
1. It is a method, A step of accessing an image depicting multiple biomarkers, wherein the image includes multiple tiles, The step of selecting a tile from the aforementioned plurality of tiles, The steps include processing the selected tile to mark one or more first positions corresponding to one or more first biomarkers among the plurality of biomarkers, and marking one or more second positions corresponding to one or more second biomarkers among the plurality of biomarkers, A step of calculating an array of distance transformations for the tile, wherein the array of distance transformations includes a set of values indicating the distance between one or more first positions corresponding to one or more first biomarkers and one or more second positions corresponding to one or more second biomarkers. The steps include selecting a subset of values in the set of values such that the selected phenotype satisfies predetermined criteria, A step of storing a subset of the aforementioned values in memory, Methods that include...
2. In the method according to claim 1, A method wherein each tile of the plurality of tiles overlaps with one or more other tiles in the plurality of tiles.
3. In the method according to claim 1, A method wherein each tile in the plurality of tiles is associated with a binary mask that indicates one or more locations of one or more biomarkers among the plurality of biomarkers.
4. In the method according to claim 1, A method comprising the step of processing the tile to mark one or more first positions corresponding to one or more first biomarkers and one or more second positions corresponding to one or more second biomarkers, wherein the step of dividing the binary mask associated with the tile into a plurality of non-overlapping cells.
5. In the method according to claim 4, A method comprising the step of processing the tiles to mark one or more first positions corresponding to one or more first biomarkers, and marking one or more second positions corresponding to one or more second biomarkers, further comprising the step of classifying each cell in the plurality of non-overlapping cells into one of a plurality of classes.
6. In the method according to claim 1, A method wherein the step of selecting a subset of values in the set of values that satisfy the predetermined criteria includes the step of selecting values that are in the set of values corresponding to phenotypic colocalization.
7. It is a system, One or more data processors, One or more non-temporary computer-readable storage media storing instructions that cause the system to perform the method according to any one of claims 1 to 6 when executed on the one or more data processors, A system that includes these features.
8. A computer program for causing one or more data processors to perform the method described in any one of claims 1 to 6.