Scalable and high precision context-guided segmentation of histological structures including ducts / glands and lumen, cluster of ducts / glands, and individual nuclei in whole slide images of tissue samples from spatial multi-parameter cellular and intracellular imaging platforms
The method and system enhance digital pathology by using Gaussian multiscale pyramid decomposition and machine learning for precise segmentation of histological structures, addressing subjectivity and inefficiencies in current workflows.
Patent Information
- Application Number
- JP2025035163
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-03-16
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-03-16
AI Technical Summary
Current digital pathology workflows for diagnosing diseases based on histopathological structures are subjective, time-consuming, and prone to errors due to manual evaluation of large amounts of patient data, leading to high disagreement in atypical cases.
A method and system for segmenting histological structures using multi-parameter cell and intracellular imaging data, involving Gaussian multiscale pyramid decomposition, superpixel division, machine learning algorithms for probability assignment, and region-based active contour algorithms to refine boundaries.
Enables scalable and high-precision segmentation of histological structures like ducts/glands, clusters, and nuclei, reducing subjectivity and improving diagnostic accuracy in digital pathology.
Smart Images

Figure 2025098060000001_ABST
Abstract
Description
Technical Field
[0001] <Government Contract> This invention was made with government support under grant #CA204826 awarded by the National Institutes of Health (NIH). The government has certain rights in this invention.
[0002] <Technical Field> The present invention relates to digital pathology, and more particularly, to scalable and high-precision context-guided segmentation of histological structures including, but not limited to, ducts / glands and lumens, clusters of ducts / glands, and individual nuclei, in multi-parameter cellular and sub-cellular imaging data of several stained tissue images such as whole slide images obtained from several patients or several multicellular in vitro models.
Background Art
[0003] The histopathological examination of diseased tissue is essential for the diagnosis and grading of diseases. Currently, pathologists usually make diagnostic decisions (e.g., the malignancy or severity of a disease) based on the visual interpretation of histopathological structures in transmitted light images of diseased tissue. Such decisions are mostly subjective, and in particular, a high level of disagreement occurs in atypical situations.
[0004] In addition, digital pathology is gaining momentum in applications such as second opinion telepathology, interpretation of immunohistochemistry, and intraoperative telepathology. Usually, digital pathology consists of several tissue slides, a large amount of patient data representing them is generated, the slides are displayed on a high-resolution monitor, and evaluated by a pathologist. Due to the inclusion of manual work, the current workflow practice in digital pathology is time-consuming, error-prone, and subjective.
Summary of the Invention
[0005] In one embodiment, a method is provided for segmenting one or more histological structures in a tissue image represented by multi-parameter cell and intracellular imaging data. The method includes receiving the coarsest level of image data of the tissue image, where the coarsest level of image data corresponds to the coarsest level of a multi-scale representation of first data corresponding to the multi-parameter cell and intracellular imaging data. The method further includes dividing the coarsest level of image data into a plurality of non-overlapping superpixels, assigning to each superpixel a probability of belonging to one or more histological structures using some pre-trained machine learning algorithms to create a probability map, applying a contour algorithm to the probability map to extract an estimated boundary of the one or more histological structures, and using the estimated boundary to yield an accurate boundary of the one or more histological structures. In one exemplary embodiment, the multi-scale representation includes Gaussian multiscale pyramid decomposition, the multi-parameter cell and intracellular imaging data includes stained tissue image data, receiving the coarsest level of image data of the tissue image includes receiving the coarsest level of normalized constituent stain image data of the stained tissue image, the coarsest level of normalized constituent stain image data pertains to a particular constituent stain of the stained tissue image and corresponds to the coarsest level of a Gaussian multiscale pyramid decomposition of first data corresponding to the stained tissue image data, and dividing the coarsest level of image data into a plurality of superpixels includes dividing the coarsest level of normalized constituent stain image data into a plurality of superpixels.
[0006] In one embodiment, a computer system is provided for segmenting one or more histological structures in a tissue image represented by multi-parameter cell and intracellular imaging data. The system includes a processing device, and the processing device comprises several components configured to implement the above method.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Best Mode for Carrying Out the Invention
[0008] As used herein, the singular forms "a", "an", and "the" include plural references unless the context clearly dictates otherwise.
[0009] As used herein, the recitation that two or more parts or components are "coupled" together means that the parts are directly or indirectly coupled, i.e., coupled through one or more intermediate parts or components, or operate together, as long as a connection occurs.
[0010] As used herein, the term "some" means one or an integer greater than one (i.e., a plurality).
[0011] As used herein, the terms "component" and "system" are intended to refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution. For example, a component may be a process running on a processor, a processor, an object, an executable file, a thread of execution, a program, and / or a computer, and is not limited thereto. For example, both an application running on a server and the server may be considered components. One or more components may reside within a process and / or thread of execution, and a component may be localized on one computer and / or distributed between two or more computers. Although several ways of displaying information to a user are shown and described herein as screenshots in specific figures or graphs, those of ordinary skill in the relevant art will recognize that various other alternative means may be employed.
[0012] As used herein, the term "multi-parameter cellular and intracellular imaging data" refers to data obtained by generating several images from a section of tissue, which provides information regarding a plurality of measurable parameters at the cellular level and / or intracellular level in the section of tissue. The multi-parameter cellular and intracellular imaging data may be created by any of several different imaging modalities, such as, but not limited to, transmitted light (e.g., a combination of H&E and / or IHC (one or more biomarkers)), fluorescence, immunofluorescence (including, but not limited to, antibodies, nanobodies), multiplexing and / or hyper-multiplexing of live cell biomarkers, and electron microscopy. Targets include, but are not limited to, tissue samples (human or animal) and in vitro models of tissues and organs (human or animal).
[0013] As used herein, the term "superpixel" refers to a connected patch or group of two or more pixels having similar image statistics defined in a suitable color space (e.g., RGB, CIELAB, or HSV).
[0014] As used herein, the term "non-overlapping superpixel" refers to a superpixel whose boundaries do not overlap with any of its neighboring superpixels.
[0015] As used herein, "Gaussian multi-scale pyramid decomposition" means repeatedly applying smoothing and subsampling by a factor of two in the x and y directions to an image using a Gaussian filter.
[0016] As used herein, the term "region-based active contour algorithm" refers to any active contour model that takes into account the image gradient to detect the boundaries of an object.
[0017] As used herein, the term "context-ML model" refers to a machine learning algorithm that can take into account the neighborhood information of superpixels.
[0018] As used herein, the term "staining-ML model" refers to a machine learning algorithm that can take into account the staining intensity of superpixels.
[0019] As used herein, the term "probability map" is meant to refer to a set of pixels having probability values in the range from 0 to 1, where these probability values represent the positional probability of whether a pixel is within a particular histological structure.
[0020] For example, terms related to directions used herein, such as up, down, left, right, upper, lower, front, back, and their derivatives, relate to the directions of the elements shown in the drawings and do not limit the scope of the claims unless explicitly stated.
[0021] Hereinafter, the disclosed concepts will be described in terms of many specific details for the purpose of explanation to provide a complete understanding of the innovation that is the subject matter. However, it will be apparent that the disclosed concepts can be practiced without these specific details without departing from the spirit and scope of the present invention.
[0022] The concepts further detailed and disclosed herein in connection with various exemplary embodiments provide a novel approach for identifying and characterizing morphological features of histopathological structures. An initial application example of such a tool is to perform scalable and high-precision context-guided segmentation of histological structures, including, for example, but not limited to, ducts / glands and lumens, clusters of ducts / glands, and individual nuclei, in images of tissue samples (e.g., whole slide images), based on spatial multi-parameter cell / intracellular imaging data representing such images. In this particular non-limiting application of the disclosed concept, as described in more detail herein, hematoxylin and eosin (H&E) image data is employed as the multi-parameter cell / intracellular imaging data, and color deconvolved hematoxylin image data is used to distinguish ducts / glands and lumens, clusters of ducts / glands, and individual nuclei. However, this is merely an example, and it will be understood that the disclosed concept may be employed to segment other histological structures using other types of data. For example, connective tissue may be segmented using color deconvolved eosin image data obtained from H&E image data. Further possibilities are contemplated within the scope of the disclosed concept.
[0023] The disclosed concepts are related to and improve upon the subject matter described in U.S. Patent Application No. 15 / 577,838, published as 2018 / 0204085, "Systems and Methods for Finding Regions of Interest in Hematoxylin and Eosin (H&E) Stained Tissue Images and Quantifying Intratumor Cellular Spatial Heterogeneity In Multiplexed / Hyperplexed Fluorescence Tissue Images", the disclosure of which is incorporated herein by reference. The disclosed concepts differ from the subject matter of the above application in at least two respects. First, the disclosed concepts are classified into the semi-supervised or weakly-supervised category in that user input exists in at least one step of the machine learning algorithm. Also, the disclosed concepts function optimally when the boundaries of the region of interest (ROI) are given as a rough estimate. The concepts disclosed and described in detail herein sharpen such rough boundaries.
[0024] FIG. 1 is a schematic diagram of an exemplary digital pathology system 5 constructed and configured to automatically segment histological structures from multiparameter cell-intracellular imaging data, based on an exemplary embodiment of the concepts disclosed herein. As seen in FIG. 1, system 5 is a computer device constructed and configured to generate and / or receive multiparameter cell-intracellular imaging data (reference numeral 25 in FIG. 1) and to process that data as described herein to segment histological structures within the tissue image represented by the multiparameter cell-intracellular imaging data 25. System 5 may be, for example, a PC, a laptop computer, a tablet computer, or other suitable computer device constructed and configured to perform the functions described herein, but is not limited thereto.
[0025] System 5 includes an input device 10 (such as a keyboard), a display 15 (such as an LCD), and a processing device 20. A user can provide an input to the processing device 20 using the input device 10, and the processing device 20 provides an output signal to the display 15 to enable the display 15 to display information to the user as detailed herein. The processing device 20 includes a processor and a memory. The processor is, for example, but not limited to, a microprocessor (μP), a microcontroller, an application-specific integrated circuit (ASIC), or other suitable processing device that interfaces with the memory. The memory, in the case of data storage such as the internal storage area of a computer, can be one or more of various types of internal and / or external storage media such as RAM, ROM, EPROM, EEPROM, FLASH (registered trademark), and others that provide storage registers, and can be volatile memory or non-volatile memory. The memory stores several routines executable by the processor and includes routines for implementing the disclosed concepts as described herein. In particular, the processing device 20 includes a histological structure segmentation component 30, and the histological structure segmentation component 30 is configured to identify and segment histological structures (such as, but not limited to, tubes / glands and lumens, clusters of tubes / glands, and individual nuclei, etc.) in several tissue images (such as H&E stained image data) represented by multi-parameter cell / intracellular imaging data 25 obtained from various imaging modalities as described herein in various embodiments.
[0026] Figures 2A and 2B are flowcharts showing a method of segmenting histological structures according to certain exemplary embodiments of the disclosed concept. The method shown in FIGS. 2A and 2B may be implemented, for example but not limited to, in the system 5 of FIG. 1 described above, and the method is described as such for illustrative purposes. Additionally, in the particular non-limiting exemplary embodiments shown in FIGS. 2A and 2B, the multi-parameter cell / intracellular imaging data used is H&E stained image data of a tissue sample, the histological structures to be segmented are based on hematoxylin image data, and include tubes / glands and lumens, clusters of tubes / glands, and individual nuclei. Again, the particular embodiments shown in FIGS. 2A and 2B and described herein are intended to be merely illustrative and are understood not to be limiting. It will be understood that other types of multi-parameter cell / intracellular imaging data may be used in connection with the disclosed concept.
[0027] Referring to FIG. 2A, the method begins at step 100, where the processing device 20 of the system 5 generates and / or receives multi-parameter cell / intracellular imaging data representing an H&E stained tissue image to be processed. A non-limiting exemplary H&E stained tissue image 35 that may be processed by the disclosed concept is shown in FIG. 3 for illustrative purposes. Additionally, in the exemplary embodiment, the multi-parameter cell / intracellular imaging data generated and / or received at step 100 is in RGB format.
[0028] Next, in step 105, the multi-parameter cell and intracellular imaging data (i.e., H&E stained tissue image data in RGB format) generated and / or received in step 100 is color deconvolved into individual staining intensities (hematoxylin and eosin), and hematoxylin image data and eosin image data are created for the processed H&E stained tissue image. FIG. 3 illustrates the color deconvolution of step 105 by showing a hematoxylin image 40 represented by the hematoxylin image data and a resulting eosin image 45 represented by the eosin image data resulting from the color deconvolution of the H&E stained tissue image 35.
[0029] Then, the method proceeds to step 110 where the staining intensity of the hematoxylin image data is normalized with a reference dataset to generate normalized hematoxylin image data. The normalization of the staining intensity in step 110 is performed such that the variation in the staining intensity is standardized for downstream processing. In an exemplary embodiment, to set the reference dataset, first, a batch of whole slide images (WSIs) is color deconvolved into hematoxylin staining intensity images and eosin staining intensity images. From this batch, a random number of 1Kx1K images are trimmed and used to create a cumulative intensity histogram of the hematoxylin channel. The test WSI first undergoes a color deconvolution operation. Next, histogram equalization is performed to match the intensity histogram of the hematoxylin channel with the histogram of the reference dataset. The staining intensity normalization of step 110 is shown in FIG. 4. FIG. 4 shows the original hematoxylin channel 50 of another (different) exemplary whole slide image and the normalized hematoxylin channel 55 of the same exemplary whole slide image. As a result, the intensity of the normalized hematoxylin channel 55 now matches that of the reference image dataset.
[0030] Next, in step 115, Gaussian multi-scale pyramid decomposition (a form of pyramid representation) is performed on the normalized hematoxylin image data to generate a multi-scale representation of the normalized hematoxylin image data. The multi-scale representation created in step 115 includes n levels, L1...L n where L1 is level data representing the full-resolution level of the decomposition, and L n is level data representing the coarsest level of the decomposition. By constructing a Gaussian pyramid, the computational load can be reduced when detecting histological structures from whole-slide images according to the method of the disclosed concept. FIG. 5 shows the Gaussian multi-scale pyramid decomposition of the exemplary normalized hematoxylin channel 55 shown in FIG. 4. In the exemplary embodiment, the size of the image is reduced from the original resolution of 30K×50K to 1K×1.5K at the coarsest level. Also, in the exemplary embodiment, the size of the image is halved at each level of the decomposition hierarchy.
[0031] Thereafter, the method proceeds to step 120 of FIG. 2B. In step 120, the data L at the coarsest level n is divided into non-overlapping superpixels, which are sets of connected pixels having similar intensity (gray) values in the exemplary embodiment. As will be appreciated, this can be done in several ways. In the simplest approach, a normal distribution of noise is assumed with a mean of zero and a standard deviation of sigma. For example, in the case of an image with 256 gray levels per pixel, the standard deviation of the noise is typically assumed to be 4 gray levels. This value may be set by the end user. In the exemplary embodiment, this is done using a simple linear iterative clustering (SLIC) algorithm, some of which are known in the art. FIG. 6 shows the data L at the coarsest level obtained from the exemplary hematoxylin image 40 described herein nShows the result of step 120 when executed against. In an exemplary embodiment, the image is segmented into superpixels of about 5K. This is recommended for 1K×1K hematoxylin channel images because the calculations are fast and effective for downstream processing. Additionally, in an exemplary embodiment, Delaunay triangulation of the superpixel centroids is performed to identify spatial neighbors for each superpixel.
[0032] Additionally, according to one aspect of the disclosed concept, in an exemplary embodiment, several machine learning algorithms / models are trained to predict superpixels belonging to a particular histological structure that is a tube / gland. Thus, following step 120, the method proceeds to step 125, where each superpixel is assigned a probability of belonging to a histological structure such as a tube / gland in the illustrated exemplary embodiment using several pre-trained machine learning algorithms. As a result, a probability map of the coarsest level of data L n is created.
[0033] In a non-limiting exemplary embodiment of the disclosed concept, several trained machine learning algorithms employed in step 125 include a context-ML model (such as a context-support vector machine (SVM) model or a context-logistic regression (LR) model) and a stain-ML model (such as a stain-support vector machine (SVM) model or a stain-logistic regression (LR) model) for predicting superpixels belonging to the structure in question. In this exemplary embodiment, the RGB color histograms of the superpixels and their neighbors are used as feature vectors. Specifically, in this embodiment, two models, namely, a context-ML model and a stain-ML model, are constructed and trained (in a supervised manner). Each of these models will be described in more detail below.
[0034] For an example embodiment's context-ML model, such as the context-SVM model, for the training set of the example embodiment, 2000 adjacent superpixel pairs in 10 different images from the reference image dataset are randomly selected. The superpixel pairs are displayed on a screen, and ground truth is collected (i.e., user input is requested) by asking a subject (e.g., an experienced / specialist pathologist) whether none of the displayed superpixels belong to a tube, or whether one or both belong to a tube. For illustrative purposes, FIG. 7 shows three such exemplary superpixel pairs of an exemplary image (labeled A, B, C), where for the pair labeled A, the number of superpixels present in the tube is 0 (class label 0), for the pair labeled B, there is 1 superpixel in the tube (class label 1), and for the pair labeled C, there are 2 superpixels in the tube (class label 2). This ground truth is used for training the context ML model as class labels. In an example embodiment, as feature vectors, the color histograms of each superpixel pair and their first neighbors (i.e., the pixel values of the R, G, B colors) are used. For each test superpixel pair, the ML model returns the probability that the pair has superpixels that do not both belong to a tube, or that one or both belong to a tube. In each case, the superpixel pair is assigned to the category with the highest probability. Note that this does not determine the actual identity of the superpixels within the structure. Instead, a second model, the stain-ML described below, is applied for that purpose.
[0035] In an exemplary embodiment, for a staining-ML model such as the staining-SVM model, ground truth is collected separately. In particular, a subject (e.g., an experienced / specialist pathologist) is asked to classify whether a particular superpixel has any of "no staining", "light staining", "moderate staining", "dark staining", or "unknown" (i.e., again, user input is required). Since some structures such as tubes are amorphous, one way to detect the boundaries of the structure is to carefully observe the change in staining color while moving from the inside of the tube to the surrounding connective tissue. Using this information, superpixels that may be part of the tube can be identified. FIG. 7 shows four such superpixels D, E, F, G classified as "no staining", "light staining", "moderate staining", and "dark staining" respectively for an exemplary image. The superpixel is assigned to the category with the highest probability.
[0036] Thus, in step 125 according to this particular exemplary embodiment, the above two machine learning models (context-ML model and staining-ML model) are sequentially applied to the non-overlapping superpixels of the coarsest level of data L n (step 120), and a probability map is created to identify pairs of superpixels that are likely to be inside the structure (a tube in this exemplary embodiment) in question. In an exemplary embodiment, all superpixels stained moderately to darkly are identified as being inside the tube. In other words, the context-ML model and the staining-ML model together assign a conditional probability belonging to the structure in question to each superpixel.
[0037] When the probability map is created as described above in step 125, the method proceeds to step 130. In step 130, a rough estimate of the boundaries of the histological structure is extracted by applying a region-based active contour algorithm to the probability map. An exemplary probability map 60 and an exemplary image 65 showing the application of the region-based active contour algorithm are provided in FIG. 8.
[0038] Next, the method proceeds to step 135, where the newly obtained estimate is used to effect segmentation of the structures in the full-resolution image. Specifically, step 135 involves sequentially refining the boundaries of the histological structures from coarse to fine by upsampling the structure boundaries from level K+1 to level K and then starting region-based active contours at level K using the upsampled boundaries. In an exemplary embodiment, upsampling first involves simply doubling the coordinates known at K+1 to determine the coordinates of the boundaries at level K+1 to level K, and then interpolating between the boundary pixels at level K.
[0039] Thus, in the exemplary method shown in FIGS. 2A and 2B, superpixels are utilized only at the coarsest level of the pyramid. As explained, region-based contours are run against the probability map of the identified superpixels. However, at successive finer scales, superpixels are not required and region-based active contours are run directly on the associated stain-separated image.
[0040] In one particular implementation of this exemplary embodiment, the region-based active contour algorithm used is the Chan-Vese segmentation algorithm that separates the foreground (tubes) from the background (the rest of the image). The cost function of the active contour varies by the difference in the average value of the hematoxylin stain in the foreground and background regions. For example, two superpixels that are likely to be inside a tube have approximately the same stain (medium to dark), and their boundaries are repeatedly merged by active contour optimization.
[0041] To construct a "cluster" of tubes, a region-based active contour may be executed on the probability maps returned by the context-ML model and the staining-ML model. In an exemplary embodiment, the probability map attributes a non-zero probability to regions that bridge tubes, and the region-based active contour model executed on the probability map may better depict a cluster of tubes. To segment tubes from the entire WSI, first, tubes and clusters of tubes are identified from the lowest-resolution pyramid image using the steps described above. These results are recursively upsampled, and the region-based active contour is re-executed at each level of the hierarchy to refine the upsampled tube boundaries. In an exemplary embodiment, the active contour image consists of a mask indicating pixels inside and outside the tube boundary. The active contour image and the hematoxylin image are upsampled together.
[0042] The disclosed concepts may be used to identify nuclei within tubes. Once a tube is identified, superpixel segmentation is executed on the region belonging to the tube. Next, the staining-ML model may be executed to further separate moderately stained superpixels and darkly stained superpixels within the tube. The darkly stained superpixels will likely correspond to the positions of nuclei within the tube. To identify nuclei outside the tube, a similar model may be developed that constructs a feature vector (histograms of each of the red, blue, and green channels) of superpixels that do not include first-layer neighbors. Without the average histogram of the superpixels and its first layer, all darkly stained superpixels corresponding to nuclei inside and outside the tube would be identified.
[0043] In the exemplary embodiments described in connection with FIGS. 2A and 2B, a color deconvolution step (105) and a staining intensity normalization step (110) are performed before the Gaussian multi-scale pyramid decomposition is performed (step 115). In another embodiment, the order of these steps is reversed. In particular, in an alternative embodiment of the disclosed concept, after receiving multi-parameter cell / intracellular imaging data (step 100), first, a Gaussian multi-scale pyramid decomposition is performed on the entire H&E stained tissue image. Next, the color deconvolution of step 105 and the normalization of the staining intensity of step 110 are performed only on the image data at the coarsest level, and normalized hematoxylin image data is generated only for the coarsest level image. Thereafter, steps 120 to 130 are performed using only the normalized hematoxylin image data for the coarsest level image, and a probability map and an estimated boundary of the histological structure are generated as described above. Thereafter, the boundary is refined sequentially using step 135 as described.
[0044] Here again, while certain embodiments of the disclosed concept use color deconvolved hematoxylin image data to distinguish tubes / glands and lumens, clusters of tubes / glands, and individual nuclei, this is meant to be illustrative only, and it is understood that the disclosed concept can be employed using other types of data to distinguish other histological structures. For example, but not limited to, connective tissue may be segmented using color deconvolved eosin image data (different from the color deconvolved hematoxylin image data) obtained from H&E image data. Still other possibilities are contemplated within the scope of the disclosed concept.
[0045] Furthermore, the description of the disclosed concepts above is based on and utilizes in situ multi-parameter cell and intracellular imaging data. However, it will be understood that this is not meant to be limiting. Rather, it will be understood that the disclosed concepts can be used in conjunction with in vitro microphysiological models for basic research and clinical translation. Multicellular in vitro models summarize the spatiotemporal cellular heterogeneity and heterocellular communication of human tissues applicable to the investigation of the mechanisms of disease progression in vitro, the testing of drugs, and the structural composition and content characterization of these models for use in transplantation, enabling the study of these aspects.
[0046] In the claims, reference signs placed between parentheses shall not be construed as limiting the claims. The words "comprising" or "including" do not exclude the presence of elements or steps other than those recited in the claims. In apparatus claims listing several means, several of these means may be embodied by one and the same item of hardware. The word "a" preceding an element does not exclude the presence of a plurality of such elements. In any apparatus claim listing several means, several of these means may be embodied by one and the same item of hardware. The mere fact that certain elements are recited in mutually different dependent claims does not indicate that these elements cannot be used in combination.
[0047] The present invention has been described in detail for the purpose of illustration based on embodiments currently considered to be the most practical and preferred. However, such details are for that purpose only, and the present invention is not limited to the disclosed embodiments. On the contrary, it is intended to cover modifications and equivalent configurations within the spirit and scope of the appended claims. For example, it is intended that the present invention can, to the extent possible, combine one or more features of any embodiment with one or more features of any other embodiment.
Claims
1. 1. A method for segmenting one or more histological structures in a tissue image represented by multi-parameter cellular and subcellular imaging data, comprising: receiving a coarsest level of image data for the tissue image, the coarsest level of image data corresponding to a coarsest level of a multi-scale representation of first data corresponding to the multi-parameter cellular and subcellular imaging data; decomposing said coarsest level image data into a plurality of non-overlapping superpixels; assigning to each superpixel a probability of belonging to said one or more histological structures using a number of pre-trained machine learning algorithms to create a probability map; extracting estimated boundaries of the one or more histological structures by applying a contour algorithm to the probability map; using the estimated boundary to provide an accurate boundary for the one or more histological structures; A method comprising:
2. 2. The method of claim 1, wherein the multi-scale representation includes a Gaussian multi-scale pyramid decomposition, the multi-parameter cellular and subcellular imaging data includes stained tissue image data, receiving the coarsest level image data for the tissue image includes receiving the coarsest level normalized constituent stain image data for the stained tissue image, the coarsest level normalized constituent stain image data being for a particular constituent stain of the stained tissue image and corresponding to a coarsest level of a Gaussian multi-scale pyramid decomposition of the first data corresponding to the stained tissue image data, and decomposing the coarsest level image data into a plurality of non-overlapping superpixels includes splitting the coarsest level normalized constituent stain image data into the plurality of superpixels.
3. 3. The method of claim 2, wherein the multi-parameter cellular and subcellular imaging data includes stained tissue image data, and receiving the coarsest level image data for the tissue image includes receiving the coarsest level normalized constituent stain image data of the stained tissue image, the coarsest level normalized constituent stain image data being generated by: (i) color deconvolving the stained tissue image data into individual staining intensities to create constituent stain image data, (ii) normalizing the constituent stain image data using a reference dataset to create normalized data, and (iii) performing a Gaussian multi-scale decomposition on the normalized data to generate a multi-scale representation of the normalized data including the coarsest level normalized constituent stain image data, and the first data corresponding to the stained tissue image data is the normalized data.
4. 3. The method of claim 2, wherein the first data corresponding to the stained tissue image data is the stained tissue image data, and the coarsest level normalized constituent stain image data is generated by: (i) performing a Gaussian multi-scale decomposition on the stained tissue image data to create a multi-scale representation of the stained tissue image data; (ii) color deconvolving the coarsest level of the Gaussian multi-scale pyramidal decomposition of the stained tissue image data into individual stain intensities to create the coarsest level constituent stain image data; and (iii) normalizing the coarsest level constituent stain image data using a reference dataset to create the coarsest level normalized constituent stain image data.
5. The method of claim 2 , wherein the number of pre-trained machine learning algorithms are a number of supervised machine learning algorithms pre-trained based on user input.
6. The method of claim 5 , wherein the several pre-trained machine learning algorithms include a context-SVM model or a context-LR model applied to the plurality of superpixels, and a staining-SVM model or a staining-LR model applied to the plurality of superpixels.
7. The method of claim 2 , wherein the contour algorithm is a region-based active contour algorithm.
8. 3. The method of claim 2, wherein using the boundary estimate to yield an accurate boundary of the one or more histological structures comprises upsampling a boundary of the one or more histological structures using a level of the Gaussian multi-scale pyramid decomposition other than a coarsest level, and yielding the accurate boundary using a region-based active contour algorithm at a finest level of the Gaussian multi-scale pyramid decomposition.
9. The method of claim 2 , wherein the structures include ducts or glands, the stained tissue image data includes H&E data, and the specific constituent stain is hematoxylin.
10. The method of claim 2 , wherein the structure comprises connective tissue, the stained tissue image data comprises H&E data, and the specific constituent stain is eosin.
11. The method of claim 1 , wherein dividing the coarsest level image data into a plurality of non-overlapping superpixels employs a Linear Iterative Clustering (SLIC) algorithm.
12. A non-transitory computer readable medium storing one or more programs comprising instructions that, when executed by a computer, cause the computer to perform the method of claim 1.
13. 1. A computer system for segmenting one or more histological structures in a tissue image represented by multi-parameter cellular and subcellular imaging data, comprising: receiving a coarsest level of image data for the tissue image, the coarsest level of image data corresponding to a coarsest level of a multi-scale representation of first data corresponding to the multi-parameter cellular and subcellular imaging data; decomposing said coarsest level image data into a plurality of non-overlapping superpixels; assigning to each superpixel a probability of belonging to said one or more histological structures using a number of pre-trained machine learning algorithms to create a probability map; extracting estimated boundaries of the one or more histological structures by applying a contour algorithm to the probability map; using the estimated boundary to provide an accurate boundary for the one or more histological structures; A system comprising a processing device including a number of components configured to perform the steps of:
14. 14. The system of claim 13, wherein the multi-scale representation includes a Gaussian multi-scale pyramid decomposition, the multi-parameter cellular and subcellular imaging data includes stained tissue image data, receiving the coarsest level image data for the tissue image includes receiving the coarsest level normalized constituent stain image data for the stained tissue image, the coarsest level normalized constituent stain image data being for a particular constituent stain of the stained tissue image and corresponding to a coarsest level of a Gaussian multi-scale pyramid decomposition of the first data corresponding to the stained tissue image data, and decomposing the coarsest level image data into a plurality of non-overlapping superpixels includes splitting the coarsest level normalized constituent stain image data into the plurality of superpixels.
15. 15. The system of claim 14, wherein the multi-parameter cellular and subcellular imaging data includes stained tissue image data, and receiving the coarsest level image data for the tissue image includes receiving the coarsest level normalized constituent stain image data of the stained tissue image, the coarsest level normalized constituent stain image data being generated by: (i) color deconvolving the stained tissue image data into individual staining intensities to create constituent stain image data, (ii) normalizing the constituent stain image data using a reference dataset to create normalized data, and (iii) performing a Gaussian multi-scale decomposition on the normalized data to generate a multi-scale representation of the normalized data including the coarsest level normalized constituent stain image data, and the first data corresponding to the stained tissue image data is the normalized data.
16. 15. The system of claim 14, wherein the first data corresponding to the stained tissue image data is the stained tissue image data, and the coarsest level normalized constituent stain image data is generated by: (i) performing a Gaussian multi-scale decomposition on the stained tissue image data to create a multi-scale representation of the stained tissue image data; (ii) color deconvolving the coarsest level of the Gaussian multi-scale pyramidal decomposition of the stained tissue image data into individual stain intensities to create the coarsest level constituent stain image data; and (iii) normalizing the coarsest level constituent stain image data using a reference dataset to create the coarsest level normalized constituent stain image data.
17. 15. The system of claim 14, wherein the number of pre-trained machine learning algorithms are a number of supervised machine learning algorithms pre-trained based on user input.
18. 20. The system of claim 17, wherein the number of pre-trained machine learning algorithms includes a context-SVM model or a context-LR model applied to the plurality of superpixels, and a staining-SVM model or a staining-LR model applied to the plurality of superpixels.
19. The system of claim 14 , wherein the contour algorithm is a region-based active contour algorithm.
20. 15. The system of claim 14, wherein using the boundary estimate to yield an accurate boundary of the one or more histological structures comprises upsampling a boundary of the one or more histological structures using a level of the Gaussian multi-scale pyramid decomposition other than a coarsest level, and yielding the accurate boundary using a region-based active contour algorithm at a finest level of the Gaussian multi-scale pyramid decomposition.
21. The system of claim 14 , wherein the structures include ducts or glands, the stained tissue image data includes H&E data, and the specific constituent stain is hematoxylin.
22. The system of claim 14 , wherein the structure comprises connective tissue, the stained tissue image data comprises H&E data, and the specific constituent stain is eosin.
23. 14. The system of claim 13, wherein dividing the coarsest level image data into a plurality of non-overlapping superpixels employs a Linear Iterative Clustering (SLIC) algorithm.
Citation Information
Patent Citations
Biological object detection
JP2019522276A