Microscopic Slide Image-Based Machine Learning Image Analysis of Inflammatory Bowel Disease
A digital pathology platform with machine learning capabilities addresses the variability in IBD assessment by segmenting intestinal tissue images to generate reproducible histological scores, enhancing the objectivity and accuracy of disease burden evaluation.
Patent Information
- Application Number
- JP2024573636
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-14
- Filing Date
- 2023-06-30
- Publication Date
- 2025-07-10
AI Technical Summary
Current methods for assessing inflammatory bowel disease (IBD) are subjective, lack granularity, and suffer from high variability, making it difficult to reliably determine the effectiveness of treatment options and evaluate disease burden objectively.
A digital pathology platform that uses machine learning to segment and analyze whole-slide images of intestinal tissue, identifying specific cells and tissue regions, and generate reproducible histological scores to quantify disease burden.
Provides unbiased, reproducible data for evaluating IBD, reducing variability and improving the accuracy of disease assessment, enabling better treatment monitoring and efficacy evaluation.
Smart Images

Figure 2025521473000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 358,019, filed on July 1, 2022, and U.S. Provisional Patent Application No. 63 / 387,467, filed on December 14, 2022, the disclosures of which are hereby incorporated by reference in their entireties.
[0002] The subject matter described herein generally relates to digital pathology, and more specifically, to machine - learning image analysis based on microscopic slide images of inflammatory bowel disease.
Background Art
[0003] Inflammatory bowel disease (IBD), which includes different entities of Crohn's disease and ulcerative colitis (UC), is a common disease with a high proportion of poor long - term outcomes in terms of steroid dependence, hospitalization, and surgery, but there are few effective treatment options. UC affects the intestine and the outermost layer of the mucosa and always includes the rectum. UC can also extend proximally to include some or all of the colon. The disease burden of IBD is difficult to clinically and objectively evaluate, making it difficult to reliably determine the effectiveness of various treatment options. Therefore, biological insights into IBD and reliable and reproducible methods for assessing the pathophysiology and disease burden of IBD remain important for identifying new effective treatments, improving patient monitoring, etc.
Summary of the Invention
[0004] A system, method, and product including a computer program product are provided for image-based (e.g., whole-slide image-based) machine learning analysis of IBD. In some exemplary embodiments, a system is provided that includes at least one processor and at least one memory. The at least one memory may include program code that provides operations when executed by the at least one processor. The operations may include determining a plurality of image patches within an image of a biological sample derived from a patient's intestine. Each image patch of the plurality of image patches depicts a portion of the biological sample. The operations may also include determining a plurality of intestinal disease syndromes based at least on the plurality of image patches. Each intestinal disease syndrome of the plurality of intestinal disease syndromes corresponds to a subset of the plurality of image patches. The operations may also include generating a group-level histological score for each intestinal disease syndrome of the plurality of intestinal disease syndromes based at least on the subset of the plurality of image patches included in each respective group. The operations may also include generating an aggregated histological score for the biological sample based on the group-level histological scores generated for each intestinal disease syndrome. The aggregated histological score indicates the disease burden in the patient's intestine.
[0005] In another aspect, a computer-implemented method includes determining a plurality of image patches within an image of a biological sample derived from a patient's intestine. Each image patch of the plurality of image patches depicts a portion of the biological sample. The method may also include determining a plurality of intestinal disease syndromes based at least on the plurality of image patches. Each intestinal disease syndrome of the plurality of intestinal disease syndromes corresponds to a subset of the plurality of image patches. The method may also include generating a group-level histological score for each intestinal disease syndrome of the plurality of intestinal disease syndromes based at least on the subset of the plurality of image patches included in each respective group. The method may also include generating an aggregated histological score for the biological sample based on the group-level histological scores generated for each intestinal disease syndrome. The aggregated histological score indicates the disease burden in the patient's intestine.
[0006] In another aspect, a computer program product is provided that includes a non-transitory computer-readable medium storing instructions. The instructions can cause an operation to be performed by at least one data processor. The operation can include determining a plurality of image patches in an image of a biological sample derived from a patient's intestine. Each image patch of the plurality of image patches depicts a portion of the biological sample. The operation also includes determining a plurality of intestinal disease syndromes based on at least the plurality of image patches. Each intestinal disease syndrome of the plurality of intestinal disease syndromes corresponds to a subset of the plurality of image patches. The operation also includes generating a group-level histological score for each intestinal disease syndrome of the plurality of intestinal disease syndromes based at least on the subset of the plurality of image patches included in each respective group. The operation also includes generating an aggregated histological score for the biological sample based on the group-level histological scores generated for each intestinal disease syndrome. The aggregated histological score indicates the disease burden in the patient's intestine.
[0007] In some variations, one or more features disclosed herein that include the following features can optionally be included in any realizable combination of a system, a method, and / or a non-transitory computer-readable medium. In some variations, the group-level histological score and the aggregated histological score are each one of a Nancy Histological Index (NHI) score, a Robart Histopathology Index (RHI) score, a Geboes Scale score, and a Global Histology Activity Score (GHAS) score.
[0008] In some variations, the group-level histological score and the aggregated histological score are at least one of a first score indicating no disease burden, a second score indicating a low disease burden in the patient's intestine, a third score indicating a moderate disease burden in the patient's intestine, and a fourth score indicating a high disease burden in the patient's intestine.
[0009] In some variations, a subset of the plurality of image patches is formed by clustering at least one or more similar image patches among the plurality of image patches based on features per one or more pixels.
[0010] In some variations, the presence of the first per-pixel feature of one or more per-pixel features in the plurality of image patches is associated with a first possible histological score. The absence of the first per-pixel feature is associated with a second possible histological score.
[0011] In some variations, generating a group-level histological score includes assigning a higher attention score to a first image patch of a subset of the plurality of image patches than to a second image patch of the subset of the plurality of image patches based at least on the presence or absence of the first per-pixel feature. The group-level histological score is generated while determining a representation encoding of the subset of the plurality of image patches.
[0012] In some variations, a higher attention score indicates that the first per-pixel feature of the first image patch contributes more to the representation encoding of the subset of the plurality of image patches than the second patch.
[0013] In some variations, the one or more per-pixel features represent the presence of at least one of tissue erosion, neutrophils, lymphatic structures, crypt abscesses, and necrotic tissue fragments within the epithelium of the tissue in the biological sample.
[0014] In some variations, the one or more per-pixel features include at least one of shape, color, size, presence of dye, and intensity associated with the pixels of the image of the biological sample.
[0015] In some variations, a subset of the plurality of image patches includes a common per-pixel feature of the one or more per-pixel features.
[0016] In some variations, the group-level histological score is generated based on at least one or more of the amount of features per pixel and the distribution of features per pixel in one or more subsets of a plurality of image patches.
[0017] In some variations, clustering is performed by applying clustering analysis techniques.
[0018] In some variations, the clustering analysis techniques include one or more of k-means clustering, mean shift clustering, density-based spatial clustering of applications with noise (DBSCAN), expectation maximization (EM) clustering using a Gaussian mixture model (GMM), and agglomerative hierarchical clustering.
[0019] In some variations, the method further includes generating a first visual representation of a reduced-dimensional representation of a plurality of image patches.
[0020] In some variations, the first visual representation includes one or more visual indicators configured to show the contribution of features per pixel to possible group-level histological scores.
[0021] In some variations, the first visual representation includes one or more visual indicators configured to provide a visual distinction between image patches of a plurality of image patches indicating different possible histological scores.
[0022] In some variations, the first visual representation is generated by applying at least a dimensionality reduction technique to the per-pixel representation of each image patch of the plurality of image patches.
[0023] In some variations, the dimensionality reduction techniques include one or more of principal component analysis (PCA), uniform manifold approximation and projection (UMAP), and t-distributed stochastic neighbor embedding (t-SNE).
[0024] In some variations, the group-level histological score and the aggregated histological score are each generated by applying at least one machine learning model trained to generate the group-level histological score and the aggregated histological score by at least determining the representation encoding of a subset of a plurality of image patches.
[0025] In some variations, the at least one machine learning model includes a multiple instance learning (MIL) model.
[0026] In some exemplary embodiments, a system is provided that includes at least one processor and at least one memory. The at least one memory may include program code that provides an operation when executed by the at least one processor. The operation may include receiving an image of a biological sample derived from a patient's intestine. The image depicts a plurality of cells of the biological sample. The operation may include segmenting the received image into a plurality of portions. Each portion of the plurality of portions corresponds to one of the plurality of cells. The operation may include identifying, based at least on the segmented image, first spatial coordinates associated with each of the plurality of cells in the image. The operation may include identifying a first cell type associated with the first spatial coordinates. The operation may include generating a visual representation that includes the image of the biological sample based at least on the first spatial coordinates and the first cell type.
[0027] In some aspects, the computer-implemented method includes receiving an image of a biological sample derived from a patient's intestine. The image depicts a plurality of cells of the biological sample. The method may include segmenting the received image into a plurality of portions. Each portion of the plurality of portions corresponds to one of the plurality of cells. The method may include identifying first spatial coordinates associated with each cell of the plurality of cells in the image, based at least on the segmented image. The method may include identifying a first cell type associated with the first spatial coordinates. The method may include generating a visual representation including the image of the biological sample, based at least on the first spatial coordinates and the first cell type.
[0028] In some aspects, a computer program product is provided that includes a non-transitory computer-readable medium storing instructions. The instructions can cause an operation to be performed by at least one data processor. The operation may include receiving an image of a biological sample derived from a patient's intestine. The image depicts a plurality of cells of the biological sample. The operation may include segmenting the received image into a plurality of portions. Each portion of the plurality of portions corresponds to one of the plurality of cells. The operation may include identifying first spatial coordinates associated with each cell of the plurality of cells in the image, based at least on the segmented image. The operation may include identifying a first cell type associated with the first spatial coordinates. The operation may include generating a visual representation including the image of the biological sample, based at least on the first spatial coordinates and the first cell type.
[0029] In some variations, one or more features disclosed herein, including the following features, may optionally be included in any realizable combination of a system, a method, and / or a non-transitory computer-readable medium.
[0030] In some variations, the identifying is further based on a plurality of annotations that identify a plurality of cell types depicted in a plurality of images of the biological sample.
[0031] In some variations, the first cell type is at least one of a neutrophil, a plasma cell, a lymphocyte, an intraepithelial lymphocyte, an eosinophil, a mast cell, a macrophage, a goblet cell, an intestinal epithelial cell, an endothelial cell, a fibroblast, a smooth muscle cell, and an endothelial cell.
[0032] In some variations, the image further depicts a plurality of tissue regions of the biological sample.
[0033] In some variations, the method further includes identifying, based at least on the segmented image, a second spatial coordinate associated with each of the plurality of tissue regions and a first tissue region type associated with the second spatial coordinate.
[0034] In some variations, the method further includes generating a second visual representation that includes an image of the biological sample, based at least on the second spatial coordinate and the first tissue region type.
[0035] In some variations, the tissue region type is at least one of epithelium, mucosa, submucosa, normal crypt, invasive crypt, lumen, blood vessel, lymphatic vessel, lamina propria, muscularis mucosa, basal plasmocytosis, ulcer, erosion, granulation tissue, invasive crypt, crypt abscess, normal collagen, abnormal collagen, stroma, stromal subtype, hyperplastic muscle, fissure, abscess, normal fat, abnormal fat, serosa, and serositis.
[0036] In some variations, identifying the second spatial coordinate and the first tissue region type is further based at least on a second plurality of annotations that identify a plurality of tissue region types depicted in a plurality of images of the biological sample.
[0037] In some variations, identifying includes generating, for each cell of the plurality of cells, a metric indicative of a confidence level associated with the identified first cell type.
[0038] In some variations, the method further includes generating spatial table data including first spatial coordinates associated with each of a plurality of cells in an image and a first cell type associated with the first spatial coordinates.
[0039] In some variations, the method includes generating a histological score of a biological sample based on at least the first spatial coordinates and a first cell type associated with the first spatial coordinates. The histological score indicates the disease burden in the patient's intestine.
[0040] In some variations, the histological score is one of a Nancy Histological Index (NHI) score, a Robart Histopathology Index (RHI) score, a Geboes Scale score, a Global Histology Activity Score (GHAS), a myofibrosis score, and a fibrosis score.
[0041] In some variations, the first cell type includes neutrophils.
[0042] In some variations, the histological score is further generated based on the spatial distribution of a plurality of cells identified as having the first cell type.
[0043] In some variations, the histological score is further generated based on the amount of cells of a plurality of cells identified as having the first cell type that match a threshold amount of cells.
[0044] In some variations, the histological score is further generated based on the tissue region types of a plurality of tissue regions depicted in the image.
[0045] In some variations, the type of tissue region is at least one of erosion and ulceration.
[0046] In some variations, the first spatial coordinates are two-dimensional.
[0047] In some variations, the method includes generating an overlay indicating the first cell type at the first spatial coordinates, based at least on the first spatial coordinates and the first cell type. The overlay includes at least one of a mask, a color, and a pattern.
[0048] In some variations, the image is segmented by applying a machine learning model trained to perform per-cell segmentation and per-tissue region segmentation by at least assigning to each pixel of the image a cell segmentation label indicating whether the pixel is associated with the cell type of the cell depicted in the image and a tissue region label indicating whether the pixel is associated with the tissue region type of the tissue region depicted in the image.
[0049] Implementations of the subject matter described herein can include, but are not limited to, methods consistent with the descriptions provided herein, and articles that include a tangible, embodied machine-readable medium operable to cause one or more machines (such as a computer) to perform one or more of the operations implementing one or more of the features described. Similarly, a computer system can be described that can include one or more processors and one or more memories coupled to the one or more processors. The memory can include a non-transitory computer-readable or machine-readable storage medium and can include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. A computer-implemented method according to one or more implementations of the subject matter can be implemented by one or more data processors in a single computing system or multiple computing systems. Such multiple computing systems can be connected through one or more connections, including, for example, connections through a network (such as the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), or through direct connections between one or more of the multiple computing systems, and can exchange data and / or instructions or other commands, etc.
[0050] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. It should be readily understood that specific features of the subject matter of the present disclosure are illustrated for purposes of exemplification in connection with IBD and UC, but such features are not limiting. The claims that follow the present disclosure define the scope of the subject matter to be protected.
Brief Description of the Drawings
[0051] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate specific embodiments of the subject matter disclosed herein and, together with the description, help to explain some of the principles associated with the disclosed implementations. This patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawings will be provided by the Patent Office upon request and payment of the necessary fee.
[0052] In the drawings,
[0053]
Figure 1
[0054]
Figure 2
[0055]
Figure 3
[0056]
Figure 4A
[0057]
Figure 4B
[0058]
Figure 4C
[0059]
Figure 4D
[0060]
Figure 5
[0061]
Figure 6
[0062]
Figure 7
[0063]
Figure 8
[0064]
Figure 9
[0065]
Figure 10A
[0066]
Figure 10B
[0067]
Figure 10C
[0068]
Figure 11
[0069]
Figure 12
[0070]
Figure 13
[0071]
Figure 14
[0072] Where practical, like reference numerals indicate like structures, features, or elements.
DETAILED DESCRIPTION OF THE INVENTION
[0073] IBD including UC has a high rate of poor long - term outcomes regarding hospitalization and surgery, but there are few effective treatment options. As described, a consistent method for biological insights into IBD and evaluating clinical disease burden remains important for identifying new effective treatments, improving patient monitoring, etc. Historically, stool frequency and rectal bleeding, as well as patient - reported outcomes such as endoscopy (e.g., sigmoidoscopy), have been used to assess the health status of the intestinal mucosa. For example, endoscopy has been used to provide a visual assessment of a patient's intestinal mucosa using a camera passed through the intestinal lumen. The appearance of the intestine on endoscopy is generally used to determine whether the intestine shows a meaningful improvement in the health status of the intestine (a concept previously called mucosal healing), but such visual assessments are highly subjective and variable. Furthermore, there are often significant discrepancies between the appearance of the intestinal tissue on endoscopy and the actual inflammatory state of the tissue as seen at the microscopic level. For example, in tissue that appears normal on endoscopy, significant microscopic inflammation of activity may persist. This separation emphasizes the importance of histological evaluation of tissue biopsies taken from IBD patients.
[0074] To address pitfalls when using an endoscope to assess intestinal health, histological scores have been developed, validated, and implemented to more objectively classify microscopic inflammation in the intestine. Although many scoring systems exist, they all assess the presence of neutrophilic inflammation in the tissue and reveal active inflammation in the intestine. The extent of epithelial damage covered by active inflammation (neutrophils) can also be a marker of tissue disease burden. For example, histological scores can be assigned to tissue samples by category. If there is epithelial erosion or ulceration in the tissue sample, the histological score by category can be classified as severe active disease. When neutrophils infiltrate the epithelium, the score can be classified as moderately active disease or mildly active disease depending on the density of neutrophils and their localization within the tissue compartment. In the absence of active disease, there may be increasing chronic (inactive) inflammation with associated category scores. The lowest score (generally a score of 0) is the case where there is no significant increase in either chronic or active inflammation.
[0075] However, conventional methods for assigning such category scores are highly subjective, yield limited ground truth data, and may lack the granularity necessary for consistent and accurate assessment of gut health. Despite the high cost and slow nature of training central readers, typically pathologists, conventional methods for assigning such category scores can result in high levels of variability, such that for the same tissue sample, scores are not readily reproducible between readers (inter-reader variability) and over time by the same reader (intra-reader variability), reducing the power of studies to identify treatment effectiveness. For example, certain cells, such as neutrophils, may be confused with eosinophils or lymphocytes, and vice versa. Similarly, the presence of ulceration can be driven by small, discrete, easily overlooked tissue fragments. Variability may be further increased when classifying lesions as either mild or moderate disease, which is a very qualitative and subjective exercise. Thus, due to exorbitant costs, high-order complexity of implementation, unmanageable variability in scoring and interpretation, and lack of unbiased data, conventional techniques for assessing gut health and assigning histological scores are not practical.
[0076] In accordance with embodiments of the present subject matter, a digital pathology platform can provide robust and reproducible data and category-specific histological scores. In particular, the digital pathology platform described herein can identify relevant cells and features and quantify the cell content of each tissue sample and the corresponding spatial localization of the features within the tissue sample. The digital pathology platform can also generate unbiased datasets for hypothesis validation in clinical trials, determination of clinical response, and correlation with, for example, endoscopy, microbiome determination, gene sequencing, biomarkers, etc. For example, the digital pathology platform can segment an image of a biological sample derived from a patient's intestine into a plurality of cells, identify spatial coordinates associated with the plurality of cells and cell types and / or tissue region types associated with the spatial coordinates, and generate a visual representation based on the identified spatial coordinates and cell types and / or tissue region types. These spatially oriented quantitative features and cell-specific segmentation can be extrapolated to category-specific scores. Additionally and / or alternatively, the digital pathology platform can perform machine-learnable predictions of intestinal disease syndromes and use an end-to-end model to predict reproducible histological scores for an entire intestinal tissue sample. Thus, the digital pathology platform can generate comprehensive and quantitative unbiased data based on tissue samples and predict reproducible histological scores based on tissue samples.
[0077] FIG. 1 depicts a system diagram showing an example of a digital pathology system 100 according to some exemplary embodiments. Referring to FIG. 1, the digital pathology system 100 may include a mucosal biopsy analysis digital pathology platform 110, an imaging system 120, and a client device 130. As shown in FIG. 1, the digital pathology platform 110, the imaging system 120, and the client device 130 may be communicatively coupled via a network 140. The network 140 may be a wired network and / or a wireless network including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc. The imaging system 120 may include one or more imaging devices including, for example, a microscope, a digital camera, a whole slide scanner, a robotic microscope, etc. The client device 130 may be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc.
[0078] The digital pathology platform 110 may include an IBD analysis engine 115 and a mucosal biopsy segmentation engine 116. The analysis engine 115 and / or the segmentation engine 116 may perform one or more of the various processes and / or workflows described herein. The analysis engine 115 and the segmentation engine 116 may communicate with each other. For example, the segmentation engine 116 may perform at least a part of a workflow, and the analysis engine 115 may perform another part of the workflow based on a first part of the workflow performed by the segmentation engine 116, and vice versa. Accordingly, one or more aspects of the analysis engine 115 described herein may be applied to the segmentation engine 116, and one or more aspects of the segmentation engine 116 may be applied to the analysis engine 115. The digital pathology platform 110 may be hosted on a cloud-based infrastructure such that the functionality of the digital pathology platform 110 is remotely accessible, for example, as part of a web-based application, a native mobile application, software as a service (SaaS), etc.
[0079] Referring again to FIG. 1, the digital pathology platform 110 may receive one or more images of a biological sample derived from a patient's intestine from, for example, an imaging system 120. The one or more images of the biological sample may be whole slide images (WSIs). In some embodiments, the one or more images of the biological sample are whole scan images of hematoxylin and eosin stained (H&E) formalin-fixed, paraffin-embedded (FFPE) tissue. The one or more images may be microscope slides and / or images of microscope slides.
[0080] One or more images may depict a plurality of cells of a biological sample that may include a mucosal biopsy. For example, a biological sample from a patient's intestine may include at least one neutrophil, plasma cell, lymphocyte, intraepithelial lymphocyte, eosinophil, mast cell, macrophage, goblet cell, intestinal epithelial cell, endothelial cell, fibroblast, smooth muscle cell, endothelial cell, etc. One or more images may additionally and / or alternatively depict a plurality of tissue regions including epithelium, mucosa, submucosa, normal crypts, invasive crypts, lumen, blood vessels, lymphatics, lamina propria, muscularis mucosa, basal plasmacytosis, ulcers, erosions, granulation tissue, invasive crypts, crypt abscesses, normal collagen, abnormal collagen, stroma, stromal subtypes, hyperplastic muscle, fissures, abscesses, normal fat, abnormal fat, serosa, and serositis. Exemplary images are described in FIGS. 2 and 3.
[0081] As described above, generally, unbiased data generated based on images such as image 200 of FIG. 2 and / or image 300 of FIG. 3 of a biological sample from a patient's intestine is rare. Unbiased data includes data generated as a result of subjective analysis. For example, as described above, the data generated may be biased based on the very subjective and variable nature of visual evaluation, categorical score assignment, and interpretation of images of tissue samples.
[0082] According to some exemplary embodiments, the segmentation engine 116 generates a plurality of unbiased data based on such images. For example, the segmentation engine 116 uses a machine learning model to detect various features of a mucosal biopsy sample including neutrophils, plasma cells, lymphocytes, intraepithelial lymphocytes, eosinophils, mast cells, macrophages, goblet cells, intestinal epithelial cells, endothelial cells, fibroblasts, smooth muscle cells, endothelial cells, epithelium, mucosa, submucosa, normal crypts, invasive crypts, lumens, blood vessels, lymphatic vessels, lamina propria, muscularis mucosa, basal plasmacytosis, ulcers, erosions, granulation tissue, invasive crypts, crypt abscesses, normal collagen, abnormal collagen, stroma, subtypes of stroma, hyperplastic muscle, fissures, abscesses, normal fat, abnormal fat, serosa, serositis, etc. For example, the segmentation engine 116 may generate spatial expression data, heatmap visualization / overlay, predicted histological scores, etc. The spatial expression data includes at least one table having a plurality of rows corresponding to specific cell types and / or tissue region types and a plurality of columns corresponding to at least one spatial coordinate associated with cells having the specific cell types and / or tissue region types. The at least one spatial coordinate may be two-dimensional spatial coordinates to identify the two-dimensional position of at least a portion of the cells having the corresponding cell types and / or tissue region types. In other words, the spatial coordinates may include at least an x coordinate and a y coordinate associated with the position of at least a portion of the cells having the corresponding cell types and / or tissue region types.
[0083] The generated spatial profiling data, visualization, predicted histological scores, etc. assist in identifying, localizing, and quantifying cell types and tissue region types, and thus disease burden, in an image with improved granularity. For example, the spatial profiling data enables localization of each cell and the associated features depicted or detected in the image. Localization of cells and associated features can yield important biomarkers indicative of disease burden in a patient's intestine. For example, localization of cells depicted in an image provides the two-dimensional location of each cell and / or tissue region, allowing determination of the distribution of cells in the image and use in efficiently and accurately assessing disease burden in a patient's intestinal biopsy. The spatial location, quantity, and distribution of cells and / or tissue regions in an image can also be used to judge the therapeutic efficacy of treatment options for treating IBD such as UC. For example, the spatial location, quantity, and distribution can be associated with a particular level of disease burden (e.g., none, mild, moderate, severe, etc.). The spatial location, quantity, and distribution are determined at various time points and can be compared to determine whether the spatial location, quantity, and distribution indicate an improvement in disease burden in the patient's intestine and to judge effective treatment options for treating the disease.
[0084] Based at least on the two-dimensional coordinates of each cell and / or tissue region, the segmentation engine 116 can generate a plurality of visual representations including an image of the biological sample. The visual representation can indicate the position of at least one cell type and / or tissue region type in the image. For example, the segmentation engine 116 can generate an overlay such as at least one of a mask, color, pattern, shadow, etc., each associated with the cell type and / or tissue region type of the cell and / or tissue region depicted in the image of the biological sample. In some embodiments, the generated visual representation can be manipulated via selection at the user interface 135 to show the spatial position of at least one cell and / or tissue region. For example, the generated visual representation can be manipulated via selection at the user interface 135 to display an overlay associated with the selected cell type and / or tissue region type. In other words, selecting a particular cell type and / or tissue region type can indicate the position of each cell having the corresponding cell type and / or tissue region having the corresponding tissue region type. Examples of such visual representations are presented in FIGS. 4A-4D.
[0085] Additionally and / or alternatively, the segmentation engine 116 can predict a histological score of the biological sample depicted in the image based at least on the two-dimensional coordinates of each cell and / or tissue region, the spatial distribution of the cells and / or tissue regions, the amount of cells and / or tissue regions that fit a threshold amount, the detection of specific cell types and / or tissue regions, and / or the localization of specific cell types and / or tissue region types in the image. In some embodiments, the segmentation engine 116 supplies spatial tabular data and / or location-specific data to the analysis engine 115 to predict the histological score. The predicted histological score can indicate the disease burden in the patient's intestine. The predicted histological score can be at least one of a Nancy Histological Index (NHI) score, a Robart Histopathology Index (RHI) score, a Geboes Scale score, a Global Histology Activity Score (GHAS), a myofibrosis score, and a fibrosis score. While a predicted score consistent with an implementation of the subject matter is referred to herein as indicating the disease burden in the patient's intestine, the predicted score can additionally and / or alternatively indicate the disease burden in a biopsy (e.g., a biological sample) or an image of a biopsy (e.g., an image of a biological sample). The predicted score can still indicate the overall disease burden in the patient's intestine. In some embodiments, the predicted score indicates the overall disease burden in the patient's intestine, such as based on the predicted score associated with one or more biopsies (e.g., biological samples) or one or more images of biopsies (e.g., images of biological samples).
[0086] In some embodiments, the category-specific histological score can be scored from 0 to 4, based at least on the spatial coordinates, distribution, etc. of the cells and / or tissue regions identified by the segmentation engine 116. For example, a grade 4 score can be predicted based on the determination of a severely active disease, such as when there are erosions or ulcers in the tissue sample. A grade 3 score can be predicted based on the determination of a moderately active disease, or a grade 2 score can be predicted based on the determination of a moderately active disease, for example, when it is determined that acute inflammatory cells are infiltrating the epithelium of the tissue sample. A grade 1 score can be predicted based on the determination of a moderate or marked increase in chronic inflammatory cells and is identified as infiltrating the lamina propria of the tissue sample without acute inflammatory infiltration. A grade 0 score can be predicted based on the determination that there is no or a mild increase in chronic inflammation when no significant disease activity is identified.
[0087] To generate spatial tabular data, visualization, and / or histological scores, the segmentation engine 116 can segment an image of a biological sample into a plurality of portions corresponding to cells and / or tissue regions. The segmentation engine 116 can segment the image to locate individual cells and / or tissue regions present in the image depicting the biological sample. For example, the segmentation engine 116 can segment the image to locate individual cells including at least one of neutrophils, plasma cells, lymphocytes, intraepithelial lymphocytes, eosinophils, mast cells, macrophages, goblet cells, intestinal epithelial cells, endothelial cells, fibroblasts, smooth muscle cells, endothelial cells, etc., and / or individual tissue regions, such as epithelium, mucosa, submucosa, normal crypts, infiltrated crypts, lumen, blood vessels, lymphatic vessels, lamina propria, muscularis mucosa, basal plasmacytosis, ulcers, erosions, granulation tissue, infiltrated crypts, crypt abscesses, normal collagen, abnormal collagen, stroma, subtypes of stroma, hyperplastic muscle, fissures, abscesses, normal fat, abnormal fat, serosa, serositis, etc.
[0088] For example, the segmentation engine 116 may perform segmentation for each cell and / or each tissue region to assign to each pixel in an image of a biological sample a cell segmentation label that identifies the cell to which the pixel belongs and / or a tissue region label that identifies the tissue region to which the pixel belongs. Further, the task of region segmentation for each cell and / or each tissue may involve operating on high-dimensional image data, particularly when the image is obtained at high resolution and / or at a high level of magnification.
[0089] In some exemplary embodiments, the segmentation engine 116 may apply various cell and / or tissue region segmentation techniques, including, for example, watershed cell segmentation, deep learning-based cell segmentation, and the like. For example, the segmentation engine 116 may train and apply a machine learning model (e.g., a convolutional neural network, etc.) to perform cell-by-cell segmentation and / or tissue-by-tissue region segmentation. The machine learning model may perform cell-by-cell segmentation by at least determining, for each pixel in the image, whether the pixel is part of the background of the image or part of a cell and / or tissue region depicted in the image (see FIG. 4A showing the segmented background). To do so, the machine learning model may further predict, for each pixel in the image, the probability or confidence level that the pixel is part of a cell and / or tissue region depicted in the image.
[0090] In some exemplary embodiments, the segmentation engine 116 may apply multiple cell segmentation techniques to perform per-cell segmentation and identify individual cells and / or tissue regions present in the image. In some cases, the segmentation label for a single pixel may be determined based on a first result of a first cell segmentation technique and a second result of a second cell segmentation technique. For example, the segmentation engine 116 may apply a first machine learning model for determining a first cell and / or tissue region segmentation label for pixels in the image and a second machine learning model for determining a second cell and / or tissue region segmentation label for pixels in the image. Further, the segmentation engine 116 may determine a third cell and / or tissue region segmentation label for a pixel based at least on the first cell and / or tissue region segmentation label and the second cell and / or tissue region segmentation label.
[0091] The segmentation engine 116 may identify at least one spatial coordinate based on the segmented image. For example, by at least identifying the cell and / or tissue region, the segmentation engine 116 may identify at least one spatial coordinate associated with each cell and / or tissue region in the image. As described herein, the spatial coordinate may include the x coordinate and / or y coordinate of the corresponding image. The identified spatial coordinates may be stored for use in generating spatial map data, visualization, histological scores, and the like.
[0092] Once the segmentation engine 116 locates the individual cells and / or tissue regions present in the image, it may identify at least one cell type and / or tissue region type associated with the identified spatial coordinates. For example, the segmentation engine 116 may identify, as associated with the identified pixels, at least one cell type including neutrophils, plasma cells, lymphocytes, intraepithelial lymphocytes, eosinophils, mast cells, macrophages, goblet cells, intestinal epithelial cells, endothelial cells, fibroblasts, smooth muscle cells, endothelial cells, etc., and / or tissue regions such as epithelium, mucosa, submucosa, normal crypts, invasive crypts, lumen, blood vessels, lymphatic vessels, lamina propria, muscularis mucosa, basal plasmacytosis, ulcers, erosions, granulation tissue, invasive crypts, crypt abscesses, normal collagen, abnormal collagen, stroma, subtypes of stroma, hyperplastic muscle, fissures, abscesses, normal fat, abnormal fat, serosa, serositis, etc. The segmentation engine 116 may associate and store the assigned cell type and / or tissue region type with the corresponding pixels. For example, the segmentation engine 116 may store the assigned cell type and / or tissue region type together with the corresponding pixels for use in generating spatial tables, visualizations, histological scores, etc.
[0093] In some embodiments, the segmentation engine 116 identifies the cell type and / or tissue region type corresponding to a particular pixel based on at least one annotation that identifies the cell types and / or tissue region types of multiple images of a biological sample. For example, the at least one annotation may be applied by a pathologist or other trained expert. The pathologist may provide annotations to various images to identify each of the cell types and / or tissue region types in the images. A machine learning model may be trained based at least in part on the annotations. In this manner, the machine learning model may include a supervised machine learning model.
[0094] In some embodiments, the segmentation engine 116 identifies the cell type and / or tissue region type corresponding to a particular pixel based on one or more features identified in the pixel. For example, a pixel may depict one or more features associated with a particular cell type and / or tissue region type. Based on the one or more features depicted in the pixel, the segmentation engine 116 assigns a cell segmentation label and / or a tissue region segmentation label.
[0095] As an example, FIG. 2 is an image 200 of a biological sample derived from the intestine of a patient. In particular, the image 200 is a FFPE tissue sectioned on a slide glass and stained with hematoxylin and eosin (H&E). The image 200 depicts normal or healthy colonic mucosa without disease. The image 200 depicts normal colonic mucosa having an epithelium that includes crypts and surface epithelium, and a lamina propria that includes sparse inflammatory cells such as plasma cells and lymphocytes. Here, the epithelial crypts have a substantially equal spatial distribution, a uniform shape and size, indicating the absence of disease or disease burden.
[0096] FIG. 3 shows a portion of the image 200 compared to an image 300 of another biological sample derived from the intestine of a patient with UC. In the image 300, the lamina propria is filled with a high density of chronic inflammatory cells such as lymphocytes and plasma cells and has features associated with basal plasmacytosis, which is an inflammatory band between the base of the crypts and the muscularis mucosa. Also, as shown in FIG. 3, the epithelium is irregular in size, shape, and distribution (e.g., architectural distortion). In addition to these features of chronic colitis, active inflammation is present as revealed by the presence of neutrophil infiltrates. Epithelial damage is indicated by the presence of cryptitis, ulcers, etc. The segmentation engine 116 can determine that the image 300 depicts chronic active colitis.
[0097] Figure 4A shows an example of a visual representation 400 that includes an image 401 of a biological sample according to some exemplary embodiments. As shown in Figure 4A, the visual representation 400 includes a first selectable element 402 and a second selectable element 404. Selection of the first selectable element 402 can cause an overlay to be displayed on the image 401 showing an artifact of the visual representation 400. Selection of the second selectable element 404 can cause an overlay to be displayed on the image 401 showing the background in the visual representation 400. Thus, the visual representation 400 can be used to more clearly show features of the image 401 for evaluation.
[0098] Figure 4B shows an example of a visual representation 410 that includes an image 411 of a biological sample (which may be the same as or different from image 401) according to some exemplary embodiments. As shown in Figure 4B, the visual representation 410 includes a first selectable element 412, a second selectable element 414, a third selectable element 416, a fourth selectable element 418, a fifth selectable element 420, a sixth selectable element 422, a seventh selectable element 424, an eighth selectable element 426, a ninth selectable element 428, and a tenth selectable element 429. In this example, the first selectable element 412, the second selectable element 414, the third selectable element 416, the fourth selectable element 418, the fifth selectable element 420, the sixth selectable element 422, the seventh selectable element 424, the eighth selectable element 426, the ninth selectable element 428, and the tenth selectable element 429 are associated with different cell types and / or tissue region types. For example, the first selectable element 412 is associated with basal cell carcinoma, the second selectable element 414 is associated with granulation tissue, the third selectable element 416 is associated with blood vessels, the fourth selectable element 418 is associated with erosion or ulceration, the fifth selectable element 420 is associated with the lamina propria, the sixth selectable element 422 is associated with the muscularis mucosa, the seventh selectable element 424 is associated with crypt abscess, the eighth selectable element 426 is associated with the crypt lumen, the ninth selectable element 428 is associated with the infiltrating epithelium, and the tenth selectable element 429 is associated with normal crypts. The first selectable element 412, the second selectable element 414, the third selectable element 416, the fourth selectable element 418, the fifth selectable element 420, the sixth selectable element 422, the seventh selectable element 424, the eighth selectable element 426, the ninth selectable element 428, and the tenth selectable element 429 can each be selected to display an overlay indicating the associated cell type and / or tissue region type.
[0099] FIG. 4C shows an example of a visual representation 430 that includes an image 431 of a biological sample according to some exemplary embodiments. As shown in FIG. 4C, the visual representation 430 includes a first selectable element 432 and a second selectable element 434. In this example, the first selectable element 432 and the second selectable element 434 are associated with different cell types and / or tissue region types. For example, the first selectable element 432 is associated with a vascular lumen, and the second selectable element 434 is associated with the cytoplasm of a goblet cell. The first selectable element 432 and the second selectable element 434 can each be selected to display an overlay indicating the associated cell type and / or tissue region type.
[0100] FIG. 4D shows an example of a visual representation 440 that includes an image 441 of a biological sample, according to some exemplary embodiments. As shown in FIG. 4B, the visual representation 440 includes a first selectable element 442, a second selectable element 444, a third selectable element 446, a fourth selectable element 448, a fifth selectable element 450, a sixth selectable element 452, a seventh selectable element 454, and an eighth selectable element 456. In this example, the first selectable element 442, the second selectable element 444, the third selectable element 446, the fourth selectable element 448, the fifth selectable element 450, the sixth selectable element 452, the seventh selectable element 454, and the eighth selectable element 456 are associated with different cell types and / or tissue region types. For example, the first selectable element 442 is associated with goblet cell nuclei, the second selectable element 444 is associated with absorptive epithelial cells of the intestinal epithelium that are not goblet cells, the third selectable element 446 is associated with intraepithelial lymphocytes, the fourth selectable element 448 is associated with non-epithelial lymphocytes, the fifth selectable element 450 is associated with plasma cells, the sixth selectable element 452 is associated with eosinophils, the seventh selectable element 454 is associated with neutrophils, and the eighth selectable element 426 is associated with other cells. The first selectable element 442, the second selectable element 444, the third selectable element 446, the fourth selectable element 448, the fifth selectable element 450, the sixth selectable element 452, the seventh selectable element 454, and the eighth selectable element 456 can each be selected to display an overlay indicating the associated cell type and / or tissue region type.
[0101] FIG. 5 shows an exemplary bar graph 500 according to some exemplary embodiments. The bar graph 500 compares cell segmentation labels and / or tissue region segmentation labels generated by the segmentation engine 116 (see 502) for a group of 240 images with manual annotations obtained by five pathologists (see 504) for those images. As shown in the bar graph 500, the cell segmentation labels and / or tissue region segmentation labels generated by the segmentation engine 116 had a positive correlation with the manual annotations for each of the cell types and / or tissue region types. FIG. 6 depicts a confusion matrix further showing the positive correlation between the cell segmentation labels and / or tissue region segmentation labels generated by the segmentation engine 116 and the manually applied annotations. Similarly, FIG. 8 shows in the second column that the segmentation engine 116 consistently and accurately identified cell segmentation labels and / or tissue region segmentation labels. For example, as shown in FIG. 8, the model used by the segmentation engine 116 to identify cell segmentation labels and / or tissue region segmentation labels associated with the pixels of an image had a weighted kappa of 0.80 and a Spearman correlation of 0.79.
[0102] FIG. 9 shows an end-to-end model 900 according to some exemplary embodiments. The end-to-end model 900 can be implemented at least in part by an analysis engine 115. The analysis engine 115 can implement the end-to-end model 900 to generate an aggregated histological score (e.g., an overall score) of a biological sample depicted in an image, such as an entire slide scan image of an H&E stained FFPE tissue. As described herein, histological scores, such as aggregated histological scores, include the Nancy Histological Index (NHI) score, the Robart Histopathology Index (RHI) score, the Geboes Scale score, and the Global Histology Activity Score (GHAS). In some examples, the histological score includes a first score indicating that the patient's intestine has no (e.g., does not have) or has a low disease burden, a second score indicating a mild disease burden in the patient's intestine, a third score indicating a moderate disease burden in the patient's intestine, a fourth score indicating a high disease burden, and the like. As described above, while a prediction score consistent with an implementation of the present subject matter is referred to herein as indicating a disease burden in a patient's intestine, the prediction score can additionally and / or alternatively indicate a disease burden in a biopsy (e.g., a biological sample) or an image of a biopsy (e.g., an image of a biological sample). The prediction score can still indicate an overall disease burden in the patient's intestine. In some implementations, the prediction score indicates an overall disease burden in the patient's intestine, such as based on a prediction score associated with one or more biopsies (e.g., biological samples) or one or more images of biopsies (e.g., images of biological samples).
[0103] As described herein, scores have conventionally been assigned to images of biological samples. However, such scores have generally been of low reliability, at least in part due to the subjectivity involved in assigning scores, the intra- and inter-reader variability in assigning scores to the same biological sample, and the variability in interpreting and defining various aspects of the biological sample. The analysis engine 115 can use the end-to-end model 900 to generate a reliable and reproducible histological score for the biological sample depicted in the image. Thus, the analysis engine 115 can reduce the variability previously associated with histological scores, which improves the ability to more accurately evaluate and / or treat an ability that depends on the histological score, as well as the patient's health status (e.g., disease burden in the patient's intestine).
[0104] The end-to-end model 900 can include one or more machine learning models 910, such as a multi-instance machine learning model, among other models. The analysis engine 115 can apply one or more machine learning models 910 to determine a histological score indicative of the level of disease burden in the portion of the intestine depicted in the image, based at least on an image of a biological sample depicting at least a portion of a patient's intestine. For example, as shown in FIG. 9, the analysis engine 115 can determine one or more intestinal disease syndromes, including, for example, a first intestinal disease syndrome 902A, a second intestinal disease syndrome 902B, a third intestinal disease syndrome 902C, etc., in images such as image 200 and / or image 300 of at least a portion of a patient's intestine.
[0105] One or more intestinal disease syndromes each include a plurality of image patches that each depict a portion of the biological sample. For example, FIG. 9 shows an image patch 905 among the plurality of image patches. The image patch 905 can depict at least a portion of the cells present in the biological sample.
[0106] Referring to FIG. 9, the first intestinal disease indication group 902A, the second intestinal disease indication group 902B, and the third intestinal disease indication group 902C may each include a subset of a plurality of image patches representing one or more portions of the biological sample depicted in the image. For example, the first intestinal disease indication group 902A includes a first plurality of image patches 904A, the second intestinal disease indication group 902B includes a second plurality of image patches 904B, and the third intestinal disease indication group 902C includes a third plurality of image patches 904C. Each of the first plurality of image patches 904A, the second plurality of image patches 904B, and the third plurality of image patches 904C includes one or more portions of the biological sample depicted in the image. Thus, the first plurality of image patches 904A, the second plurality of image patches 904B, and the third plurality of image patches 904C may each be a subset of the plurality of image patches.
[0107] In some embodiments, the analysis engine 115 may exclude image patches that do not depict an amount of cells and / or portions of cells present in the biological sample that exceeds a threshold from the generation of the histological score. Examples of image patches excluded from the generation of the histological score may include image patches in which the proportion of the background of the image exceeds a threshold, image patches in which the average color channel dispersion is below a threshold (e.g., gray tiles), and the like.
[0108] The analysis engine 115 may form a subset of a plurality of image patches (e.g., a first plurality of image patches 904A, a second plurality of image patches 904B, a third plurality of image patches 904C, etc.) by at least clustering one or more similar image patches among the plurality of image patches based on a per-pixel representation of the plurality of image patches including features for at least one or more pixels. The analysis engine 115 may form a first intestinal disease syndrome group 902A, a second intestinal disease syndrome group 902B, and a third intestinal disease syndrome group 902C by at least clustering one or more image patches having similar per-pixel features. In some embodiments, the analysis engine 115 forms an intestinal disease load group by at least clustering image patches based on per-pixel features such that each intestinal disease syndrome group includes a distribution of image patches representing the overall image of the biological sample.
[0109] In some embodiments, the analysis engine 115 (e.g., the machine learning model 910) includes a feature extraction mechanism that extracts one or more per-pixel features from a plurality of image patches, e.g., from a per-pixel representation of the plurality of image patches. The per-pixel features may be associated with a particular pixel of the image. The per-pixel features may include features associated with a particular cell and / or tissue region. For example, the per-pixel features may include shape, color, size, presence of a dye, intensity, etc., associated with a particular pixel of an image of a biological sample. The per-pixel features may indicate the presence in a biological sample of tissue region types and / or cell types such as, for example, tissue erosion, neutrophils, lymphatic structures, crypt abscesses, and necrotic tissue fragments within the epithelium of the tissue.
[0110] The analysis engine 115 may form a subset of multiple image patches for each intestinal disease burden group by applying clustering analysis methods such as k-means clustering, mean shift clustering, density-based spatial clustering of applications with noise (DBSCAN), expectation maximization (EM) clustering using a Gaussian mixture model (GMM), and agglomerative hierarchical clustering to the per-pixel representation of the images of the biological sample. In so doing, the analysis engine 115 may identify the number of clusters that maximizes the intra-cluster correlation among the members of each cluster.
[0111] Additionally and / or alternatively, the analysis engine 115 may identify an intestinal disease indication group by applying dimensionality reduction techniques such as principal component analysis (PCA), uniform manifold approximation and projection (UMAP), t-distributed stochastic neighbor embedding (t-SNE), etc. to the per-pixel representation of each image patch of the image. The resulting reduced-dimensional representation of the image patches of the image may correspond to the projection onto a low-dimensional subspace of the per-pixel representation.
[0112] The analysis engine 115 may generate one or more visual representations of the reduced-dimensional representation of the multiple image patches. The one or more visual representations may include one or more visual indicators (e.g., overlays, colors, masks, etc.) that show per-pixel features that are different and / or show the contribution of per-pixel features to possible group-level histological scores (see FIGS. 10A-10C). The one or more visual indicators result in a visual distinction between image patches showing different possible histological scores.
[0113] Referring back to FIG. 9, the analysis engine 115 generates a group-level histological score for each of the plurality of intestinal disease indication groups, based at least on a subset of the plurality of image patches included in each respective group, via the machine learning model 910. In other words, the analysis engine 115 can generate a group-level histological score by applying at least one or more machine learning models to predict the histological score for each intestinal disease indication group of the image patches. The one or more machine learning models can be trained to generate a group-level histological score by at least determining a representation encoding of a subset of the plurality of image patches. In some embodiments, at least one machine learning model is trained based on cell segmentation labels and tissue region segmentation labels generated by the segmentation engine 116 and / or annotations obtained by a pathologist. Referring to FIG. 9, the machine learning model 910 can generate a first group-level histological score 908A corresponding to the first intestinal disease indication group 902A, a second group-level histological score 908B corresponding to the second intestinal disease indication group 902B, and a third group-level histological score 908C corresponding to the third intestinal disease indication group 902C.
[0114] In some embodiments, the group-level histological score can be generated while determining a representation encoding of a subset of the plurality of image patches. For example, the machine learning model 910 can include an attention mechanism configured to assign an attention score to each image patch, representing the contribution (e.g., relevance) of each pixel feature of the image patch to the group-level histological score of the corresponding intestinal disease indication group. Thus, the attention mechanism of the machine learning model 910 can assign a higher attention score to a first image patch of the subset of the plurality of image patches than to a second image patch of the subset of the plurality of image patches, based at least on the presence or absence of a first pixel feature. The higher attention score indicates that the first pixel feature of the first image patch contributes more to the representation encoding of the subset of the plurality of image patches than the second patch.
[0115] In some exemplary embodiments, the analysis engine 115 can generate one or more visual representations, e.g., showing the contribution of per-pixel features to a particular histological score and associated attention score at a particular group level, for display on the user interface 135 of the client device 130. The one or more visual representations can include an overlay showing the contribution at the image patch level to each group level histological score. In general, the prediction of group level histological scores and / or aggregated histological scores can be driven by more image patches that contribute more to a given histological score.
[0116] For example, FIG. 10A shows a first visual representation 1002 corresponding to a first histological score indicating low disease burden, a second visual representation 1004 corresponding to a second histological score indicating a mild or significant increase in disease burden, a third visual representation 1006 corresponding to a third histological score indicating mild disease burden, a fourth visual representation 1008 corresponding to a fourth histological score indicating moderate disease burden, and a fifth visual representation 1010 corresponding to a fifth histological score indicating high disease burden. As shown, the second visual representation 1004 shows a greater number of image patches that contribute more than the first visual representation 1002, the third visual representation 1006 shows a greater number of image patches that contribute more than the second visual representation 1004 and / or the first visual representation 1002, the fourth visual representation 1008 shows a greater number of image patches that contribute more than the first visual representation 1002, the second visual representation 1004, and the third visual representation, and the fifth visual representation 1010 shows a greater number of image patches that contribute more than the first visual representation 1002, the second visual representation 1004, the third visual representation 1006, and the fourth visual representation 1008.
[0117] As another example, FIG. 10B shows a first visual representation 1012 corresponding to a first histological score indicating a low disease burden, a second visual representation 1014 corresponding to a second histological score indicating a mild or significant increase in disease burden, a third visual representation 1016 corresponding to a third histological score indicating a mild disease burden, a fourth visual representation 1018 corresponding to a fourth histological score indicating a moderate disease burden, and a fifth visual representation 1020 corresponding to a fifth histological score indicating a high disease burden. As shown, the second visual representation 1014 shows a greater number of image patches with a higher contribution than the first visual representation 1012, the third visual representation 1016 shows a greater number of image patches with a higher contribution than the second visual representation 1014 and / or the first visual representation 1012, the fourth visual representation 1018 shows a greater number of image patches with a higher contribution than the first visual representation 1012, the second visual representation 1014, and the third visual representation, and the fifth visual representation 1020 shows a greater number of image patches with a higher contribution than the first visual representation 1012, the second visual representation 1014, the third visual representation 1016, and the fourth visual representation 1018.
[0118] As another example, FIG. 10C shows a first visual representation 1022 corresponding to a first histological score indicating a low disease burden, a second visual representation 1024 corresponding to a second histological score indicating a mild or significant increase in disease burden, a third visual representation 1026 corresponding to a third histological score indicating a mild disease burden, a fourth visual representation 1028 corresponding to a fourth histological score indicating a moderate disease burden, and a fifth visual representation 1030 corresponding to a fifth histological score indicating a high disease burden. As shown, the second visual representation 1024 shows a greater number of image patches with a higher contribution than the first visual representation 1022, the third visual representation 1026 shows a greater number of image patches with a higher contribution than the second visual representation 1024 and / or the first visual representation 1022, the fourth visual representation 1028 shows a greater number of image patches with a higher contribution than the first visual representation 1022, the second visual representation 1024, and the third visual representation, and the fifth visual representation 1030 shows a greater number of image patches with a higher contribution than the first visual representation 1022, the second visual representation 1024, the third visual representation 1026, and the fourth visual representation 1028.
[0119] Additionally and / or alternatively, the machine learning model 910 generates a group-level histological score based at least on, for example, the amount of features per one or more pixels within a subset of the plurality of image patches, the distribution of features per one or more pixels within a subset of the plurality of image patches, and the like.
[0120] Referring to FIG. 9, at 911, the analysis engine 115 can generate an aggregated histological score 912 for a biological sample based on the group-level histological scores generated for each intestinal disease indication group. The aggregated histological score indicates the disease burden in the patient's intestine and can be an overall score predicted for the overall image of the biological sample. In some embodiments, the analysis engine 115 generates the aggregated histological score 912 by at least applying a machine learning model, such as the machine learning model 910 or another machine learning model trained to determine an overall histological score based on the generated group-level histological scores. In some embodiments, the analysis engine 115 generates the aggregated histological score 912 by at least determining the average of the generated group-level histological scores (e.g., the first group-level histological score 908A, the second group-level histological score 908B, the third group-level histological score 908C, etc.). In some embodiments, the analysis engine 115 averages the group-level histological scores by applying a weight to at least one of the group-level histological scores. For example, the analysis engine 115 can weight one or more group-level histological scores associated with one or more intestinal disease indication groups and / or one or more image patches, including one or more per-pixel features shown to have a higher contribution to the aggregated histological score and / or a threshold amount of per-pixel features. The analysis engine 115 can aggregate the group-level histological scores to generate the aggregated histological score 912 via one or more other aggregation techniques.
[0121] FIG. 7 depicts a confusion matrix further showing a positive correlation between the aggregated histological scores generated by the analysis engine 115 and the manually assigned histological scores. Similarly, FIG. 8 shows in the third column that the analysis engine 115 consistently and accurately generated histological scores at the slide level and the intestinal disease burden group level. For example, as shown in FIG. 8, the model used by the analysis engine 115 to generate the aggregated histological scores had a weighted kappa of 0.83 and a Spearman correlation of 0.80. Further, FIG. 11 shows a positive correlation between the end-to-end model used by the analysis engine 115 and scores, biomarkers, and other tests, such as the manual NHI score, the machine NHI score, the endoscopic subscore (ES), the physician's global assessment (PGA), the stool frequency (SF) biomarker, the rectal bleeding (RB) biomarker, the c-reactive protein (CRP) biomarker, and the fecal calprotectin (FCP) biomarker.
[0122] FIG. 12 depicts a flowchart showing an example of a process 1200 for cell segmentation and / or tissue region segmentation according to some exemplary embodiments. The process 1200 may be implemented by the analysis engine 115, the segmentation engine 116, the digital pathology platform 110, and / or other components therein.
[0123] At 1202, the segmentation engine 116 receives an image of a biological sample derived from a patient's intestine. The image may depict a plurality of cells of the biological sample. The image may further depict a plurality of tissue regions of the biological sample. The image may include a whole slide image, such as a scanned image of H&E stained FFPE tissue.
[0124] At 1204, the segmentation engine 116 segments the received image into a plurality of parts. Each part of the plurality of parts corresponds to one cell of the plurality of cells and / or one tissue region of the plurality of tissue regions. For example, the segmentation engine 116 may segment the image by applying a machine learning model trained to perform cell-by-cell segmentation and tissue-by-tissue segmentation. The segmentation engine 116 may assign to each pixel of the image a cell segmentation label indicating whether the pixel is associated with the cell type of the cell depicted in the image, and a tissue region label indicating whether the pixel is associated with the tissue region type of the tissue region depicted in the image. In some embodiments, the image is manually annotated with a plurality of annotations that identify a plurality of cell types and a plurality of tissue region types.
[0125] At 1206, the segmentation engine 116 identifies first spatial coordinates associated with each cell of the plurality of cells in the image, based at least on the segmented image. The segmentation engine 116 may additionally and / or alternatively identify second spatial coordinates associated with each tissue region of the plurality of tissue regions in the image, based at least on one segmented image. The first spatial coordinates and / or the second spatial coordinates may be two-dimensional.
[0126] At 1208, the segmentation engine 116 identifies a first cell type associated with a first spatial coordinate. The first cell type is at least one of a neutrophil, a plasma cell, a lymphocyte, an intraepithelial lymphocyte, an eosinophil, a mast cell, a macrophage, a goblet cell, an intestinal epithelial cell, an endothelial cell, a fibroblast, a smooth muscle cell, and an endothelial cell. Additionally and / or alternatively, the segmentation engine 116 identifies a first tissue region type associated with a second spatial coordinate. The tissue region type is at least one of an epithelium, a mucosa, a submucosa, a normal crypt, an infiltrated crypt, a lumen, a blood vessel, a lymphatic vessel, a lamina propria, a muscularis mucosa, a basal plasmocytosis, an ulcer, an erosion, granulation tissue, an infiltrated crypt, a crypt abscess, normal collagen, abnormal collagen, a stroma, a stromal subtype, a hypertrophic muscle, a fissure, an abscess, normal fat, abnormal fat, a serosa, and a serositis.
[0127] In some embodiments, the segmentation engine 116 identifies the first cell type based on a plurality of annotations that identify a plurality of cell types depicted in a plurality of images of a biological sample. For example, the plurality of annotations may be manually applied by one or more pathologists and / or may be generated by the digital pathology platform 110. Similarly, the segmentation engine 116 may identify the first tissue region type based on a second plurality of annotations that identify a plurality of tissue region types depicted in a plurality of images of a biological sample. The second plurality of annotations may be manually applied by one or more pathologists and / or may be generated by the digital pathology platform 110. In some embodiments, the segmentation engine 116 generates a metric indicative of a confidence level associated with the identified first cell type for each cell of the plurality of cells and / or the identified first tissue region type for each tissue region of the plurality of tissue regions.
[0128] In 1210, the segmentation engine 116 generates a visual representation including an image of a biological sample based on at least a first spatial coordinate and a first cell type. The segmentation engine 116 may further generate a visual representation based on at least a second spatial coordinate and a first tissue region type. The segmentation engine 116 may generate an overlay indicating the first cell type at the first spatial coordinate and / or the first tissue region type at the second spatial coordinate. The overlay may include at least one mask, color, pattern, etc.
[0129] Additionally and / or alternatively, the segmentation engine 116 generates spatial table data including the first spatial coordinates associated with each cell of a plurality of cells in the image and the first cell type associated with the first spatial coordinates. The segmentation engine 116 may generate spatial table data including the second spatial coordinates associated with each tissue region of a plurality of tissue regions in the image and the first tissue region type associated with the second spatial coordinates. The spatial table data may provide localization of each of the cells and / or tissue regions in the image. The spatial table data may be stored and used to generate a visual representation and / or a histological score.
[0130] Additionally and / or alternatively, the segmentation engine 116 generates a histological score of the biological sample based on at least the first spatial coordinate and the first cell type associated with the first spatial coordinate, and / or the second spatial coordinate and the first tissue region type associated with the second spatial coordinate. In some embodiments, the segmentation engine 116 communicates with the analysis engine 115 to generate a histological score.
[0131] The histological score indicates the disease burden in the patient's intestine. The histological score includes one of the Nancy Histological Index (NHI) score, Robart Histopathology Index (RHI) score, Geboes Scale score, Global Histology Activity Score (GHAS), myofibrosis score, fibrosis score, etc. The histological score can be generated based on the spatial distribution of a plurality of cells identified as having a first cell type and / or the spatial distribution of a plurality of tissue regions identified as having a first tissue region type. Additionally and / or alternatively, the histological score is generated based on the amount of cells of a plurality of cells identified as having a first cell type that conforms to a threshold amount of cells. Additionally and / or alternatively, the histological score is generated based on the tissue region type of a plurality of tissue regions depicted in an image, such as when the tissue region type is erosion and / or ulceration.
[0132] Accordingly, the digital pathology platform 110 can be used to generate unbiased locational data, such as spatial tabular data and other data, that describe an image of a biological sample to generate visual representations, predict cells and / or tissue regions in the image of the biological sample, generate histological scores, and the like.
[0133] In some embodiments, category-specific histological scores can be calculated using the spatial table data, visualization, and / or histological scores generated by the segmentation engine 116. For example, a graph neural network (GNN) can be implemented by the analysis engine 115 to receive spatial table data, visualization, and / or predicted histological scores (e.g., Nancy score) and output category-specific histological scores (e.g., grades from 0 to 4). In one or more examples, the GNN can be trained to learn the association between the spatial table data and the category-specific scores. The analysis engine 115 can obtain the spatial table data, visualization, and predicted histological scores, and input this data into the trained GNN to obtain the category-specific scores for a given image. Thus, the extracted features such as spatial table data can be connected to the molecular data of the tissue sample depicted by the image. Subsequently, the molecular data / markers can be related to the phenotype to perform an assessment regarding inflammation, which can further support the disease burden in IBD analysis. In some embodiments, the spatial table data can be clustered together using the trained GNN to predict category-specific scores. The spatial table data includes multi-factor data, and the GNN can be used to aggregate the data by identifying patterns and relationships between the data.
[0134] FIG. 13 depicts a flowchart illustrating an example of a process 1300 for histological score generation according to some exemplary embodiments. The process 1300 can be implemented by the analysis engine 115, the segmentation engine 116, the digital pathology platform 110, and / or other components therein.
[0135] At 1302, the analysis engine 115 can determine a plurality of image patches in an image of a biological sample derived from a patient's intestine. Each image patch of the plurality of image patches depicts a part of the biological sample. The image can include a whole slide image such as a scanned image of H&E-stained FFPE tissue.
[0136] At 1304, the analysis engine 115 may determine a plurality of intestinal disease syndromes based on at least a plurality of image patches. Each intestinal disease syndrome of the plurality of intestinal disease syndromes corresponds to a subset of the plurality of image patches.
[0137] In some embodiments, the subset of the plurality of image patches includes common per-pixel features of features per one or more pixels. For example, the subset of the plurality of image patches may be formed by clustering at least one or more similar image patches among the plurality of image patches based on at least one or more per-pixel features. The analysis engine 115 may perform the clustering by applying a clustering analysis method. The clustering analysis techniques may include one or more of k-means clustering, mean shift clustering, density-based spatial clustering of applications with noise (DBSCAN), expectation maximization (EM) clustering using a Gaussian mixture model (GMM), and agglomerative hierarchical clustering. In some embodiments, a GNN may be used to cluster spatio-table data and predict category-specific scores.
[0138] The per-pixel feature may be associated with a specific pixel of the image. The per-pixel feature may include shape, color, size, presence of a dye, intensity, etc. associated with a specific pixel of the image of the biological sample. The per-pixel feature may indicate the presence of at least one of tissue erosion, neutrophils, lymphatic structures, crypt abscesses, and necrotic tissue fragments within the epithelium of the tissue in the biological sample.
[0139] At 1304, the analysis engine 115 generates a group-level histological score for each of the plurality of intestinal disease indication groups, based at least on a subset of the plurality of image patches included in each respective group. The group-level histological score is at least one of a Nancy Histological Index (NHI) score, a Robart Histopathology Index (RHI) score, a Geboes Scale score, and a Global Histology Activity Score (GHAS) score. Additionally and / or alternatively, the group-level histological score is at least one of a first score indicating the absence of disease burden in the patient's intestine, a second score indicating a mild disease burden in the patient's intestine, and a third score indicating a moderate disease burden in the patient's intestine. The score can be in the range of 0 to 4 (or other scale depending on the scoring system). In the example, the first score is 0, the second score is 1, and the third score is 2.
[0140] The analysis engine 115 can generate the group-level histological score by applying at least one machine learning model at least to predict the histological score for each intestinal disease indication group of the image patches. At least one machine learning model, which may include a multiple instance learning (MIL) model among other models, can be trained to generate the group-level histological score by at least determining the representation encoding of a subset of the plurality of image patches. In some embodiments, the at least one machine learning model is trained with cell segmentation labels and tissue region segmentation labels generated by the segmentation engine 116. For example, process 1200 or a similar process can be used for cell segmentation and / or tissue region segmentation.
[0141] In some embodiments, the group-level histological score is generated while determining the representational encoding of a subset of the plurality of image patches. For example, the analysis engine 115 may assign a higher attention score to a first image patch of the subset of the plurality of image patches than to a second image patch of the subset of the plurality of image patches based at least on the presence of a feature per first pixel or the absence of a feature per first pixel. The higher attention score indicates that the feature per first pixel of the first image patch contributes more to the representational encoding of the subset of the plurality of image patches than the second patch.
[0142] Additionally and / or alternatively, the analysis engine 115 generates a group-level histological score based at least on one or more amounts of features per one or more pixels within the subset of the plurality of image patches and the distribution of features per one or more pixels within the subset of the plurality of image patches. In some embodiments, the presence of a feature per first pixel of one or more pixels in the plurality of image patches is associated with a first possible histological score, and the absence of a feature per first pixel is associated with a second possible histological score. Thus, the presence, amount, and / or distribution of features per pixel within the subset of the plurality of image patches can be used to generate a histological score.
[0143] At 1306, the analysis engine 115 generates an aggregated histological score for the biological sample based on the group-level histological scores generated for each intestinal disease indication group. The aggregated histological score indicates the disease burden in the patient's intestine. The aggregated histological score is the overall score predicted for the image of the microscopic slide.
[0144] The aggregated histological score is at least one of a Nancy Histological Index (NHI) score, a Robart Histopathology Index (RHI) score, a Geboes Scale score, and a Global Histology Activity Score (GHAS) score. Additionally and / or alternatively, the aggregated histological score is at least one of a first score indicating a low disease burden in the patient's intestine, a second score indicating a moderate disease burden in the patient's intestine, and a third score indicating a high disease burden in the patient's intestine. The score can be in the range of 0 to 4 (or other scales depending on the scoring system). In an example, the first score is 1, the second score is 2 or 3, and the third score is 4.
[0145] In some embodiments, the analysis engine 115 generates a first visual representation of the reduced-dimensional representations of a plurality of image patches. The first visual representation can include one or more visual indicators (e.g., overlays, colors, masks, etc.). The one or more visual indicators can indicate features for each pixel. The one or more visual indicators can indicate the contribution of per-pixel features to possible group-level histological scores. The one or more visual indicators result in a visual distinction between image patches of a plurality of image patches indicating different possible histological scores.
[0146] The analysis engine 115 may generate the first visual representation by applying at least dimensionality reduction techniques to the per-pixel representations of each image patch of the plurality of image patches. The dimensionality reduction techniques include one or more of principal component analysis (PCA), uniform manifold approximation and projection (UMAP), and t-distributed stochastic neighbor embedding (t-SNE). Furthermore, the dimension of the spatial tabular data can be reduced by using a GNN to detect patterns and relationships in the multi-dimensional dataset.
[0147] Accordingly, the digital pathology platform 110 can efficiently generate a reproducible, consistent, and aggregated histological score for biological samples. The reproducible, consistent, and aggregated histological score can be generated according to the embodiments described herein, while limiting or eliminating the subjectivity that previously reduced the reliability of such histological scores.
[0148] FIG. 14 is a block diagram depicting an example of a computing system 1400 according to some exemplary embodiments. Referring to FIGS. 1 and 14, the computing system 1400 can be used to implement the digital pathology platform 110, the client device 130, the analysis engine 115, the segmentation engine 116, and / or any of its components.
[0149] As shown in FIG. 14, the computing system 1400 can include a processor 1410, a memory 1420, a storage device 1430, and an input / output device 1440. The processor 1410, the memory 1420, the storage device 1430, and the input / output device 1440 can be interconnected via a system bus 1450. The processor 1410 is capable of processing instructions for execution within the computing system 1400. Such executed instructions can implement one or more components, such as, for example, the digital pathology platform 110, the client device 130, the analysis engine 115, the segmentation engine 116, etc. In some exemplary embodiments, the processor 1410 can be a single-threaded processor. Alternatively, the processor 1410 can be a multi-threaded processor. The processor 1410 is capable of processing instructions stored in the memory 1420 and / or the storage device 1430 to display graphical information for a user interface provided via the input / output device 1440.
[0150] Memory 1420 is a computer-readable medium, such as volatile or non-volatile, that stores information within computing system 1400. Memory 1420 can store, for example, a data structure representing a configuration object database. Storage device 1430 can provide persistent storage for computing system 1400. Storage device 1430 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. Input / output device 1440 provides input / output operations for computing system 1400. In some exemplary embodiments, input / output device 1440 includes a keyboard and / or a pointing device. In various embodiments, input / output device 1440 includes a display device for displaying a graphical user interface.
[0151] According to some exemplary embodiments, input / output device 1440 can provide input / output operations for network devices. For example, input / output device 1440 can include an Ethernet port or other networking port to communicate with one or more wired and / or wireless networks (e.g., local area network (LAN), wide area network (WAN), Internet).
[0152] In some exemplary embodiments, computing system 1400 can be used to execute various interactive computer software applications that can be used for the compilation, analysis, and / or storage of various forms of data. Alternatively, computing system 1400 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., generation, management, editing of spreadsheet documents, word processing documents, and / or other objects), computing functions, communication functions, and the like. The applications can include various add-in functions or can be stand-alone computing products and / or functions. When activated within an application, the functionality can be used to generate a user interface provided via input / output device 1440. The user interface can be generated by computing system 1400 and presented to the user (e.g., on a computer screen monitor, etc.).
[0153] One or more aspects or features of the subject matter described in this specification can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can be included in an implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a memory system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally located at remote locations and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other.
[0154] These computer programs, which may also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer program product, apparatus, and / or device, such as, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium can store such machine instructions non-transitorily, for example, in non-transitory solid-state memory, magnetic hard drive, or any equivalent storage medium. A machine-readable medium can alternatively or additionally store such machine instructions transiently, for example, in a processor cache or other random access memory associated with one or more physical processor cores.
[0155] To provide for interaction with a user, one or more aspects or features of the subject matter described in this specification may be implemented on a computer having, for example, a display device such as a cathode ray tube (CRT), liquid crystal display (LCD), or light emitting diode (LED) monitor for displaying information to the user, and a keyboard, and a pointing device such as a mouse or trackball by which the user can provide input to the computer. Other types of devices may also be used to provide for interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback such as, for example, visual feedback, auditory feedback, or tactile feedback, and the input received from the user can be received in any form including acoustic input, voice input, or tactile input. Other possible input devices include a touch screen, or other touch sensor-based devices such as a single or multi-point resistive or capacitive trackpad, speech recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.
[0156] In the above specification and claims, phrases such as "at least one of ~" or "one or more of ~" may appear before a list of consecutive elements or features. The term "and / or" may also be used in the listing of two or more elements or features. Such phrases are intended to mean either each of the recited elements or features individually, or any of the recited elements or features in combination with any of the other recited elements or features, as long as there is no implicit or explicit contradiction depending on the context in which they are used. For example, the phrases "at least one of A and B", "one or more of A and B", and "A and / or B" are each intended to mean "only A, only B, or A and B together". A similar interpretation is intended for lists containing three or more items. For example, the phrases "at least one of A, B, and C", "one or more of A, B, and C", and "A, B, and / or C" are each intended to mean "only A, only B, only C, A and B together, A and C together, B and C together, or A, B, and C together". The use of the term "based on" in the above and the claims means "at least partially based on" and means that features or elements not recited are also allowed.
[0157] The subject matter described in this specification can be embodied in a system, apparatus, method, and / or article, according to a desired configuration. The implementations described in the foregoing description do not necessarily represent all implementations according to the subject matter described in this specification. Instead, the implementations described in the foregoing description are only some examples according to aspects related to the described subject matter. Several variations have been described in detail above, but other modifications or additional forms are possible. In particular, further features and / or variations can be provided in addition to those described in this specification. For example, the implementations described above can be directed to various combinations and sub - combinations of the disclosed features, and / or combinations and sub - combinations of several further features disclosed above. Additionally, the logical flows depicted in the accompanying figures and / or described in this specification do not necessarily require the particular order or sequential order shown to achieve a desired result. Other implementations can be within the scope of the following claims.
Claims
**Claim 1** A computer-implemented method comprising: determining a plurality of image patches in an image of a biological sample derived from a patient's intestine, each image patch of the plurality of image patches depicting a part of the biological sample; determining a plurality of intestinal disease syndromes based at least on the plurality of image patches, each intestinal disease syndrome of the plurality of intestinal disease syndromes corresponding to a subset of the plurality of image patches; generating a group-level histological score for each intestinal disease syndrome of the plurality of intestinal disease syndromes based at least on the subset of the plurality of image patches included in each respective group; and generating an aggregated histological score for the biological sample based on the generated group-level histological scores for each intestinal disease syndrome, the aggregated histological score indicating a disease burden in the patient's intestine. A computer-implemented method as described above. **Claim 2** The method according to claim 1, wherein the group-level histological score and the aggregated histological score are each one of a Nancy Histological Index (NHI) score, a Robart Histopathology Index (RHI) score, a Geboes Scale score, and a Global Histology Activity Score (GHAS) score. **Claim 3** The method according to claim 1 or 2, wherein the group-level histological score and the aggregated histological score are each at least one of a first score indicating no disease burden, a second score indicating a low disease burden in the patient's intestine, a third score indicating a moderate disease burden in the patient's intestine, and a fourth score indicating a high disease burden in the patient's intestine. **Claim 4** The method according to any one of claims 1 to 3, wherein the subset of the plurality of image patches is formed by clustering at least one or more similar image patches of the plurality of image patches based on features per at least one or more pixels. **Claim 5** The method according to claim 4, wherein the presence of the feature per first pixel among the features per one or more pixels in the plurality of image patches is associated with a first possible histological score, and the absence of the feature per first pixel is associated with a second possible histological score.
6. The generating of the group-level histological score includes assigning a higher attention score to a first image patch of the subset of the plurality of image patches than to a second image patch of the subset of the plurality of image patches based at least on the presence of the feature per first pixel or the absence of the feature per first pixel, and the group-level histological score is generated while determining a representation encoding of the subset of the plurality of image patches. The method according to claim 5.
7. The method according to claim 6, wherein the higher attention score indicates that the feature per first pixel of the first image patch contributes more significantly to the representation encoding of the subset of the plurality of image patches than the second patch.
8. The method according to claim 4, wherein the feature per one or more pixels represents the presence of at least one of tissue erosion, neutrophils, lymphoid structures, crypt abscesses, and necrotic tissue fragments within the epithelium of the tissue in the biological sample.
9. The method according to claim 4, wherein the feature per one or more pixels includes at least one of shape, color, size, presence of dye, and intensity associated with the pixels of the image of the biological sample.
10. The method according to claim 4, wherein the subset of the plurality of image patches includes a common feature per pixel of the feature per one or more pixels.
11. The method according to claim 4, wherein the group-level histological score is generated based on at least one or more of the amount of the feature per one or more pixels in the subset of the plurality of image patches and the distribution of the feature per one or more pixels in the subset of the plurality of image patches.
12. The method according to claim 4, wherein the clustering is performed by applying a clustering analysis technique.
13. The clustering analysis technique includes one or more of k-means clustering, mean shift clustering, density-based spatial clustering of applications with noise (DBSCAN), expectation maximization (EM) clustering using a Gaussian mixture model (GMM), and agglomerative hierarchical clustering, the method according to claim 12.
14. The method according to any one of claims 1 to 13, further comprising generating a first visual representation of the reduced dimensional representations of the plurality of image patches.
15. The method according to claim 14, wherein the first visual representation includes one or more visual indicators configured to show the contribution of per-pixel features to possible group-level histological scores.
16. The method according to claim 14, wherein the first visual representation includes one or more visual indicators configured to provide a visual distinction between the image patches of the plurality of image patches indicating different possible histological scores.
17. The method according to claim 14, wherein the first visual representation is generated by applying at least a dimensionality reduction technique to the per-pixel representation of each image patch of the plurality of image patches.
18. The method according to claim 17, wherein the dimensionality reduction technique includes one or more of principal component analysis (PCA), uniform manifold approximation and projection (UMAP), and t-distributed stochastic neighbor embedding (t-SNE).
19. The method according to any one of claims 1 to 18, wherein the group-level histological score and the aggregated histological score are each generated by applying at least one machine learning model trained to generate the group-level histological score and the aggregated histological score by at least determining the representation encoding of the subset of the plurality of image patches.
20. The method according to claim 19, wherein the at least one machine learning model includes a multiple instance learning (MIL) model.
21. A system, at least one data processor, and at least one memory storing instructions that, when executed by the at least one data processor, cause the at least one data processor to Determining a plurality of image patches in an image of a biological sample derived from a patient's intestine, each of the plurality of image patches depicting a part of the biological sample; Determining a plurality of intestinal disease syndrome groups based at least on the plurality of image patches, each intestinal disease syndrome group of the plurality of intestinal disease syndrome groups corresponding to a subset of the plurality of image patches; Generating a group-level histological score for each intestinal disease syndrome group of the plurality of intestinal disease syndrome groups based at least on the subset of the plurality of image patches included in each respective group, and Generating an aggregated histological score for the biological sample based on the generated group-level histological scores for each intestinal disease syndrome group, the aggregated histological score indicating the disease burden in the patient's intestine, a memory for causing the generation of the aggregated histological score for the biological sample; A system comprising. Claim 22 A non-transitory computer-readable medium storing instructions, which when executed by at least one data processor, Determining a plurality of image patches in an image of a biological sample derived from a patient's intestine, each of the plurality of image patches depicting a part of the biological sample; Determining a plurality of intestinal disease syndrome groups based at least on the plurality of image patches, each intestinal disease syndrome group of the plurality of intestinal disease syndrome groups corresponding to a subset of the plurality of image patches; Generating a group-level histological score for each intestinal disease syndrome group of the plurality of intestinal disease syndrome groups based at least on the subset of the plurality of image patches included in each respective group, and Generating an aggregated histological score for the biological sample based on the generated group-level histological scores for each intestinal disease syndrome group, the aggregated histological score indicating the disease burden in the patient's intestine, generating the aggregated histological score; A non-transitory computer-readable medium causing an operation including. Claim 23 A computer-implemented method, Receiving an image of a biological sample derived from a patient's intestine, wherein the image depicts a plurality of cells of the biological sample. Segmenting the received image into a plurality of parts, wherein each part of the plurality of parts corresponds to one of the plurality of cells. Identifying first spatial coordinates associated with each cell of the plurality of cells in the image, based at least on the segmented image. Identifying a first cell type associated with the first spatial coordinates, and Generating a visual representation including the image of the biological sample, based at least on the first spatial coordinates and the first cell type. A computer-implemented method comprising the above steps. **Claim 24** The method according to claim 23, wherein the identifying is further based on a plurality of annotations identifying a plurality of cell types depicted in a plurality of images of the biological sample. **Claim 25** The method according to claim 23 or 24, wherein the first cell type is at least one of a neutrophil, a plasma cell, a lymphocyte, an intraepithelial lymphocyte, an eosinophil, a mast cell, a macrophage, a goblet cell, an intestinal epithelial cell, an endothelial cell, a fibroblast, a smooth muscle cell, and an endothelial cell. **Claim 26** The method according to any one of claims 23 to 25, wherein the image further depicts a plurality of tissue regions of the biological sample. **Claim 27** The method according to claim 26, further comprising identifying second spatial coordinates associated with each tissue region of the plurality of tissue regions, and a first tissue region type associated with the second spatial coordinates, based at least on the segmented image. **Claim 28** The method according to claim 27, further comprising generating a second visual representation including the image of the biological sample, based at least on the second spatial coordinates and the first tissue region type. **Claim 29** The method according to claim 27, wherein the tissue region type is at least one of epithelium, mucosa, submucosa, normal crypt, invasive crypt, lumen, blood vessel, lymphatic vessel, lamina propria, muscularis mucosa, basal plasmacytosis, ulcer, erosion, granulation tissue, invasive crypt, crypt abscess, normal collagen, abnormal collagen, stroma, stroma subtype, hyperplastic muscle, fissure, abscess, normal fat, abnormal fat, serosa, and serositis. **Claim 30** The specifying of the second spatial coordinate and the first tissue region type is further based at least on a second plurality of annotations that specify a plurality of tissue region types depicted in a plurality of images of a biological sample, the method of claim 27.
31. The specifying includes generating, for each cell of the plurality of cells, a metric indicative of a confidence level associated with the identified first cell type, the method according to any one of claims 23 to 30.
32. The method according to any one of claims 23 to 31 further includes generating spatial table data including the first spatial coordinates associated with each cell of the plurality of cells in the image and the first cell type associated with the first spatial coordinates.
33. Generating a histological score of the biological sample based at least on the first spatial coordinates and the first cell type associated with the first spatial coordinates, wherein the histological score indicates a disease burden in the intestine of the patient, the method according to any one of claims 23 to 32.
34. The method of claim 33, wherein the histological score is one of a Nancy Histological Index (NHI) score, a Robart Histopathology Index (RHI) score, a Geboes Scale score, a Global Histology Activity Score (GHAS), a myofibroplasia score, and a fibrosis score.
35. The method of claim 33, wherein the first cell type includes neutrophils.
36. The method of claim 33, wherein the histological score is further generated based on a spatial distribution of the plurality of cells identified as having the first cell type.
37. The method of claim 33, wherein the histological score is further generated based on a quantity of cells of the plurality of cells identified as having the first cell type that conforms to a threshold quantity of cells.
38. The method of claim 33, wherein the histological score is further generated based on tissue region types of a plurality of tissue regions depicted in the image.
39. The method of claim 38, wherein the tissue region type is at least one of erosion and ulceration.
40. The method according to any one of claims 23 to 39, wherein the first spatial coordinates are two-dimensional.
41. Generating an overlay indicating the first cell type at the first spatial coordinates based at least on the first spatial coordinates and the first cell type, the overlay including at least one of a mask, a color, and a pattern, the method according to any one of claims 23 to 40, further comprising generating the overlay.
42. The method according to any one of claims 23 to 41, wherein the image is segmented by applying a machine learning model trained to perform cell-by-cell segmentation and tissue-by-region segmentation by at least assigning to each pixel of the image a cell segmentation label indicating whether the pixel is associated with the cell type of a cell depicted in the image and a tissue region label indicating whether the pixel is associated with the tissue region type of a tissue region depicted in the image.
43. A system comprising: at least one data processor, and at least one memory storing instructions that, when executed by the at least one data processor, cause the at least one data processor to receive an image of a biological sample derived from a patient's intestine, the image depicting a plurality of cells of the biological sample, segment the received image into a plurality of portions, each portion of the plurality of portions corresponding to one of the plurality of cells, identify first spatial coordinates associated with each of the plurality of cells in the image based at least on the segmented image, identify a first cell type associated with the first spatial coordinates, and generate a visual representation including the image of the biological sample based at least on the first spatial coordinates and the first cell type. A system comprising the above.
44. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, Receiving an image of a biological sample derived from a patient's intestine, wherein the image depicts a plurality of cells of the biological sample, receiving an image of the biological sample, Segmenting the received image into a plurality of portions, wherein each of the plurality of portions corresponds to one of the plurality of cells, segmenting the received image, Identifying first spatial coordinates associated with each cell of the plurality of cells in the image, based at least on the segmented image, Identifying a first cell type associated with the first spatial coordinates, and Generating a visual representation including the image of the biological sample, based at least on the first spatial coordinates and the first cell type, A non-transitory computer-readable medium that causes operations including the above to occur.