Method for an image analysis pipeline
Patent Information
- Application Number
- EP2024800880
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-28
- Publication Date
- 2026-09-09
AI Technical Summary
Current methods for diagnosing and prognosing diseases based on cell morphology are highly manual and time-consuming, leading to bottlenecks in diagnosis due to a shortage of pathologists.
A computer-implemented image analysis pipeline that processes image datasets of tissue samples to extract phenotypic signatures by analyzing morphometric, neighborhood, intensity, and texture parameters, enabling classification of tissue samples into similarity groups and analysis of patient outcome deviations.
The method provides accurate and efficient identification of unique phenotypic signatures in tissue samples, aiding in the classification of patients and the identification of novel biomarkers for personalized therapies.
Smart Images

Figure FI2024050578_08052025_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR AN IMAGE ANALYSIS PIPELINE
[0002] TECHNICAL FIELD
[0003] The present disclosure relates to the field of image analysis of biological samples. Some example embodiments relate to a multiplexed fluorescent immunohistochemical image analysis platform.
[0004] BACKGROUND
[0005] Different types of cells have characteristic shapes, and shape may indicate health of a cell or a tissue. Irregular shapes may be an indication of an injury, a condition, or a disease affecting the cell. A dying cell may appear shrunken, whereas some diseases may enlarge cells. These features have been used by pathologists as a diagnostic and prognostic tool for decades, mainly based on manual, microscopy-based inspection of basic histological stains, or staining for one or two validated disease-specific biomarkers. These methods are both highly manual and time-consuming, leading to bottlenecks in diagnosis that, due to a mounting shortage of pathologists, may have repercussions throughout healthcare systems.
[0006] There is a need for an image analysis pipeline to enable better molecular understanding of diseases and for development of prognostic indicators to aid treatment decisions.
[0007] SUMMARY
[0008] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0009] Example embodiments of the present disclosure enable an image analysis pipeline enabling providing information on unique phenotypic signatures associated with tissue samples. This benefit may be achieved by the features of the independent claims. Further example embodiments are provided in the dependent claims, the detailed description, and the drawings. According to a first aspect, a computer-implemented method for an image analysis pipeline comprises: Receiving an image dataset corresponding to a plurality of tissue samples and outcome information corresponding to the plurality of tissue samples, wherein the image dataset comprises at least one image stack corresponding to at least one staining cycle, and wherein the at least one image stack comprises one or more single-channel images associated with one or more biological stains. Extracting, from the image dataset, a plurality of regions of interest, at least one region of interest per tissue sample, wherein a region of interest within the plurality of regions of interest comprises at least one region of interest image stack corresponding to the at least one image stack. Determining, per region of interest within the plurality of regions of interest, based on at least one single-channel image within the at least one region of interest image stack corresponding to the region of interest, a plurality of cells within the region of interest. Determining, per region of interest within the plurality of regions of interest, per cell within the plurality of cells within the region of interest, values for a plurality of morphometric parameters, values for a plurality of neighborhood parameters, and values for a plurality of intensity and texture parameters. Determining, based on the plurality of morphometric parameter values, the plurality of neighborhood parameter values, and the plurality of the intensity and texture parameter values, a plurality of single-cell clusters. Quantifying, per region of interest within the plurality of regions of interest, per single-cell cluster within the plurality of single-cell clusters, a proportion of a subset of cells within the plurality of cells within the region of interest associated with the single-cell cluster. Determining, per tissue sample within the plurality of tissue samples, based on the quantified proportions, a phenotypic signature associated with the tissue sample. Classifying, based on the phenotypic signatures, the plurality of tissue samples into at least two similarity groups. Analyzing, based on the outcome information corresponding to the plurality of tissue samples, deviation in patient outcome between the at least two similarity groups. Providing output indicating the phenotypic signatures, the at least two similarity groups, and the deviation in patient outcome. With such a method, information may be provided on unique phenotypic signatures associated with tissue samples combined with patient outcome. The method may allow a person skilled in the art to draw more detailed conclusions from large patient datasets. The method may allow pharmaceutical companies, biotechnology firms, or research organizations to identify novel sets of biomarkers and to understand which patients would benefit from emerging personalized therapies.
[0010] According to an example embodiment of the first aspect, the providing the output comprises visualizing the output on a display. With such a method, the output may be provided, e.g., in full, as a summary, as images, graphs, text, parameters, or figures, which may aid the person skilled in the art to understand and / or utilize the output.
[0011] According to an example embodiment of the first aspect, a computer- implemented method further comprises: Identifying, based on the phenotypic signatures, the at least two similarity groups, and the deviation in patient outcome, one or more clinically significant parameters from within at least the plurality of morphometric parameters, the plurality of neighborhood parameters and the plurality of intensity and texture parameters. Providing further output indicating the one or more clinically significant parameters. With such a method, clinically significant parameters may enable understanding which biological features contribute most to patient outcomes.
[0012] According to an example embodiment of the first aspect, the image dataset comprises two or more image stacks corresponding to two or more staining cycles, and a region of interest within the plurality of regions of interest comprises two or more region of interest image stacks corresponding to the two or more images stacks. A computer-implemented method further comprises: Aligning, per region of interest within the plurality of regions of interest, the two or more region of interest image stacks comprised in the region of interest, to obtain one aligned region of interest image stack. With such a method, errors on alignment that may have occurred during obtaining the image dataset may be corrected to ensure an accurate image analysis.
[0013] According to an example embodiment of the first aspect, a shared staining agent has been used in the at least two staining cycles, and the aligning the two or more region of interest image stacks is performed based on the shared staining agent. With such a method, errors on alignment during obtaining the image dataset may be corrected based on the shared staining agent. According to an example embodiment of the first aspect, a computer- implemented method further comprises: Determining, per region of interest within the plurality of regions of interest, whether the region of interest comprises a presence of adequate data. Removing, in response to the region of interest not comprising the presence of adequate data, the region of interest from within the plurality of regions of interest. With such a method, further image analysis may be accelerated. Additionally, image artifacts being incorrectly identified as cells may be prevented.
[0014] According to an example embodiment of the first aspect, a computer- implemented method further comprises: Assigning, per cell within the plurality of cells within the plurality of regions of interest, based on the plurality of single-cell clusters, an identity within at least two identities with the cell. Quantifying, per identity within the at least two identities, per region of interest within the plurality of regions of interest, per single-cell cluster within the plurality of single-cell clusters, the proportion of a subset of cells assigned with the identity within the plurality of cells within the region of interest associated with the single-cell cluster. With such a method, assigning cells with identities may enable determining more accurate phenotypic signatures.
[0015] According to an example embodiment of the first aspect, a computer- implemented method further comprises: Calculating, per cell within the plurality of cells within the plurality of regions of interest, distance to a nearest cell, wherein the cell has a first identity from within the at least two identities and the nearest cell has a second identity from within the at least two identities. Modifying, based on the calculated distances, the determined plurality of singlecell clusters. With such a method, spatial metrics may be taken into consideration in the single-cell clustering.
[0016] According to an example embodiment of the first aspect, the at least two identities comprise at least a tumor identity and a stroma identity. With such a method, tumor tissue and its surrounding stroma may be treated separately in the image analysis pipeline.
[0017] According to an example embodiment of the first aspect, a computer- implemented method further comprises: Determining, based on similarities in the phenotypic signatures, the at least two similarity groups using a hierarchical clustering method. With such a method, the similarity groups may be determined based on their common features.
[0018] According to an example embodiment of the first aspect, the determining the plurality of cells comprises: Segmenting a plurality of subcellular structures comprising at least one of: a nucleus, a cytoplasm, a nuclear envelope, and an endoplasmatic reticulum. With such a method, the plurality of cells may be determined accurately.
[0019] According to an example embodiment of the first aspect, the outcome information comprises information on at least one of: survival, remission, and recurrence. With such a method, different types of patient outcomes may be analyzed.
[0020] According to an example embodiment of the first aspect, the image dataset has been obtained by multiplexed fluorescence immunohistochemical imaging.
[0021] According to a second aspect, an apparatus comprises: At least one processor and at least one memory including computer program code. The at least one memory and computer program code are configured to, with the at least one processor, cause the apparatus at least to perform: Receiving an image dataset corresponding to a plurality of tissue samples and outcome information corresponding to the plurality of tissue samples, wherein the image dataset comprises at least one image stack corresponding to at least one staining cycle, and wherein the at least one image stack comprises one or more single-channel images associated with one or more biological stains. Extracting, from the image dataset, a plurality of regions of interest, at least one region of interest per tissue sample, wherein a region of interest within the plurality of regions of interest comprises at least one region of interest image stack corresponding to the at least one image stack. Determining, per region of interest within the plurality of regions of interest, based on at least one single-channel image within the at least one region of interest image stack corresponding to the region of interest, a plurality of cells within the region of interest. Determining, per region of interest within the plurality of regions of interest, per cell within the plurality of cells within the region of interest, values for a plurality of morphometric parameters, values for a plurality of neighborhood parameters, and values for a plurality of intensity and texture parameters. Determining, based on the plurality of morphometric parameter values, the plurality of neighborhood parameter values, and the plurality of the intensity and texture parameter values, a plurality of single-cell clusters. Quantifying, per region of interest within the plurality of regions of interest, per single-cell cluster within the plurality of single-cell clusters, a proportion of a subset of cells within the plurality of cells within the region of interest associated with the single-cell cluster. Determining, per tissue sample within the plurality of tissue samples, based on the quantified proportions, a phenotypic signature associated with the tissue sample. Classifying, based on the phenotypic signatures, the plurality of tissue samples into at least two similarity groups. Analyzing, based on the outcome information corresponding to the plurality of tissue samples, deviation in patient outcome between the at least two similarity groups. Providing output indicating the phenotypic signatures, the at least two similarity groups, and the deviation in patient outcome.
[0022] According to a third aspect, a computer program product comprises a computer-readable program code configured to, when read and executed by a computer system, cause the computer system at least to perform: Receiving an image dataset corresponding to a plurality of tissue samples and outcome information corresponding to the plurality of tissue samples, wherein the image dataset comprises at least one image stack corresponding to at least one staining cycle, and wherein the at least one image stack comprises one or more singlechannel images associated with one or more biological stains. Extracting, from the image dataset, a plurality of regions of interest, at least one region of interest per tissue sample, wherein a region of interest within the plurality of regions of interest comprises at least one region of interest image stack corresponding to the at least one image stack. Determining, per region of interest within the plurality of regions of interest, based on at least one single-channel image within the at least one region of interest image stack corresponding to the region of interest, a plurality of cells within the region of interest. Determining, per region of interest within the plurality of regions of interest, per cell within the plurality of cells within the region of interest, values for a plurality of morphometric parameters, values for a plurality of neighborhood parameters, and values for a plurality of intensity and texture parameters. Determining, based on the plurality of morphometric parameter values, the plurality of neighborhood parameter values, and the plurality of intensity and texture parameter values, a plurality of single-cell clusters. Quantifying, per region of interest within the plurality of regions of interest, per single-cell cluster within the plurality of single-cell clusters, a proportion of a subset of cells within the plurality of cells within the region of interest associated with the single-cell cluster. Determining, per tissue sample within the plurality of tissue samples, based on the quantified proportions, a phenotypic signature associated with the tissue sample. Classifying, based on the phenotypic signatures, the plurality of tissue samples into at least two similarity groups. Analyzing, based on the outcome information corresponding to the plurality of tissue samples, deviation in patient outcome between the at least two similarity groups. Providing output indicating the phenotypic signatures, the at least two similarity groups, and the deviation in patient outcome.
[0023] Any example embodiment may be combined with one or more other example embodiments. Many of the attendant features will be more readily appreciated as they become better understood by reference to the following detailed description considered in connection with the accompanying drawings.
[0024] DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are included to provide a further understanding of the example embodiments and constitute a part of this specification, illustrate example embodiments and together with the description help to understand the example embodiments. In the drawings:
[0026] FIG. 1 illustrates a schematic block diagram of an example environment suitable for an image analysis platform;
[0027] FIG. 2 illustrates a schematic block diagram of an apparatus configured to practice one or more example embodiments;
[0028] FIG. 3 illustrates a schematic flow chart of an image analysis pipeline according to an example embodiment;
[0029] FIG. 4 illustrates a schematic flow chart of an image analysis pipeline according to another example embodiment;
[0030] FIG. 5 illustrates a schematic flow chart of an image analysis pipeline according to another example embodiment; FIG. 6 illustrates a schematic flow chart of an image analysis pipeline according to another example embodiment;
[0031] FIG. 7 illustrates a schematic flow chart of an image analysis pipeline according to another example embodiment; and
[0032] FIG. 8 illustrates a schematic flow chart of an image analysis pipeline according to another example embodiment.
[0033] DETAILED DESCRIPTION
[0034] Reference will now be made in detail to example embodiments, examples of which are illustrated in the accompanying drawings. The detailed description provided below in connection with the appended drawings is intended as a description of the present examples and is not intended to represent the only forms in which the present example may be constructed or utilized. The description sets forth the functions of the example and the sequence of steps for constructing and operating the example. However, the same or equivalent functions and sequences may be accomplished by different examples.
[0035] Although the specification may refer to “an”, “one”, or “some” embodiment(s) in several locations, this does not necessarily mean that each such reference is to the same embodiment(s), or that the feature may not apply to other embodiments. Single features of different embodiments may also be combined to provide other embodiments. Furthermore, words “comprising” and “including” should be understood as not limiting the described embodiments / examples, and the embodiments / examples may contain also features / structures that have not been specifically mentioned.
[0036] Furthermore, although the numerative terminology, such as “first”, “second”, etc., may be used herein to describe various embodiments, elements, or features, it should be understood that these embodiments, elements, or features should not be limited by this numerative terminology. This numerative terminology is used herein only to distinguish one embodiment, element, or feature from another embodiment, element, or feature. For example, a first identity discussed below could be called a second identity, and vice versa, without departing from the teachings of the present disclosure.
[0037] According to an example embodiment, a computer-implemented method for an image analysis pipeline is disclosed. An image dataset corresponding to a plurality of tissue samples is received, processed, and analyzed to generate phenotypic signatures (tumor fingerprints) associated with the plurality of tissue samples. The image dataset may have been obtained by multiplexed fluorescence immunohistochemical imaging. Alternatively, the image dataset may have been obtained by a different imaging technique such as, e.g., imaging mass cytometry. The plurality of tissue samples may be clinical tissue samples such as, e.g., surgical biopsies in the form of e.g., whole-slide sections or tumor microarray (TMA) sections. The phenotypic signatures may be understood as personalized descriptors of patients’ unique tumors. The phenotypic signatures are determined based on single-cell clusters that are determined using methods adapted from single-cell transcriptomic methods based on various parameter values determined in the image dataset. Outcome information corresponding to the plurality of tissue samples is used to classify the tissue samples into similarity groups, and deviation in patient outcome between the similarity groups is analyzed. Output indicating at least the phenotypic signatures, the similarity groups, and the deviation in patient outcome is provided.
[0038] In the context of this specification, the term “multiplexed” may be understood as simultaneous or sequential detection, analysis, or representation of multiple targets, signals, or entities.
[0039] In the context of this specification, the term "histochemistry" may be understood as a study and analysis of chemical composition of biological tissues using color-reactive chemicals, such as dyes or reagents, that produce a visually discernible reaction, such as a change in color, in reaction with a particular molecule. Histochemistry may be understood as in-situ analysis of biological tissues by use of light microscopy.
[0040] In the context of this specification, the term “light microscopy” may be understood as a technique that uses visible light and a system of optical lenses to magnify and produce detailed images of small objects or samples. The basic principle involves transmission or reflection of light through or off a sample, respectively, which then passes through a series of lenses that magnify the image for observation by, e.g., human eye, camera, or other detection system. Variations of light microscopy include bright-field, dark-field, phase-contrast, differential interference contrast, and fluorescence microscopy, among others. In the context of this specification, the term “immunohistochemistry” may be understood as a specialized technique within histochemistry that couples color-reactive chemicals to antibodies to specifically detect and localize proteins, peptides, or other antigens in tissue samples. Immunohistochemistry may be understood as an in-situ analysis of protein expression by light microscopy.
[0041] In the context of this specification, the term “fluorescence histochemistry” or “fluorescence histochemical” may be understood as a subset of histochemistry wherein fluorescent dyes or markers are used to label and visualize specific chemical or biological components within tissue samples. Fluorescence may be understood as in-situ analysis of biological tissues by use of fluorescence microscopy. In fluorescence microscopy, the sample is illuminated with a specific wavelength of light that excites fluorophores present. These fluorophores then emit light of a different wavelength, which is separated from excitation light using optical filters and then captured to form an image. Fluorescence histochemistry may allow for visualization of specific molecules, cells, or structures in a sample that has been labelled with fluorophores.
[0042] In the context of this specification, the term “fluorophores” or “fluorescent markers” may be understood as chemical compounds or single molecules that can absorb light at a specific wavelength and subsequently reemit light at a different, typically longer, wavelength. Absorption, excitation, and emission wavelengths may be specific for each fluorophore, and while these wavelengths are discrete for monoatomic fluorophores, polyatomic fluorophores exhibit broad excitation and emission spectra. The fluorophores may be conjugated to antibodies to detect and visualize specific proteins, peptides, or other antigens in a tissue sample. The fluorophores may be organic dyes, molecules with intrinsic binding affinity for native cellular biomolecules, genetically encoded fluorescent proteins, affinity-based tag-ligand systems, or enzyme-based tag-ligand systems. The diverse range in fluorescence emissions may allow for multiplexed imaging, where multiple fluorophores may be used concurrently to study different targets within same sample or same field of view.
[0043] In the context of this specification, the term “histochemical imaging” may be understood as a process of creating visual representations or images of tissue samples based on their chemical properties as revealed through histochemical techniques. Histochemical imaging may encompass colorimetric or fluorescence-based techniques. Histochemical imaging may be a label dependent or a label-free method.
[0044] In the context of this specification, the term "multiplexed fluorescence histochemical imaging" may be understood as an imaging technique wherein multiple fluorophores are utilized simultaneously or sequentially to label and visualize distinct chemical or biological components within a tissue sample or a field of view. By detecting signals from multiple fluorophores, the imaging technique may allow for spatial resolution of multiple targets within a single tissue section or a field of view and may allow acquisition of information of spatial and temporal distribution of multiple components, molecules, or structures within the tissue sample or field of view.
[0045] In the context of this specification, the term “whole-slide imaging” may be understood as an imaging technique capturing high-resolution images of entire pathology or histology glass slides. This method may employ bright-field or fluorescence microscopy to capture digital images of tissue samples.
[0046] In the context of this specification, the term “imaging mass cytometry” may be understood as referring to an analytical technique that combines cytometry and time-of-flight mass spectrometry. This technique may be used to analyze and visualize expression and distribution of multiple proteins at a single-cell resolution in tissue sections through use of metal-conjugated antibodies and ultraviolet (UV) laser ablation.
[0047] In the context of this specification, the term “single-cell transcriptomics” may be understood as a study, analysis, or quantification of a plurality of messenger ribonucleic acid (mRNA) transcripts in an individual cell at a time of mRNA extraction. Varied methods for single-cell transcriptomics exist, however some form of cell isolation and unique molecular identifiers to identify a cell of origin for each mRNA after sequencing may be required. Single-cell transcriptomics data analysis involves detecting cell-to-cell variation, within a typically large cell population, for the expression level of each transcript or gene. Single-cell methods are characterized by the individual cell being an only sample, without a possibility of replicates, which may present unique challenges. Data may have a large number of missing values and can include more technical noise (uneven amplification or batch effects) to be corrected for, and typical analysis requires clustering to other cells for a comprehensive analysis of a cell withing a cell cluster.
[0048] In the context of this specification, the term “single-cell” may be understood as an individual, isolated biological unit or entity, typically referring to a singular prokaryotic or eukaryotic cell, distinguished from multi-cellular aggregates or entities.
[0049] In the context of this specification, the term “cluster” may be understood as a collection or grouping of closely associated entities, which can be cells, data points, data objects, or other related units, that coalesce or come together based on similarity of certain characteristics or properties and are in some way dissimilar to other cells, data points, data objects or other related units in the data set. The entities within a cluster exhibit greater similarity to one another compared to those in other clusters.
[0050] In the context of this specification, the term “single-cell cluster” may be understood as a grouping (cluster) of individual cells that are more similar based on specific characteristics or properties than cells in other groupings (clusters), wherein a cell within the cluster maintains its distinct identity. The specific characteristics or properties used to cluster cells may be different in each experiment or condition. The individual cell may be clustered with different cells in another experiment, if the characteristics or properties used to cluster the cells are different or the method for clustering is different.
[0051] In the context of this specification, the term “clustering” may be understood as a process or a method by which entities, which can be cells, data points, or other related units, are grouped (clustered) or classified together based on shared characteristics, attributes, or patterns. Clustering may be done with varied statistical clustering algorithms. An appropriate clustering algorithm may be chosen experimentally for each type of data in each experiment.
[0052] In the context of this specification, the term "morphometric parameters" may be understood as quantitative measurements that describe a shape, a size, and / or structural features of an object or specimen.
[0053] In the context of this specification, the term "neighborhood parameters" may be understood as measurements and / or characteristics that describe a position of a particular object within a specimen or sample, or a surrounding context or an environment of the particular object or the specimen or a region within a given dataset or image.
[0054] In the context of this specification, the term “intensity” may be understood as a measure of emitted wavelengths from a substance or sample within an object, or it may be understood as a measure of quantity of staining. The intensity may be related to a concentration or an abundance of molecules and / or markers present in the substance or sample. The intensity may be measured, e.g., as average intensity within the object.
[0055] In the context of this specification, the term “texture” may be understood as a measure of spatial distribution of intensity within an object.
[0056] In the context of this specification, the term "phenotypic signature" may be understood as a unique combination of observable traits or characteristics of a tissue, a cell, or a subcellular structure such as, e.g., a nucleus, a nuclear lamina, an endoplasmatic reticulum, a plasma membrane, or a cell adhesion. The unique combination may result from interactions between genotype, history, and / or environment or the tissue or cell.
[0057] In the context of this specification, the term “biomarker” may refer to, e.g., a morphometric, a phenotypic, a genotypic, a chemical, or a molecular marker, or a combination thereof, that may be associated with a disease or a condition or a risk of having or developing thereof. It does not necessarily refer to a biomarker that would be statistically fully validated as having a specific effectiveness in a clinical setting. The biomarker may be understood as a cell shape, a cell feature, a morphometric parameter, an intensity parameter, a texture parameter, a neighborhood parameter, a compound, a protein, a gene, a moiety, a functional group, a composition, a combination of two or more genes and / or compound features and / or parameters, a measurable or measured quantity thereof, a ratio or other value derived thereof, or in principle any measurement reflecting a chemical and / or biological component that may be found associated with a disease or condition or a risk of having or developing thereof. The biomarkers and any combinations thereof, optionally in combination with further analyses and / or measures, may be used to measure a biological process indicative of disease risk, disease occurrence, disease progression, or future risk of disease. In the context of this specification, the term "clinically significant parameters" may be understood as specific measurable factors, values, or characteristics related to health, a condition, a disease, or a physiological state, which are associated with or have an impact on the diagnosis, prognosis, treatment, or understanding of a patient's medical condition or health status. The parameters may, for example, be values of biomarkers or compound biomarkers that have previously shown to be associated with a disease or prognosis.
[0058] In the context of this specification, the term "image channel" may be understood as a distinct layer within an image that contains specific data or information, such as color spaces or wavelengths, e.g., such as red, green, and / or blue channels within a color image. The image channel may be understood to be a fluorescence channel, wherein a range of wavelengths may be used to capture or measure emitted light from a fluorescent substance or sample.
[0059] In the context of this specification, the term "image stack" may be understood as a collection of two-dimensional images arranged sequentially.
[0060] In the context of this specification, the term "biological stain" may be understood as a specific dye or pigment applied to a biological specimen or sample, such as a tissue or a cell, to enhance contrast and visualize particular cell populations, components, or structures in, e.g., light microscopy.
[0061] In the context of this specification, the term "staining cycle" may be understood as a process or series of steps wherein a specimen, e.g., a biological sample, is treated with one or more stains or dyes to enhance visibility or differentiation of specific structures or components within the specimen.
[0062] In the context of this specification, the term "region of interest" may be understood as a central or primary component of a larger system or structure. This may refer to a tumor or a stroma, a main portion of a biological sample, such as a tissue sample, or an area within the tissue sample. The region of interest may be a circular area within the tissue sample, that may be obtained by needle aspiration from a larger paraffin embedded tissue section or that may be a circled area indicated by a pathologist.
[0063] In the context of this specification, the term "tissue sample" may be understood as a small portion or section of a biological tissue, obtained from an organism by, e.g., a biopsy. According to the example embodiment, generating phenotypic signatures based on single-cell clusters determined using various parameter values determined in the image dataset representing, e.g., intensity, texture, morphology, and / or neighborhood features may provide more accurate information for the user than using staining intensity values of the image dataset alone. Patients may be classified based on biological features of their tissue samples. With such a method, novel biological disease subtypes differing in their prognosis may be identified. With such a method, new biological targets may be identified. The method may enable classifying patients into patient groups that are responders / non-responders to certain treatments. The method may enable identifying diagnostic biomarkers for determining which patient should be given which treatments. According to some example embodiments, the method may be used to identify disease subtypes with novel diagnostic biomarkers that differ in their prognosis and / or treatment options.
[0064] Different example embodiments are described below using single units, models, equipment, and memory, without restricting the example embodiments to such a solution. Concepts called cloud computing and / or virtualization may be used. The virtualization may allow a single physical computing device to host one or more instances of virtual machines that appear and operate as independent computing devices, so that a single physical computing device can create, maintain, delete, or otherwise manage virtual machines in a dynamic manner. It is also possible that device operations will be distributed among a plurality of servers, nodes, devices, or hosts. In cloud computing network devices, computing devices and / or storage devices provide shared resources. Some other technology advancements, such as Software-Defined Networking (SON), may cause one or more of the functionalities described below to be migrated to any corresponding abstraction or apparatus or device. Correspondingly, Web 3.0, also known as the third-generation internet, implementing for example blockchain technology, may cause one or more of the functionalities described below to be distributed across a plurality of apparatuses or devices. Therefore, all words and expressions should be interpreted broadly, and they are intended to illustrate, not to restrict, the example embodiment. FIG. 1 illustrates an example embodiment of a general exemplary environment that is suitable for analyzing large image datasets and would benefit from the image analysis pipeline. FIG. 1 presents a simplified environment showing some devices, apparatuses, and functional entities, all being logical units whose implementation and / or number may differ from what is shown. Any suitable protocols, elements, equipment, functions, and / or structures may be used to implement the environment. It is apparent to a person skilled in the art that the environment comprises any number of shown elements, other equipment, other functions, and structures that are not illustrated. They, as well as the protocols used, are well-known by persons skilled in the art, and are irrelevant to the actual invention. Therefore, they need not to be discussed in more detail here.
[0065] In the example embodiment illustrated in FIG. 1, the environment 100 comprises at least one user device 110 connectable over one or more networks 120 to at least one backend equipment 130 comprising imaging data 132.
[0066] A user device 110 refers to a computing device (equipment, apparatus) that may be a portable device or a desktop device such as personal computer, and it may also be referred to as a user terminal or a user apparatus. Portable computing devices (apparatuses) include wireless mobile communication devices operating with or without a subscriber identification module (SIM) in hardware or in software, including, but not limited to, the following types of devices: laptop computer, touch screen computer and tablet (tablet computer). The user device 110 may comprise one or more user interfaces. The one or more user interfaces may be any kind of a user interface, e.g., a screen, a keypad, a loudspeaker, a microphone, a touch user interface, an integrated display device, and / or external display device. The user device 110 is configured to support receiving and analyzing imaging data, e.g., by carrying out methods described in more detail below. For that purpose, the user device 110 may be capable of installing applications. The user device may comprise imaging data 132a, permanently or temporarily.
[0067] A network 120 may be any wired or wireless network, or a combination thereof, enabling transmission of information between different apparatuses / devices over the network. These include, but are not limited to, local area networks (LAN), cellular networks, and wireless local area networks (WLAN).
[0068] A backend equipment 130 is configured to comprise imaging data 132 and may send the imaging data 132 or part of the imaging data 132a to the user device 110 and / or receive the imaging data 132 or part of the imaging data 132a from the user device 110. For that purpose, the backend equipment 130 comprises a memory and a processor coupled to the memory. The memory may be any kind of conventional or future data repository, including distributed and centralized storing of data, managed by any suitable management system forming part of the backend equipment 130. An example of distributed storing includes a cloud-based storage in a cloud environment (which may be, e.g., a public cloud, a community cloud, a private cloud, or a hybrid cloud). Cloud storage services may be accessed through a co-located cloud computer service, a web service application programming interface (API) or by applications that utilize API, such as cloud desktop storage, a cloud storage gateway or Webbased content management systems. Further, the backend equipment 130 may comprise several servers with databases, which may be integrated to be visible to a user device as one database and one database server. However, the manner in which data structures are stored and retrieved, and the location where different pieces of data to be obtained for the image analysis pipeline are irrelevant to the invention. The processor may comprise one or more processing cores containing circuitry configured to execute instructions. The memory stores processor executable instructions to be executed by the processor to analyze the imaging data 132 as will be described in detail below. The imaging data 132 may be encoded in the backend equipment 130 to a format suitable for the user device 110, the backend equipment may obtain the data from another device or server in a format suitable for the user device 110, or the data may be encoded to a suitable format in the user device 110.
[0069] FIG. 2 illustrates an example embodiment of an apparatus 200 configured to perform operations of one or more example embodiments, e.g., functionalities described below with reference to FIG. 3 to 8. The apparatus 200 may be, e.g., used to implement the user device 110. The apparatus 200 may comprise at least one processor 202. The at least one processor 202 may comprise, for example, one or more of various processing devices or processor circuitry, such as for example a co-processor, a microprocessor, a controller, a digital signal processor (DSP), a processing circuitry with or without an accompanying DSP, or various other processing devices including integrated circuits such as, for example, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a microcontroller unit (MCU), a hardware accelerator, a special-purpose computer chip, or the like.
[0070] The apparatus 200 may further comprise at least one memory 204. The at least one memory 204 may be configured to store, for example, computer program code or the like, for example operating system software and application software. The at least one memory 204 may comprise one or more volatile memory devices, one or more non-volatile memory devices, and / or a combination thereof. For example, the at least one memory 204 may be embodied as magnetic storage devices (such as hard disk drives, floppy disks, magnetic tapes, etc.), optical magnetic storage devices, or semiconductor memories (such as mask ROM, PROM (programmable ROM), EPROM (erasable PROM), flash ROM, RAM (random access memory), etc.).
[0071] The apparatus 200 may further comprise a communication interface 208 configured to enable apparatus 200 to transmit and / or receive information to / from other devices, functions, or entities. In one example, the apparatus 200 may use communication interface 208 to transmit or receive information over a service-based interface (SBI) message bus of the network 120. The communication interface 208 may therefore comprise a data communication interface and be configured for communication between devices, for example according to one or more data communication protocols. The apparatus 200 may further comprise a user interface 210, for example for providing user output by the apparatus 200, such as for example visual and / or audible signal(s), for example by speaker(s), display(s), light(s), or the like. The user interface 210 may be any kind of a user interface, e.g., a screen, a keypad, a loudspeaker, a microphone, a touch user interface, an integrated display device, and / or external display device. The user interface 210 may be used for example for outputting indications of phenotypic signatures and patient outcomes to a human user.
[0072] When the apparatus 200 is configured to implement some functionality, some component and / or components of the apparatus 200, such as for example the at least one processor 202 and / or the at least one memory 204, may be configured to implement this functionality. Furthermore, when the at least one processor 202 is configured to implement some functionality, this functionality may be implemented using program code 206 comprised, for example, in the at least one memory 204.
[0073] The functionality described herein may be performed, at least in part, by one or more computer program product components such as for example software components. According to an example embodiment, the apparatus 200 comprises a processor or processor circuitry, such as for example a microcontroller, configured by the program code when executed to execute the embodiments of the operations and functionality described. A computer program or a computer program product may therefore comprise instructions for causing, when executed, the apparatus 200 to perform the method(s) described herein. Alternatively, or in addition, the functionality described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), application-specific Integrated Circuits (ASICs), application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), Graphics Processing Units (GPUs).
[0074] The apparatus 200 comprises means for performing at least one method described herein. In one example, the means comprises the at least one processor 202, the at least one memory 204 including the program code 206 configured to, when executed by the at least one processor, cause the apparatus 200 to perform the method. Although the apparatus 200 is illustrated as a single device it is appreciated that, wherever applicable, functions of the apparatus 200 may be distributed to a plurality of devices, e.g., to implement example embodiments as a cloud computing service.
[0075] FIG. 3 illustrates a flowchart according to an example embodiment of a computer-implemented method for an image analysis pipeline. The apparatus 200 illustrated with FIG. 2 is configured to perform the operation of the method in the example environment 100 illustrated with FIG. 1.
[0076] Referring to FIG. 3, an image dataset corresponding to a plurality of tissue samples and outcome information corresponding to the plurality of tissue samples is received in operation 301. The plurality of tissue samples may comprise patient biopsies such as, e.g., tumor biopsies. The image dataset may have been obtained by, e.g., multiplexed fluorescence histochemical imaging.
[0077] In an example embodiment, the outcome information may comprise clinical patient information on at least one of: survival, remission, and recurrence. Outcome information corresponding to a patient may be associated with a tissue sample corresponding to the patient on a later date than when the tissue sample is obtained from the patient, and the outcome information may comprise information not yet available when obtaining the tissue sample. The image dataset comprises at least one image stack corresponding to at least one staining cycle. The at least one image stack comprises one or more singlechannel images associated with one or more biological stains. In an example embodiment, the one or more single-channel images may be fluorescence channel images.
[0078] Referring to FIG. 3, a plurality of regions of interest (ROIs) is extracted in operation 302 from the image dataset. The plurality of regions of interest corresponds to the plurality of tissue samples, at least one region of interest per tissue sample. A region of interest within the plurality of regions of interest comprises at least one region of interest image stack corresponding to the at least one image stack. A region of interest within the plurality of regions of interest may correspond to the tissue sample or a part of the tissue sample.
[0079] Referring to FIG. 3, per region of interest within the plurality of regions of interest, a plurality of cells within the region of interest is determined in operation 303, based on at least one single-channel image within the at least one region of interest image stack corresponding to the region of interest. In an example embodiment, the determining the plurality of cells comprises segmenting a plurality of subcellular structures comprising at least one of: a nucleus, a cytoplasm, a nuclear envelope, and an endoplasmatic reticulum. The determining the plurality of cells may comprise first segmenting a plurality of nuclei and then expanding the segmented plurality of nuclei by a pre-determined number of pixels, to obtain a plurality of nuclear lamina and / or a plurality of cytoplasms.
[0080] Referring to FIG. 3, values for a plurality of morphometric parameters, values for a plurality of neighborhood parameters, and values for a plurality of intensity and texture parameters are determined in operation 304, per region of interest within the plurality of regions of interest, per cell within the plurality of cells within the region of interest. The plurality of morphometric parameters may comprise, e.g., perimeter, area, circularity, roundness, solidity, compactness, aspect ratio, and / or a series of a plurality of harmonic ellipses describing nuclear contour which are summarized by deriving an elliptic Fourier coefficient (EFC) ratio. The plurality of neighborhood parameters may comprise, e.g., a position of the region of interest within the tissue sample, a position of a cell within the plurality of cells within the region of interest, an alignment of the plurality of cells, a number of neighbors within a given radius, and / or a distance to nearest neighboring nucleus. The plurality of intensity and texture parameters may comprise, e.g., fluorescence intensity and texture parameters and / or staining quantity and texture parameters.
[0081] In an example embodiment, image artefacts corresponding to, e.g., a wrinkled and / or folded tissue sample or part of a tissue sample within the plurality of tissue samples, and / or aberrant image data such as, e.g., fluorescence intensity parameter values corresponding to autofluorescence may be detected and removed from the at least one region of interest image stack. In an example embodiment, correlated and / or redundant parameters within the plurality of morphometric parameters, the plurality of neighborhood parameters, and the plurality of intensity and texture parameter may be removed using, e.g., Pearson’s correlation analysis. In an example embodiment, the plurality of morphometric parameter values, the plurality of neighborhood parameter values, and the plurality of intensity and texture parameter values may be normalized to remove sampling effect and scaled to obtain relative parameter expression between the plurality of cells.
[0082] Referring to FIG. 3, a plurality of single-cell clusters is determined in operation 305 based on the plurality of morphometric parameter values, the plurality of neighborhood parameter values, and the plurality of the intensity and texture parameter values. The single-cell clusters are determined using methods adapted for imaging data based on single-cell transcriptomic methods. In an example embodiment, the methods may comprise principal component analysis (PCA). In an example embodiment, the methods may comprise Uniform Manifold Approximation and Projection (UMAP) dimensionality reduction. This may enable visualizing similarities and / or differences between the plurality of cells.
[0083] Referring to FIG. 3, per region of interest within the plurality of regions of interest, per single-cell cluster within the plurality of determined single-cell clusters, a proportion of a subset of cells within the plurality of cells within the region of interest associated with the single-cell cluster is quantified in operation 306. The proportion of the subset of cells may be understood as the percentage of the plurality of cells within the region of interest that are associated with the single-cell cluster within the plurality of single-cell clusters, and thus the total amount of the proportions quantified for one region of interest within the plurality of regions of interest should be 100%. A phenotypic signature is determined in operation 307, per tissue sample within the plurality of tissue samples, based on the quantified proportions. If a tissue sample corresponds to one region of interest within the plurality of regions of interest, the phenotypic signature associated with the tissue sample corresponds to the quantified proportions of the one region of interest. If the tissue sample corresponds to two or more regions of interest within the plurality of regions of interest, the phenotypic signature associated with the tissue sample corresponds to the quantified proportions of the two or more regions of interest. The plurality of tissue samples is classified in operation 308 into at least two similarity groups, based on the phenotypic signatures associated with the tissue samples.
[0084] In an example embodiment, values for tissue shape pattern parameters may be determined, per region of interest within the plurality of regions of interest. The tissue shape pattern parameters may be clustered to obtain a plurality of tissue architecture classes. The tissue samples may be associated with a tissue architecture class within the plurality of tissue architecture classes. The plurality of tissue samples may be classified within operation 308 into the at least two similarity groups, based on the phenotypic signatures associated with the tissue samples and the tissue architecture classes associated with the tissue samples.
[0085] Referring to FIG. 3, deviation in patient outcome between the at least two similarity groups is analyzed in operation 309, based on the outcome information corresponding to the plurality of tissue samples. Output indicating the phenotypic signatures, the at least two similarity groups, and the deviation in patient outcome is provided in operation 310. In an example embodiment, the output may be visualized on a display. The visualization may enable the user to compare differences across, e.g., tumors or patients.
[0086] FIG. 4 illustrates a flowchart according to an example embodiment of a computer-implemented method for an image analysis pipeline comprising identifying clinically significant parameters. The apparatus 200 illustrated with FIG. 2 is configured to perform the operation of the method in the example environment 100 illustrated with FIG. 1.
[0087] Referring to FIG. 4, the process continues from operation 310 in FIG. 3. One or more clinically significant parameters from within at least the plurality of morphometric parameters, the plurality of neighborhood parameters, and the plurality of intensity and texture parameters are identified in operation 401, based on the phenotypic signatures, the at least two similarity groups, and the deviation in patient outcome. The one or more clinically significant parameters may be understood as parameters that contribute most to, e.g., the patient outcome. Further output indicating the one or more clinically significant parameters is provided in operation 402.
[0088] In an example embodiment, phenotypic interaction between a pair of single-cell clusters within the plurality of single-cell clusters is determined, per pair of single-cell clusters within the plurality of single-cell clusters. The further output provided in operation 402 may additionally comprise indicating the phenotypic interactions determined.
[0089] FIG. 5 illustrates a flowchart according to an example embodiment of a computer-implemented method for an image analysis pipeline wherein the image dataset comprises two or more image stacks. The apparatus 200 illustrated with FIG. 2 is configured to perform the operation of the method in the example environment 100 illustrated with FIG. 1. The functionalities illustrated in FIG. 5 may be carried out between operations 302 and 303 in FIG. 3.
[0090] Referring to FIG. 5, the two or more image stacks correspond to two or more staining cycles. A region of interest within the plurality of regions of interest comprises two or more region of interest image stacks corresponding to the two or more images stacks. The two or more region of interest image stacks are aligned in operation 501, per region of interest within the plurality of regions of interest, to obtain one aligned region of interest image stack. The process is then continued in operation 502 to operation 303 in FIG. 3.
[0091] In an example embodiment, a shared staining agent has been used in the two or more staining cycles. The aligning the two or more region of interest image stacks may be performed based on the shared staining agent.
[0092] FIG. 6 illustrates a flowchart according to an example embodiment of a computer-implemented method for an image analysis pipeline comprising quality analysis for region of interest extraction. The apparatus 200 illustrated with FIG. 2 is configured to perform the operation of the method in the example environment 100 illustrated with FIG. 1. The functionalities illustrated in FIG.
[0093] 6 may be carried out between operations 302 and 303 in FIG. 3.
[0094] Referring to FIG. 6, operations within block 601 are performed per region of interest within the plurality of regions of interest. It is determined in operation 602, whether the region of interest comprises a presence of adequate data. The adequate data may be understood as image data corresponding to a sample. An absence of adequate data may occur, e.g., if a tissue sample has become detached during the imaging process. If the region of interest comprises the presence of the adequate data (602: yes), the region of interest is kept in operation 603 within the plurality of regions of interest. If the region of interest does not comprise the presence of the adequate data (602: no), the region of interest is removed in operation 604 from within the plurality of regions of interest. When operations 602 to 604 have been performed for the plurality of regions of interest, the process is then continued in operation 605 to operation 303 in FIG. 3.
[0095] FIG. 7 illustrates a flowchart according to an example embodiment of a computer-implemented method for an image analysis pipeline comprising assigning cell identities to the plurality of cells. The apparatus 200 illustrated with FIG. 2 is configured to perform the operation of the method in the example environment 100 illustrated with FIG. 1. The functionalities illustrated in FIG.
[0096] 7 may be carried out between operations 305 and 307 in FIG. 3, as an alternative to operation 306 in FIG. 3.
[0097] Referring to FIG. 7, per cell within the plurality of cells within the plurality of regions of interest, an identity within at least two identities is assigned in operation 701 with the cell, based on the plurality of single-cell clusters. In an example embodiment, the at least two identities comprise at least a tumor identity and a stroma identity. Per identity within the at least two identities, per region of interest within the plurality of regions of interest, per single-cell cluster within the plurality of determined single-cell clusters, a proportion of a subset of cells assigned with the identity within the plurality of cells within the region of interest associated with the single-cell cluster is quantified in operation 702. The process is then continued in operation 703 to operation 307 in FIG. 3.
[0098] FIG. 8 illustrates a flowchart according to an example embodiment of a computer-implemented method for an image analysis pipeline wherein the single-cell clusters are modified based on calculated distances to nearest cells. The apparatus 200 illustrated with FIG. 2 is configured to perform the operation of the method in the example environment 100 illustrated with FIG. 1. The functionalities illustrated in FIG. 8 may be carried out between operations 701 and 702 in FIG. 7.
[0099] Referring to FIG. 8, distance to a nearest cell is calculated in operation 801, per cell within the plurality of cells within the plurality of regions of interest, wherein the cell has a first identity from within the at least two identities and the nearest cell has a second identity from within the at least two identities. In an example embodiment, the at least two identities comprise at least a tumor identity and a stroma identity. The determined plurality of singlecell clusters is modified in operation 802, based on the calculated distances. The modifying the plurality of single-cell clusters may be understood as reclustering the plurality of cells. The process is then continued in operation 803 to operation 702 in FIG. 7.
[0100] Although the subject matter has been described in language specific to structural features and / or acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example embodiments of implementing the claims and other equivalent features and acts are intended to be within the scope of the claims.
[0101] It will be understood that the benefits and advantages described above may relate to one example embodiment or may relate to several example embodiments. The example embodiments are not limited to those that solve any or all of the stated problems or those that have any or all of the stated benefits and advantages. It will further be understood that reference to 'an' item may refer to one or more of those items.
[0102] The steps or operations of the methods described herein may be carried out in any suitable order, or simultaneously where appropriate. Additionally, individual blocks may be deleted from any of the methods without departing from the scope of the subject matter described herein. Aspects of any of the example embodiments described above may be combined with aspects of any of the other example embodiments described to form further example embodiments without losing the effect sought.
[0103] It will be understood that the above description is given by way of example embodiments only and that various modifications may be made by those skilled in the art. The above specification, example embodiments, and data provide a complete description of the structure and use of exemplary embodiments. Although various example embodiments have been described above with a certain degree of particularity, or with reference to one or more individual embodiments, those skilled in the art could make numerous alterations to the disclosed example embodiments without departing from the scope of this specification.
Claims
CLAIMS1. A computer-implemented method for an image analysis pipeline, comprising: receiving an image dataset corresponding to a plurality of tissue samples and outcome information corresponding to the plurality of tissue samples, wherein the image dataset comprises at least one image stack corresponding to at least one staining cycle, and wherein the at least one image stack comprises one or more single-channel images associated with one or more biological stains; extracting, from the image dataset, a plurality of regions of interest, at least one region of interest per tissue sample, wherein a region of interest within the plurality of regions of interest comprises at least one region of interest image stack corresponding to the at least one image stack; determining, per region of interest within the plurality of regions of interest, based on at least one single-channel image within the at least one region of interest image stack corresponding to the region of interest, a plurality of cells within the region of interest; determining, per region of interest within the plurality of regions of interest, per cell within the plurality of cells within the region of interest, values for a plurality of morphometric parameters, values for a plurality of neighborhood parameters, and values for a plurality of intensity and texture parameters; determining, based on the plurality of morphometric parameter values, the plurality of neighborhood parameter values, and the plurality of the intensity and texture parameter values, a plurality of single-cell clusters; quantifying, per region of interest within the plurality of regions of interest, per single-cell cluster within the plurality of single-cell clusters, a proportion of a subset of cells within the plurality of cells within the region of interest associated with the single-cell cluster; determining, per tissue sample within the plurality of tissue samples, based on the quantified proportions, a phenotypic signature associated with the tissue sample; classifying, based on the phenotypic signatures, the plurality of tissue samples into at least two similarity groups; analyzing, based on the outcome information corresponding to theplurality of tissue samples, deviation in patient outcome between the at least two similarity groups; and providing output indicating the phenotypic signatures, the at least two similarity groups, and the deviation in patient outcome.
2. A computer-implemented method according to claim 1, wherein the providing the output comprises visualizing the output on a display.
3. A computer-implemented method according to claim 1 or 2, further comprising: identifying, based on the phenotypic signatures, the at least two similarity groups, and the deviation in patient outcome, one or more clinically significant parameters from within at least the plurality of morphometric parameters, the plurality of neighborhood parameters, and the plurality of intensity and texture parameters; and providing further output indicating the one or more clinically significant parameters.
4. A computer-implemented method according to any of the preceding claims, wherein the image dataset comprises two or more image stacks corresponding to two or more staining cycles, and wherein a region of interest within the plurality of regions of interest comprises two or more region of interest image stacks corresponding to the two or more images stacks, further comprising: aligning, per region of interest within the plurality of regions of interest, the two or more region of interest image stacks comprised in the region of interest, to obtain one aligned region of interest image stack.
5. A computer-implemented method according to claim 4, wherein a shared staining agent has been used in the at least two staining cycles, and wherein the aligning the two or more region of interest image stacks is performed based on the shared staining agent.
6. A computer-implemented method according to any of the preceding claims, further comprising:determining, per region of interest within the plurality of regions of interest, whether the region of interest comprises a presence of adequate data; and removing, in response to the region of interest not comprising the presence of the adequate data, the region of interest from within the plurality of regions of interest.
7. A computer-implemented method according to any of the preceding claims, further comprising: assigning, per cell within the plurality of cells within the plurality of regions of interest, based on the plurality of single-cell clusters, an identity within at least two identities with the cell; and quantifying, per identity within the at least two identities, per region of interest within the plurality of regions of interest, per single-cell cluster within the plurality of single-cell clusters, the proportion of a subset of cells assigned with the identity within the plurality of cells within the region of interest associated with the single-cell cluster.
8. A computer-implemented method according to claim 7, further comprising: calculating, per cell within the plurality of cells within the plurality of regions of interest, distance to a nearest cell, wherein the cell has a first identity from within the at least two identities and the nearest cell has a second identity from within the at least two identities; and modifying, based on the calculated distances, the determined plurality of single-cell clusters.
9. A computer-implemented method according to claim 7 or 8, wherein the at least two identities comprise at least a tumor identity and a stroma identity.
10. A computer-implemented method according to any of the preceding claims, further comprising: determining, based on similarities in the phenotypic signatures, the at least two similarity groups using a hierarchical clustering method.
11. A computer-implemented method according to any of the preceding claims, wherein the determining, per region of interest within the plurality of regions of interest, the plurality of cells within the region of interest comprises: segmenting a plurality of subcellular structures comprising at least one of: a nucleus, a cytoplasm, a nuclear envelope, and an endoplasmatic reticulum.
12. A computer-implemented method according to any of the preceding claims, wherein the outcome information comprises information on at least one of: survival, remission, and recurrence.
13. A computer-implemented method according to any of the preceding claims, wherein the image dataset has been obtained by multiplexed fluorescence immunohistochemical imaging.
14. An apparatus comprising: at least one processor; and at least one memory including computer program code, the at least one memory and computer program code being configured to, with the at least one processor, cause the apparatus at least to perform: receiving an image dataset corresponding to a plurality of tissue samples and outcome information corresponding to the plurality of tissue samples, wherein the image dataset comprises at least one image stack corresponding to at least one staining cycle, and wherein the at least one image stack comprises one or more single-channel images associated with one or more biological stains; extracting, from the image dataset, a plurality of regions of interest, at least one region of interest per tissue sample, wherein a region of interest within the plurality of regions of interest comprises at least one region of interest image stack corresponding to the at least one image stack; determining, per region of interest within the plurality of regions of interest, based on at least one single-channel image within the at least one region of interest image stack corresponding to the region of interest, a plurality of cells within the region of interest; determining, per region of interest within the plurality of regions ofinterest, per cell within the plurality of cells within the region of interest, values for a plurality of morphometric parameters, values for a plurality of neighborhood parameters, and values for a plurality of intensity and texture parameters; determining, based on the plurality of morphometric parameter values, the plurality of neighborhood parameter values, and the plurality of the intensity and texture parameter values, a plurality of single-cell clusters; quantifying, per region of interest within the plurality of regions of interest, per single-cell cluster within the plurality of single-cell clusters, a proportion of a subset of cells within the plurality of cells within the region of interest associated with the single-cell cluster; determining, per tissue sample within the plurality of tissue samples, based on the quantified proportions, a phenotypic signature associated with the tissue sample; classifying, based on the phenotypic signatures, the plurality of tissue samples into at least two similarity groups; analyzing, based on the outcome information corresponding to the plurality of tissue samples, deviation in patient outcome between the at least two similarity groups; and providing output indicating the phenotypic signatures, the at least two similarity groups, and the deviation in patient outcome.
15. A computer program product comprising a computer-readable program code configured to, when read and executed by a computer system, cause the computer system at least to perform: receiving an image dataset corresponding to a plurality of tissue samples and outcome information corresponding to the plurality of tissue samples, wherein the image dataset comprises at least one image stack corresponding to at least one staining cycle, and wherein the at least one image stack comprises one or more single-channel images associated with one or more biological stains; extracting, from the image dataset, a plurality of regions of interest, at least one region of interest per tissue sample, wherein a region of interest within the plurality of regions of interest comprises at least one region of interest image stack corresponding to the at least one image stack; determining, per region of interest within the plurality of regions ofinterest, based on at least one single-channel image within the at least one region of interest image stack corresponding to the region of interest, a plurality of cells within the region of interest; determining, per region of interest within the plurality of regions of interest, per cell within the plurality of cells within the region of interest, values for a plurality of morphometric parameters, values for a plurality of neighborhood parameters, and values for a plurality of intensity and texture parameters; determining, based on the plurality of morphometric parameter values, the plurality of neighborhood parameter values, and the plurality of the intensity and texture parameter values, a plurality of single-cell clusters; quantifying, per region of interest within the plurality of regions of interest, per single-cell cluster within the plurality of single-cell clusters, a proportion of a subset of cells within the plurality of cells within the region of interest associated with the single-cell cluster; determining, per tissue sample within the plurality of tissue samples, based on the quantified proportions, a phenotypic signature associated with the tissue sample; classifying, based on the phenotypic signatures, the plurality of tissue samples into at least two similarity groups; analyzing, based on the outcome information corresponding to the plurality of tissue samples, deviation in patient outcome between the at least two similarity groups; and providing output indicating the phenotypic signatures, the at least two similarity groups, and the deviation in patient outcome.