A system and method of categorizing protein expression in biological cells

The system addresses the inefficiencies in analyzing whole-slide images by employing a stain separation algorithm and machine-learning models to enhance and categorize protein expression in digital pathology, improving the visibility and accuracy of faint protein markers.

WO2026028203A1PCT designated stage Publication Date: 2026-02-05NUCLEAI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IL2025/050655
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-30
Filing Date
2025-07-30
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional methods for analyzing whole-slide images in digital pathology lack efficiency and accuracy in extracting meaningful information, particularly in discerning faintly stained assays, which are crucial for identifying subtle cellular markers and pathological features, often obscured by background noise or variations in staining intensity.

Method used

A system and method that utilizes a stain separation algorithm to generate enhanced color images from multi-channel Whole Slide Images (WSIs) of Immunohistochemistry assays, applying machine-learning-based classification models to categorize protein expression by separating and enhancing faint protein markers, and transforming images into a second color space for improved visibility and analysis.

Benefits of technology

Enhances the visibility and accuracy of protein expression analysis in biological cells, allowing for precise classification and interpretation of faint stains, thereby optimizing digital pathology workflows.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050655_05022026_PF_FP_ABST
    Figure IL2025050655_05022026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method of categorizing protein expression in biological cells may include receiving a Whole Slide Image (WSI) of an Immunohistochemistry (IHC) assay of cells, stained with two or more color-specific markers, of two or more respective proteins, and applying a stain separation algorithm on the WSI image, to obtain two or more single¬ channel marker images. Based on the two or more marker images, embodiments may generate an enhanced color image, representing expression of the two or more proteins in the assay, and prompt a user to assign at least one label to at least one respective location in the enhanced color image. Embodiments may then use the at least one label as supervisory information, to train a machine-learning (ML)-based classification model, to classify at least one cell -representative location according to categories of expression of the at least two proteins.
Need to check novelty before this filing date? Find Prior Art

Description

A SYSTEM AND METHOD OF CATEGORIZING PROTEIN EXPRESSION IN BIOLOGICAL CELLSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority of U.S. Patent Application No. 63 / 676,935, filed July 30, 2024, titled “A SYSTEM AND METHOD OF CATEGORIZING PROTEIN EXPRESSION IN BIOLOGICAL CELLS”, which is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION

[0002] The present invention relates generally to automated digital pathology. More specifically, the present invention relates to categorizing protein expression in biological cells.BACKGROUND OF THE INVENTION

[0003] In the era of digital pathology, the analysis of whole-slide images (WSIs) presents a significant challenge due to the complexity and scale of these datasets. Conventional methods often lack efficiency and accuracy in extracting meaningful information from WSIs, hindering diagnostic and research efforts. There is a critical need for innovative solutions that enhance the analysis of WSIs, empowering pathologists and researchers with tools to efficiently navigate, annotate, and extract insights from these vast image datasets.SUMMARY OF THE INVENTION

[0004] Pathologists often encounter challenges when observing faintly stained assays within W Sis. These assays, crucial for identifying subtle cellular markers and pathological features, can be obscured by background noise or variations in staining intensity. The intricacies of discerning faint signals amidst complex tissue structures present a formidable task, demanding meticulous examination and analysis. Current methods for enhancing visibility, such as manual adjustments or image processing algorithms, may be time-consuming or lack precision. Consequently, the accurate interpretation of faint stains remains a bottleneck in pathology workflows, underscoring the pressing need for innovative solutions to optimize detection and analysis processes in digital pathology. The present invention addresses these challenges by introducing novel techniques for enhanced analysis and interpretation of WSIs, and improving digital pathology workflows.

[0005] Embodiments of the invention may include a method of categorizing protein expression in biological cells by at least one processor.

[0006] According to some embodiments, the at least one processor may receive at least one first, multi-channel, Whole Slide Image (WSI) of an Immunohistochemistry (IHC) assay of cells. The assay may be stained with two or more color-specific markers, of two or more respective proteins.

[0007] The at least one processor may apply a stain separation algorithm on the WSI image as elaborated herein, to obtain two or more single-channel marker images. Each singlechannel marker image may represent concentration of a respective marker in the assay.

[0008] Based on the two or more marker images, the at least one processor may generate an enhanced color image, representing expression of the two or more proteins in the assay.

[0009] The at least one processor may subsequently display the enhanced color image to a user via a User Interface (UI), and prompt the user to assign at least one label to at least one respective location in the enhanced color image. The assigned label may indicate expression of the at least two proteins in a cell of the assay.

[0010] The at least one processor may use the at least one label as supervisory information, to train a machine-learning (ML)-based classification model, so as to (i) identify at least one cell-representative location in the at least one first WSI, and (ii) classify the at least one cellrepresentative location according to categories of expression of the at least two proteins.

[0011] The categories of expression may include, for example (i) lack of expression of the two or more proteins at the cell-representative location, (ii) expression of a single protein of the two or more proteins at the cell-representative location, and (iii) co-expression of the at least two proteins at the cell-representative location.

[0012] The at least one processor may (e.g., during an inference stage) receive a target WSI image, depicting cells stained with the color-specific markers of the two or more proteins. The at least one processor may infer, or apply the trained classification model on the target WSI image, to classify at least one cell-representative location in the target WSI image according to the categories of expression.

[0013] According to some embodiments, the channels of the at least one first WSI may span a first color space (e.g., an RGB color space). The at least one processor may apply the stain separation algorithm by calculating a first marker color vector, representing a first marker of the two or more markers in the first color space; calculating a second marker color vector, representing a second marker of the two or more markers in the first color space; transforming, and normalizing the first marker color vector and second marker color vectorin a second color space; calculating a third marker color vector, such that the first, second and third marker color vectors span the second color space; and calculating the two or more marker images based on the first and second marker color vectors, in the second color space.

[0014] According to some embodiments, for each of a plurality of pixels of the at least one first WSI, the at least one processor may calculate a first coefficient value, representing a weight of the first marker color vector in that pixel; and generate a first marker image of the two or more marker images, corresponding to the first marker, based on the first coefficient values of the plurality of pixels.

[0015] Additionally, or alternatively, for each of a plurality of pixels of the at least one first WSI, the at least one processor may calculate a second coefficient value, representing a weight of the second marker color vector in that pixel; and generate a second marker image of the two or more marker images, corresponding to the second marker, based on the second coefficient values of the plurality of pixels.

[0016] For each of the two or more marker images, the at least one processor may: obtain a distribution data element, representing statistical distribution of the corresponding marker in the marker image; based on the distribution data element, compute at least one histogram parameter value; and based on the at least one histogram parameter value, apply a non-linear histogram adaptation algorithm on the marker image, to produce an enhanced, singlechannel image.

[0017] The enhanced single-channel image may include more representations of the corresponding marker, than those depicted in the marker image.

[0018] Additionally, or alternatively, the enhanced single-channel image may have a dynamic range that is wider than that of the marker image.

[0019] According to some embodiments, the at least one processor may compute the at least one histogram parameter value by: receiving one or more reference WSI images of H4C assays of cells, wherein said assays may be stained with a specific, single marker of the two or more color-specific markers. The at least one processor may subsequently apply the stain separation algorithm on the one or more reference WSI images, to obtain respective singlechannel, reference marker images, representing concentration of a specific protein corresponding to the specific marker. The at least one processor may proceed to calculate one or more distribution data elements, representing statistical distribution of the specificmarker in the one or more reference marker images, and compute the at least one histogram parameter value as a function of a percentile of said statistical distribution in the one or more reference marker images.

[0020] Additionally, or alternatively, the at least one processor may compute the at least one histogram parameter value by applying a pretrained ML-based optimization model on the distribution data element, to predict the at least one histogram parameter value.

[0021] According to some embodiments, the at least one processor may receive a training WSI image depicting an assay stained with at least one marker of the two or more colorspecific markers. The training WSI image may be associated with one or more annotations, and the annotations may indicate expression of proteins corresponding to the at least one marker in the assay. The at least one processor may apply the stain separation algorithm on the training WSI image, to obtain a training marker image representing concentration of a respective marker in said training WSI image, and calculate a training distribution data element, representing statistical distribution of the marker in the training marker image. The at least one processor may then use the annotated WSI image as supervisory information, to train the ML-based optimization model, so as to predict the at least one histogram parameter value based on the training distribution data element.

[0022] According to some embodiments, the at least one processor may recombine the enhanced, single-channel images pertaining to each of the two or more marker images, to generate the enhanced color image.

[0023] Embodiments of the invention may include a system for categorizing protein expression in biological cells. Embodiments of the system may include a non-transitory memory device, wherein modules of instruction code may be stored, and at least one processor associated with the memory device, and configured to execute the modules of instruction code. Upon execution of said modules of instruction code, the at least one processor may be configured to: receive at least one first, RGB WSI of an IHC assay of cells, wherein said assay may be stained with two or more color-specific markers of two or more respective proteins; apply a stain separation algorithm on the WSI image, to obtain two or more single-channel marker images, representing concentration of two or more respective proteins in said assay; based on the two or more marker images, generate an enhanced color image, representing expression of the two or more proteins in said assay; display theenhanced color image to a user via a UI; and prompt the user to assign at least one label to at least one respective location in the enhanced color image, wherein said label indicates expression of the at least two proteins in a cell of the assay.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The subject matter regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. The invention, however, both as to organization and method of operation, together with objects, features, and advantages thereof, may best be understood by reference to the following detailed description when read with the accompanying drawings in which:

[0025] Fig. 1 is a block diagram, depicting a computing device which may be included in a system for categorizing protein expression in biological cells according to some embodiments;

[0026] Fig. 2 is a block diagram, depicting an overview of a system for categorizing protein expression in biological cells, according to some embodiments;

[0027] Fig. 3 A shows a segment of a whole-slide image (WSI), depicting cells in a slice of a biological tissue;

[0028] Fig. 3B shows another version of the segment of Fig. 3 A, having been enhanced by embodiments of the invention;

[0029] Fig. 4 is a block diagram, depicting components of an image enhancement module that may be included in the system of Fig. 2, according to some embodiments of the invention;

[0030] Fig. 5A shows a segment of a WSI, depicting cells in a slice of a biological tissue, in a Red, Green and Blue (RGB) color space;

[0031] Figs. 5B and 5C are single-channel, deconvoluted, or unmixed images of the same WSI segment depicted in Fig. 5 A, each corresponding to a specific stain color vector, in a second color space, according to some embodiments of the invention;

[0032] Fig. 6 is a block diagram, depicting another version of the image enhancement module, that may be included in the system of Fig. 2, according to some embodiments of the invention;

[0033] Fig. 7 is a block diagram, depicting functionality of a deconvolution module or algorithm, that may be included in the system of Fig. 2, according to some embodiments of the invention;

[0034] Fig. 8 is a flow diagram, depicting a method of categorizing protein expression in biological cells, according to some embodiments.

[0035] It will be appreciated that for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity. Further, where considered appropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements.DETAILED DESCRIPTION OF THE PRESENT INVENTION

[0036] One skilled in the art will realize the invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The foregoing embodiments are therefore to be considered in all respects illustrative rather than limiting of the invention described herein. Scope of the invention is thus indicated by the appended claims, rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein.

[0037] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be understood by those skilled in the art that the present invention may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the present invention. Some features or elements described with respect to one embodiment may be combined with features or elements described with respect to other embodiments. For the sake of clarity, discussion of same or similar features or elements may not be repeated.

[0038] Although embodiments of the invention are not limited in this regard, discussions utilizing terms such as, for example, “processing,” “computing,” “calculating,” “determining,” “establishing”, “analyzing”, “checking”, or the like, may refer to operation(s) and / or process(es) of a computer, a computing platform, a computing system, or other electronic computing device, that manipulates and / or transforms data represented as physical (e.g., electronic) quantities within the computer’s registers and / or memories intoother data similarly represented as physical quantities within the computer’s registers and / or memories or other information non-transitory storage medium that may store instructions to perform operations and / or processes.

[0039] Although embodiments of the invention are not limited in this regard, the terms “plurality” and “a plurality” as used herein may include, for example, “multiple” or “two or more”. The terms “plurality” or “a plurality” may be used throughout the specification to describe two or more components, devices, elements, units, parameters, or the like. The term “set” when used herein may include one or more items.

[0040] Unless explicitly stated, the method embodiments described herein are not constrained to a particular order or sequence. Additionally, some of the described method embodiments or elements thereof can occur or be performed simultaneously, at the same point in time, or concurrently.

[0041] Reference is now made to Fig. 1, which is a block diagram depicting a computing device, which may be included within an embodiment of a system for categorizing protein expression in biological cells, according to some embodiments.

[0042] Computing device 1 may include a processor or controller 2 that may be, for example, a central processing unit (CPU) processor, a chip or any suitable computing or computational device, an operating system 3, a memory 4, executable code 5, a storage system 6, input devices 7 and output devices 8. Processor 2 (or one or more controllers or processors, possibly across multiple units or devices) may be configured to carry out methods described herein, and / or to execute or act as the various modules, units, etc. More than one computing device 1 may be included in, and one or more computing devices 1 may act as the components of, a system according to embodiments of the invention.

[0043] Operating system 3 may be or may include any code segment (e.g., one similar to executable code 5 described herein) designed and / or configured to perform tasks involving coordination, scheduling, arbitration, supervising, controlling or otherwise managing operation of computing device 1, for example, scheduling execution of software programs or tasks or enabling software programs or other modules or units to communicate. Operating system 3 may be a commercial operating system. It will be noted that an operating system 3 may be an optional component, e.g., in some embodiments, a system may include a computing device that does not require or include an operating system 3.

[0044] Memory 4 may be or may include, for example, a Random-Access Memory (RAM), a read only memory (ROM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a double data rate (DDR) memory chip, a Flash memory, a volatile memory, a nonvolatile memory, a cache memory, a buffer, a short term memory unit, a long term memory unit, or other suitable memory units or storage units. Memory 4 may be or may include a plurality of possibly different memory units. Memory 4 may be a computer or processor non-transitory readable medium, or a computer non-transitory storage medium, e.g., a RAM. In one embodiment, a non-transitory storage medium such as memory 4, a hard disk drive, another storage device, etc. may store instructions or code which when executed by a processor may cause the processor to carry out methods as described herein.

[0045] Executable code 5 may be any executable code, e.g., an application, a program, a process, task, or script. Executable code 5 may be executed by processor or controller 2 possibly under control of operating system 3. For example, executable code 5 may be an application that may categorize protein expression in biological cells as further described herein. Although, for the sake of clarity, a single item of executable code 5 is shown in Fig. 1, a system according to some embodiments of the invention may include a plurality of executable code segments similar to executable code 5 that may be loaded into memory 4 and cause processor 2 to carry out methods described herein.

[0046] Storage system 6 may be or may include, for example, a flash memory as known in the art, a memory that is internal to, or embedded in, a micro controller or chip as known in the art, a hard disk drive, a CD-Recordable (CD-R) drive, a Blu-ray disk (BD), a universal serial bus (USB) device or other suitable removable and / or fixed storage unit. Data pertaining to biological cells may be stored in storage system 6 and may be loaded from storage system 6 into memory 4 where it may be processed by processor or controller 2. In some embodiments, some of the components shown in Fig. 1 may be omitted. For example, memory 4 may be a non-volatile memory having the storage capacity of storage system 6. Accordingly, although shown as a separate component, storage system 6 may be embedded or included in memory 4.

[0047] Input devices 7 may be or may include any suitable input devices, components, or systems, e.g., a detachable keyboard or keypad, a mouse and the like. Output devices 8 may include one or more (possibly detachable) displays or monitors, speakers and / or any other suitable output devices. Any applicable input / output (VO) devices may be connected toComputing device 1 as shown by blocks 7 and 8. For example, a wired or wireless network interface card (NIC), a universal serial bus (USB) device or external hard drive may be included in input devices 7 and / or output devices 8. It will be recognized that any suitable number of input devices 7 and output device 8 may be operatively connected to Computing device 1 as shown by blocks 7 and 8.

[0048] A system according to some embodiments of the invention may include components such as, but not limited to, a plurality of central processing units (CPU) or any other suitable multi-purpose or specific processors or controllers (e.g., similar to element 2), a plurality of input units, a plurality of output units, a plurality of memory units, and a plurality of storage units.

[0049] The term neural network (NN) or artificial neural network (ANN), e.g., a neural network implementing a machine learning (ML) or artificial intelligence (Al) function, may be used herein to refer to an information processing paradigm that may include nodes, referred to as neurons, organized into layers, with links between the neurons. The links may transfer signals between neurons and may be associated with weights. A NN may be configured or trained for a specific task, e.g., pattern recognition or classification. Training a NN for the specific task may involve adjusting these weights based on examples. Each neuron of an intermediate or last layer may receive an input signal, e.g., a weighted sum of output signals from other neurons, and may process the input signal using a linear or nonlinear function (e.g., an activation function). The results of the input and intermediate layers may be transferred to other neurons and the results of the output layer may be provided as the output of the NN. Typically, the neurons and links within a NN are represented by mathematical constructs, such as activation functions and matrices of data elements and weights. At least one processor (e.g., processor 2 of Fig. 1) such as one or more CPUs or graphics processing units (GPUs), or a dedicated hardware device may perform the relevant calculations.

[0050] Reference is now made to Fig. 2, which depicts a system 10 for categorizing protein expression in biological cells, according to some embodiments.

[0051] According to some embodiments of the invention, system 10 may be implemented as a software module, a hardware module, or any combination thereof. For example, system 10 may be, or may include a computing device such as element 1 of Fig. 1, and may beadapted to execute one or more modules of executable code (e.g., element 5 of Fig. 1) to categorize protein expression in biological cells, as further described herein.

[0052] As shown in Fig. 2, arrows may represent flow of one or more data elements to, and from system 10 and / or among modules or elements of system 10. Some arrows have been omitted in Fig. 2 for the purpose of clarity.

[0053] According to some embodiments, system 10 may be configured to receive, e.g., from 20 an Immunohistochemistry (IHC) slide scanner 20 at least one multi-channel (e.g., RGB) 20C WSI image 201 of a multiplex IHC assay of cells.

[0054] The WSI image 201 may be referred to as a multi-channel 20C image in a sense that it may include multiple (e.g., three) layer or channels 20C, such as Red, Green and Blue (RGB) channels 20C, spanning a first color space (e.g., an RGB color space).

[0055] The multiplex IHC WSI image 201 may be stained with two or more color-specific markers 20M of two or more respective proteins. The terms “stains” and “markers” may be used herein interchangeably, referring to visual (e.g., color-marked) indication of expression of one or more specific biological entities (e.g., proteins) in sections of WSI images.

[0056] System 10 may include an image enhancement module 100, configured to generate an enhanced color version of WSI image 201, herein referred to as “enhanced image 100E”. As elaborated herein, enhanced image 100E may represent an enhanced expression of the two or more proteins in the assay of WSI image 201.

[0057] Reference is also made to Figs. 3A and 3B. Fig. 3A shows a segment of a wholeslide image, depicting cells in a slice of a biological tissue. Fig. 3B shows another version of the segment of Fig. 3 A, having been enhanced by embodiments of the invention.

[0058] Fig. 3 A corresponds to WSI image element 201 of Fig. 2, depicting a WSI (or portion thereof) of an Immunohi stochemi stry (IHC) assay of cells. The assay was stained with two or more color-specific markers, corresponding to two or more respective proteins. The first marker may be observed as a purple hue in Fig. 3 A (201). The second marker has a yellow hue, and is hardly recognizable in Fig. 3 A (201).

[0059] Fig. 3B corresponds to enhanced image 100E of Fig. 2, provided by image enhancement module 100. Fig. 3B (100E) depicts the same WSI section as in Fig. 3A (201). By comparing Fig. 3A (201) with Fig. 3B (100E), it may be appreciated that both colorspecific markers, e.g., the purple and yellow hues, are clearly recognizable in Fig. 3B (100E).Image element 100E is therefore referred to as “enhanced” in a sense that it provides an improved expression of all color-specific markers of the original WSI image 201 version.

[0060] Additionally, Fig. 3B (100E) depicts finer, sharper details in relation to original WSI image 201 (Fig. 3A). This is due to a non-linear histogram algorithm performed by system 10, resulting in improved contrast (wider dynamic range) in image 100E in relation to WSI image 201, while maintaining representation of regions in which the purple and yellow markers are shown.

[0061] Reference is now made to Fig. 4, which is a block diagram, depicting components of image enhancement module 100 (or enhancement 100, for short) that may be included in the system 10 of Fig. 2, according to some embodiments of the invention.

[0062] As shown in Fig. 4, enhancement 100 may include a deconvolution module 110, configured to implement, or apply a stain separation algorithm on WSI image 201, as elaborated herein. Deconvolution module 110 may thus obtain two or more single-channel (also referred to as gray -level) marker images 110M, where each single-channel marker image 110M respectively represents a concentration of a specific marker (associated with a respective protein) of the two or more markers in the assay of WSI 201.

[0063] Reference is now made to Figs. 5 A, 5B and 5C. Fig. 5 A shows a segment of a WSI image (e.g., image 201 of Fig. 2), depicting an H&E (Hematoxylin and Eosin) dyed slice of cells of a biological tissue, in a Red, Green and Blue (RGB) color space.

[0064] Figs. 5B and 5C are single-channel (gray -level), deconvoluted, or unmixed images of the same WSI segment (201) depicted in Fig. 5 A. Each of Figs. 5B and 5C corresponds to a specific single-marker image 110M of Fig. 4, represented by a specific stain color vector, in a second, “unmixed” color space.

[0065] As elaborated herein, each of Figs. 5B and 5C (and thus single-marker images 110M1, 110M2) may be referred to as “deconvoluted”, “unmixed” or “separated” in a sense that they may each represent a specific stain color (and thereby concentration of a specific protein) in an original image (e.g., WSI image 201) such as depicted in Fig. 5 A. In the example of Figs. 5A-5C, Fig. 5B is a single-marker image 110M (e.g., 110M1) depicting concentration of Hematoxylin in the dyed WSI of Fig. 5A. In a complementary manner, Fig. 5C is a single-marker image 110M (e.g., 110M2) depicting concentration of Eosin in the dyed WSI of Fig. 5 A.

[0066] As explained herein, the channels 20C of WSI image 201 may span a first multichannel color space (e.g., an RGB color space), allowing color representation of WSI image 201 in that color space.

[0067] According to some embodiments, deconvolution module 110 may calculate a first marker color vector 110SV (e.g., 110SV1), representing a first marker of the two or more markers in the color space of WSI 201, and calculate a second marker color vector 110SV (e.g., 110SV2), representing a second marker of the two or more markers in the color space of WSI 201.

[0068] For example, marker color vector 110SV1 may initially include R, G and B values, representing a hue of a first color-specific marker of the at least two color-specific markers of WSI image 201. Pertaining to the example of Fig. 5B, marker color vector 110S VI may represent hue of the first marker, e.g., Hematoxylin.

[0069] In a complementary manner, marker color vector 110SV2 may initially include R, G and B values, representing a hue of a second color-specific marker of the at least two color-specific markers of WSI image 201. Pertaining to the example of Fig. 5C, marker color vector 110SV2 may represent hue of the second marker, e.g., Eosin.

[0070] Marker color vectors 110SV1 and 110SV2 may be transformed and normalized as a unit-vectors (e.g., having a unit length) in a second color space, also referred to herein as the “unmixed” color space. In some embodiments, marker color vectors 110SV1 and 110SV2 may be orthogonal (or orthonormal) unit vectors in the unmixed color space. In other words, marker color vectors 110SV1 and 110SV2 may be represented by Eqs. 1A and IB respectivelyEq, 1A:110SV1 = {xl, yl, zl}, andEq, IB:110SV2 = {x2, y2, z2}.

[0071] Deconvolution module 110 may subsequently calculate a third marker color vector 110SV (e.g., 110SV3) also referred to herein as a “residual” color vector 110SV3, such that the first (110SV1), second (110SV2) and third (110SV3) marker color vectors 110SV span the unmixed color space. In some embodiments, the third (110SV3) marker color vector may be orthogonal to the first marker color vector (110SV1) and second marker color vector (110SV2).

[0072] It may be appreciated by a person skilled in the art that the representation of stain concentration is not linear in the first, RGB color space. In other words, twice the amount of stain concentration will not necessarily result in color intensity that is twice as strong in the RGB representation. Therefore, it may be incorrect to represent color intensity in the RGB space as in Eq. 2 below:Eq, 2A:RGB = V * S, where RGB is the color intensity (in each of the RGB channels), S is the stain concentration vector 3x1, and V is the stain color matrix 3x3.

[0073] According to some embodiments, the second, unmixed color space may provide a substantially linear relation between stain concentration and color intensity (e.g., twice the concentration of stain producing a stain color that is twice as intense). In other words, the unmixed color space may compensate for (i) an Optical Density (OD) of the studied biological sample and (ii) a thickness of the studied biological sample, to represent stain concentration as color intensity in a substantially linear relation.

[0074] Therefore, the transformation of color vectors from a first (e.g., RGB) color space to a second, “unmixed” color space may be presented as an operator OD( ), according to Eq. 2B, below:Eq, 2B:OD(RGB) = V * OD(S).

[0075] Embodiments of the invention may now represent WSI image 201 in the unmixed color space. In other words, a color (WSI COLOR) of each pixel in WSI 201 may be represented by a weighted linear combination of 110SV1, 110SV2 and 110SV3, as in equation Eq. 3 below:Eq, 3:OD(WSI_COLOR) = wl*110SVl + w2*110SV2 + w3*110SV3, where wl, w2 and w3 represent intensity weights of each of the marker color vectors 110S V in the unmixed color space.

[0076] According to some embodiments, operator OD( ) may include a logarithmic operator. The logarithmic operator may transfer a non-linear (e.g., exponential) relation of physical properties (e.g., thickness, density) of the observed biological sample on absorption of transmitted light, into a linear (e.g., product) relation. In other words, the logarithmic operator of OD( ) may produce a substantially linear relationship between color intensity inthe unmixed color space and the concentration of stain (and corresponding level of protein expression) in WSI image 201.

[0077] As shown in Eq. 3, for one or more (e.g., each) of the plurality of pixels of at least one WSI image 201, deconvolution module 110 may calculate a first coefficient value wl, representing a weight of the first marker color vector in that pixel, and calculate a second coefficient value w2, representing a weight of the second marker color vector in that pixel.

[0078] Deconvolution module 110 may then generate, or calculate the two or more marker expression images 110M based on the first (110SV1) and second (110SV2) marker color vectors 110SV, and the coefficient values (wl, w2) of the plurality of pixels.

[0079] For example, a first marker expression image 110M (e.g., 110M1) may correspond to the first marker color vector 110SV1 in a sense that pixels of image 110M1 may consist of the coefficient weight values wl of the first marker color vector 110SV1. In other words, first marker image 110M1 may be a grey level image, whose intensity is determined by the weight value wl in the unmixed color space. It may be appreciated that wl may represent a concentration of the first marker (e.g., Hematoxylin) in a respective location in WSI image 201.

[0080] In a complementary manner, a second marker image 110M (e.g., 110M2) may correspond to the second marker color vector 110SV2 in a sense that pixels of image 110M2 consist of coefficient weight values w2 of the second marker color vector 110SV2. In other words, second marker image 110M2 may be a grey level image, whose intensity is determined by the weight value w2 in the unmixed color space. It may be appreciated that w2 may represent a concentration of the first marker (e.g., Eosin) in a respective location in WSI image 201.

[0081] According to some embodiments, the third, residual (110SV3) marker color vector 110SV may be used to evaluate success in unmixing image 201 into marker images 110M1 and 110M2.

[0082] It may be appreciated that color vectors 110SV1 and 110SV2 should effectively span the color space of WSI 201, that was originally stained by corresponding color-specific markers 20M of the two or more respective proteins. Therefore, vector 110SV3 may primarily represent coloring that was obtained from artifacts and noise. A strong signal (weight w3) of the residual channel 110SV3 may thus indicate a poor result of unmixing WSI 201 by deconvolution module 110 into vectors 110S VI and 110SV2.

[0083] According to some embodiments, image enhancement module may calculate a deconvolution confidence score 110CF, based on signal w3 of the residual channel 110SV3, e.g., where a poor confidence score 110CF corresponds to a strong signal w3. System 20 may subsequently choose to ignore one or more cell -representative areas in WSI 201 (now in images 110M), corresponding with a poor (e.g., below a predetermined threshold) confidence score 11 OCF for the purpose of classifying such regions according to their protein expression.

[0084] Image enhancement module 100 may include a reference distribution data generation module 120 (or “distribution module 120”, for short), configured to receive (e.g., via input 7 of Fig. 1), obtain or calculate a reference distribution data element 120RD (or distribution data 120RD for short).

[0085] Distribution data 120RD may be implemented as a data structure (e.g., a table), representing statistical characteristics and / or distribution of concentration of a marker in a corresponding marker image (e.g., 110M1, 110M2). Additionally, or alternatively, distribution data 120RD may represent statistical characteristics pertaining to expression of proteins associated with specific markers, as depicted in WSI image(s) 201.

[0086] For example, distribution data element 120RD may be, or may include a hypothesis, or definition of a statistical model (e.g., Gaussian distribution model, long-tailed distribution model, Poisson distribution model, and the like), that is hypothesized to represent distribution of pixel brightness in marker images 110M1, 110M2 (corresponding to marker concentrations) in one or more (e.g., a plurality of) WSI images 201. Additionally, or alternatively, distribution data element 120RD may include a hypothesis of a statistical model (e.g., Gaussian distribution model) in one or more WSI images 201.

[0087] Additionally, or alternatively, distribution data element 120RD may be, a distribution, or a histogram of pixel brightness in marker images 110M1, 110M2 (corresponding to marker concentrations) in one or more (e.g., a plurality of) WSI images 201.

[0088] As shown elaborated herein, image enhancement module 100 may include a histogram parameters calculation module 130 (or “module 130” for short), and a histogram adjustment module 140 (or “module 140” for short). As explained herein, module 130 may be adapted to compute at least one histogram parameter value 13 OH based on distribution data element 120RD. Module 140 may in turn apply a non-linear histogram adaptationalgorithm on at least one marker image 110M (e.g., 110M1, 110M2), to produce at least one respective enhanced, single-channel image 140E (e.g., 140E1, 140E2).

[0089] Histogram parameter values 130H may, for example include an upper histogram threshold and / or lower histogram threshold, that define a histogram range of pixel brightness (corresponding to coefficient weight values wl, w2) in respective marker images 110M1, 110M2. As elaborated herein, module 130 may calculate an optimal value for histogram parameter values (e.g., histogram thresholds) I30H, thereby allowing module 140 to enhance marker images 110M1, 110M2, and produce respective, enhanced , single-channel images 140E1, 140E2, in an optimal manner.

[0090] The term “optimal” is used in this context in a sense that images 140E1, 140E2 may be enhanced in relation to marker images 110M1, 110M2 according to a combination of one or more (optionally contradicting) criteria.

[0091] For example, enhanced , single-channel image MOE (a) may include at least the same number of representations of the corresponding marker, than those depicted in the respective marker image 110M, and (b) may have a dynamic range that is wider than that of the respective marker image 110M (e.g., thereby depicting more details than in marker image 110M).

[0092] Additionally, or alternatively, enhanced , single-channel image MOE (a) may include more representations of the corresponding marker, than those depicted in the respective marker image 110M, while (b) may have a similar dynamic range to that of the respective marker image 110M.

[0093] Additionally, or alternatively, enhanced , single-channel image MOE (a) may include more representations of the corresponding marker, than those depicted in the respective marker image 110M, and (b) have a wider dynamic range in relation to that of the respective marker image 110M.

[0094] It may be appreciated that such optimization or enhancement of marker images 110M may allow maximal definition of details by a human observer, while preserving or improving classification of elements in WSI image 201, as elaborated herein.

[0095] In another example, one or more (e.g., a cohort of) multiplex IHC WSI image(s) 201 may include reference image(s) 20IR of IHC assays of cells, stained with a specific, single marker of the two or more color-specific markers. As elaborated herein image enhancement module 100 may apply the stain separation algorithm on the one or more reference WSIimage(s) 20IR, to obtain respective single-channel, reference marker image(s) 110M (denoted 11 OMR). Reference marker images 11 OMR may represent concentration of a specific protein corresponding to the specific marker in the assay of reference WSI image(s) 20IR.

[0096] Histogram parameters calculation module 130, may calculate one or more distribution data elements 120RD representing statistical distribution of the specific marker in the one or more reference marker images 20IR. For example, distribution data elements 120RD may be, or may include histograms representing distribution of brightness levels in reference marker images 11 OMR, corresponding to statistical distribution of concentration of the specific marker in the one or more reference marker images.

[0097] Module 130 may compute histogram parameter value 130Hbased on the distribution data elements 120RD of the reference marker images 110MR. For example, histogram parameter value 130H may be a function (e.g., a weighted average) of a predetermined percentile (e.g., 10%) of the statistical distribution of the specific marker in the one or more reference marker images 100MR.

[0098] It may be appreciated that reference marker images 100MR may pertain to a cohort of subjects, thereby improving the robustness and accuracy of histogram parameter values I 30H.

[0099] Additionally, or alternatively, histogram parameters’ calculation module 130 may compute the at least one histogram parameter value (e.g., histogram threshold) by applying a pretrained ML-based optimization model 135 on distribution data element 120RD, ML optimization model 135 may thereby predict the at least one histogram parameter value 135 based on distribution data element 120RD.

[0100] According to some embodiments, ML-based optimization model 135 may be trained by a supervised training scheme, using expert-annotated multiplex IHC WSI images 201 as supervisory data to predict the at least one histogram parameter value 135 based on distribution data element 120RD.

[0101] For example, image enhancement module 100 may (e.g., during a training stage) obtain or receive (e.g., via input 7 of Fig. 1) at least one (e.g., a plurality of) training multiplex IHC WSI images 201 (denoted 20IT) depicting an assay stained with at least one marker of the two or more color-specific markers.

[0102] Image enhancement module 100 may further receive plurality of expert annotations 120 AN associated with training multiplex IHC WSI image(s) 20IT. Expert annotations 120AN may be provided by a human or automated annotator, and may indicate expression of proteins corresponding to specific markers, in specific locations or regions (e.g., within cells, cell regions or sub-cellular organelles) of the associated training WSI image(s) 20IT.

[0103] For example, training multiplex IHC WSI images 20IT may depict cell lines which are known to a human experts as such that (i) do not express proteins corresponding to the two color-unique markers, (ii) express proteins corresponding to a first marker of the two color-unique markers, (iii) express proteins corresponding to a second marker of the two color-unique markers, and (iv) express proteins corresponding to both color-unique markers. Based on this knowledge, the expert annotator may provide annotations 120AN regardless to their ability to distinguish a relevant marker in training image 20IT (e.g., even when an expected marker signal is too weak to be perceived by a human observer).

[0104] Deconvolution module 110 may apply the stain separation, or unmixing algorithm on training WSI image 20IT as elaborated herein, to obtain at least one training marker image 110M (denoted 110MT) representing concentration of a respective, specific marker in training WSI image 20IT.

[0105] Distribution data generator module 120 may, intum, calculate at least one training distribution data element 120RD (denoted 120RDT) based on the at least one training marker image 110M, as elaborated herein. Training distribution data element 120RDT may represent statistical distribution (e.g., a histogram) of pixel brightness values (e.g., corresponding to marker concentration) in the training marker image 110MT.

[0106] According to some embodiments, histogram parameters calculation module 130 may provide training distribution data element 120RDT as input for ML optimization model 135. Histogram parameters calculation module 130 may use the annotations of training multiplex IHC WSI image 20IT as supervisory information, to train the ML-based optimization model 135, so as to predict the at least one histogram parameter value based on the training distribution data element 120RDT.

[0107] Additionally, or alternatively, histogram parameters calculation module 130 may collaborate with histogram adjustment module 140, to train ML model 135 so as to predict an optimal value of histogram parameters 13 OH.

[0108] Module 130 may calculate a value of a loss function 135L, based on enhanced, single-channel image 140E, and optimize the predicted histogram parameters 130H based on the calculated loss function value 135L.

[0109] For example, loss function value 135L may be positively correlative to the dynamic range (e.g., ability to detect details) in enhanced image 140E. In another example, loss function value 135L may be positively correlative to the number of regions where at least one marker appears in enhanced image 140E. In another example, loss function value 135L may implement a penalty on loss of a marker signal, e.g., when existence of a marker is depicted in a specific location in marker image 110M, but is absent in a respective location in enhanced image MOE.

[0110] In another example, module 130 may collaborate with reference data generator 120 to determine histogram parameter value 130H based on distribution data element 120RD, by employing a distribution fitting algorithm.

[0111] Reference is now made to Fig. 6, which is a block diagram, depicting another example of implementation of image enhancement module 100, that may be included in the system 10 of Fig. 2, according to some embodiments of the invention.

[0112] As explained herein, distribution data element 120RD may initially be set, or dedicated to representing a specific type, of statistical model (e.g., Gaussian distribution model, long-tailed distribution model, and the like), that is hypothesized to represent distribution of a specific marker concentration in multiplex IHC WSI images 201.

[0113] As known in the art, the term “distribution fitting” may refer to choosing a probability model for an unknown population, and calibrating that model using a representative sample from that population. Distribution data generation module 120 may apply a distribution fitting algorithm on one or more (e.g., each) marker image 110M as known in the art, to calibrate the hypothesized distribution model. In the example of a Gaussian distribution model, distribution module 120 may utilize a distribution fitting algorithm to determine a mean value, and a standard deviation value of brightness levels in marker images 110M which, as explained above, correspond to concentration of marker colors in multiplex IHC WSI image 201.

[0114] Distribution data generation module 120 may subsequently collaborate with histogram parameters’ calculation module 130 and histogram adjustment module 140 to generate enhanced, single-channel images MOE (e.g., 140E1, 140E2) as explained herein.

[0115] As shown in Figs. 4 and 6, image enhancement module 100 may further include a recombination module, configured to recombine the enhanced, single-channel images 140E1, 140E2 pertaining to each of the two or more marker images 110M1, 110M2, to generate an enhanced color image 110E.

[0116] Enhanced color image 110E may represent expression of the two or more proteins in the assay of multiplex IHC WSI image(s) 201. Color image 110E may be enhanced in relation to original WSI image 201 in a sense that it may depict more regions of markers in the imaged assay where proteins of interest are expressed, as shown in the example of Figs. 3A and 3B.

[0117] Additionally, or alternatively, color image 110E may be enhanced in relation to original W SI image 201 in a sense that it may have a dynamic range that is wider (and thereby depict finer details) than that of original WSI image 201.

[0118] As shown in Fig. 2, system 10 may include, or may be associated with a User Interface (UI) 200 such as input and output devices 7, 8 of Fig. 1.

[0119] According to some embodiments, system 10 may display enhanced color image 110E to a user via UI 200. System 10 may thus facilitate an improved presentation of WSI images 201 for the purpose of evaluating a condition of a subject from which the assay was extracted.

[0120] Additionally, or alternatively, system 10 may prompt the user (e.g., an expert pathologist) to assign at least one label or annotation 20A to at least one respective location in the enhanced color image 110E.

[0121] A value of labels 20A may indicate expression of the at least two proteins in a cell of the assay. For example, labels 20A may include binary values {00, 01, 10, 11}, representing (i) no protein expression, (ii) expression of only a first protein, (iii) expression of only a second protein, and (iv) expression of both proteins.

[0122] A location of labels 20 A in image 100E may indicate a region of the assay to which the label pertains. This region may be, for example, a location in a specific cell, a region within a specific the cell (e.g., cytoplasm, nucleus, organelles, membrane, etc.), and the like.

[0123] As shown in Fig. 1, system 10 may further include an ML-based classification model 300, adapted to classify one or more regions of multiplex IHC WSI image 201 .

[0124] System 10 may (e.g., during a training stage) receive at least one training WSI 20IT, generate therefrom an enhanced image 100E, and obtain labels 20 A as elaborated herein. System 10 may subsequently use the one or more labels as supervisory information, to train ML-based classification model 300 to identify at least one cell-representative location (E.g., cytoplasm, organelles, membrane, etc.) in the at least one test WSI 20IT. Additionally, or alternatively, system 10 may use the one or more labels as supervisory information, to train ML-based classification model 300 to classify at least one cellrepresentative location according to categories of expression of the at least two proteins. The categories of expression for ML model 300 may include (i) lack of expression of the two or more proteins at the cell-representative location, (ii) expression of a single protein of the two or more proteins at the cell -representative location, and (iii) co-expression of the at least two proteins at the cell-representative location. In other words, categories of expression for ML model 300 may be defined as {00, 01, 10, 11}, as elaborated herein.

[0125] System 10 may (e.g., at a subsequent inference stage) receive a target WSI image 201 of interest. Target WSI image 201 may depict cells stained with the color-specific markers of the two or more proteins. System 10 may provide Target WSI image 201 as input to ML classifier 300. System 10 may thereby infer the classification model 300 on the target WSI image 201 based on the training, to produce a classification 310C of at least one cellrepresentative location in the target WSI image 201. As elaborated herein, classification 310C may indicate a category of protein expression (e.g., lack of expression, expression of a single protein of the two or more proteins, and / or co-expression of the at least two proteins) at the cell-representative location.

[0126] Additionally, or alternatively, system 10 may be used as an assistive tool for pathologists, in an effort to classify cell -representative locations in target WSI images 201.

[0127] In such embodiments, a pathologist may use UI 200 to manually assign labels or scores 20L for specific cell-representative locations in target WSI images 201. Such labels may include indication of expression of one or more proteins (e.g., lack of expression, expression of a single protein of the two or more proteins, and / or co-expression of the at least two proteins) in a cell-representative location of a specific WSI image 201. System 10 may search for a discrepancy between the expert’s label 20L and classification 310C. When such a discrepancy is found, UI 200 may present the relevant cell -representative location, highlighting the located discrepancy. Additionally, or alternatively, UI 200 may present therelevant cell -representative location in enhanced image 100E, and prompt the human expert to reevaluate their preliminary assessment, e.g., by reassigning a value to label 20L.

[0128] System 10 may thus fine-tune the labeling of cell-representative locations in W SI images 201, to improve pathologist training and precision. Additionally, or alternatively, system 10 may thus improve accuracy of training datasets, and thus improve training of ML classifier 300.

[0129] Reference is now made to Fig. 7, which is a block diagram depicting functionality of a deconvolution module or algorithm 110, that may be included in the system 10 for categorizing protein expression in biological cells, according to some embodiments of the invention. Deconvolution module 110 may be the same as deconvolution module 110 of Figs. 4 and 6.

[0130] Deconvolution module 110 may include a stain-specific color vector estimator 112 (or “Estimator 112” for short). According to some embodiments, estimator 112 may calculate a multichannel histogram 112MH (e.g., including a red channel histogram, a green channel histogram and a blue channel histogram) of one or more regions of WSI image 201.

[0131] Based on the multichannel histogram 112MH, estimator 112 may identify the two or more color-specific markers or respective two or more proteins.

[0132] Estimator 112 may subsequently calculate statistics (e.g., mean and standard deviation) for each of the two or more color-specific markers depicted in the one or more regions of WSI image 201, also based on the multichannel histogram 112MH. based on the calculated statistics, estimator 112 may calculate an initial version of stain color vectors 110SV (e.g., 110SV1, 110SV2) in the RGB color space.

[0133] Estimator 112 may proceed to transform, and normalize stain color vectors 110SV1 and 110SV2 to be linearly independent, and (together with a residual, third orthonormal vector 110SV3) span a color space that is referred to herein as the unmixed color space.

[0134] Deconvolution module 110 may include a stain concentration estimator module 115. Stain concentration estimator module 115 may evaluate an intensity of each of the colors represented by color vectors 110SV1 and 110SV2 in WSI image 201. It may be appreciated that these intensity values may represent concentration of respective colorspecific stains in WSI image 201. The intensity values of color vectors 110S VI and 110SV2 may be the same as weight values wl and w2 of Eq. 3.

[0135] Additionally, or alternatively, deconvolution module 110 may include an image generator module 117, adapted to produce marker images 110M (e.g., 110M1, 110M2) as 2-dimensional (2D) arrays of intensity values wl and w2 respectively.

[0136] Reference is now made to Fig. 8, which is a flow diagram, depicting a method of categorizing protein expression in biological cells by at least one processor (e.g., processor 2 of Fig. 1), according to some embodiments of the invention.

[0137] As shown in step S1005, processor 2 may receive at least one first, multi-channel, WSI image (e.g., 201 of Fig. 2) of an Immunohistochemistry (IHC) assay (e.g., 20 of Fig. 2) of cells. The assay may be stained with two or more color-specific markers, or stains 20S of two or more respective proteins.

[0138] As shown in step S1010 and explained herein (e.g., in relation to Fig. 4), the at least one processor 2 may apply a stain separation algorithm on the WSI image 201, to obtain two or more single-channel marker images (e.g., 110M of Fig. 4). Each marker image 110M (e.g., 110M1, 110M2) may represent a concentration of a respective marker in the assay.

[0139] As shown in step S 1015, based on the two or more marker images 110M, the at least one processor 2 may generate an enhanced color image (e.g., 100E of Fig. 4). Enhanced color image 100E may represent expression of the two or more proteins in assay 20.

[0140] As shown in step S1020 and S1025, processor 2 may display the enhanced color image 100E to a user via a UI (e.g., UI 200 of Fig. 2), and then prompt a user to assign at least one label 20 A to at least one respective location in enhanced color image 1000E. Label 20A may indicate expression of the at least two proteins in a cell of assay 20.

[0141] As elaborated herein, processor 2 may subsequently use label 20 A as supervisory information, for training a machine learning model (e.g., ML classifier 300 of Fig. 2) to classify expression of proteins in one or more regions of WSI image 201.

[0142] For example, processor 2 may determine expression, lack of expression, and / or co-expression of the two or more proteins in one or more compartments (e.g., cells, nuclei, organelles, membranes, etc.). As shown in Fig. 2, processor 2 may produce an indication 310C of compartment classification (e.g., cell types, cell states, etc.) based on the expression of the two or more proteins.

[0143] In another example, processor 2 may produce a diagnostic indication 320D of a subject associated with WSI image 201, based on the expression of the two or more proteins.Additionally, or alternatively, processor 2 may produce a recommendation 330SR for adjusting, or fixing the staining process of analyzed WSI image 201.

[0144] In other words, based on individual marker cell staining (e.g., wl, w2) a WSI level protein marker co-expression score can be computed. The score can be used as a biomarker to assess the chances of a patient responding to a drug that is targeting one or more of the relevant proteins. Alternatively it can be used as a selection criteria for a clinical trial, or be used in an exploratory effort.

[0145] Additionally, or alternatively, using known cell lines data the performance of the assay can be assessed to determine limits of co-expression detection.

[0146] Embodiments of the invention may include a practical application in the technological field of assistive diagnostics.

[0147] As explained herein, embodiments of the invention may provide an improvement in identifying and labeling cell-representative locations in a WSI image according to specific levels of protein expression. Embodiments of the invention may subsequently use the improved labeling to train ML-based models to automatically classify cell-representative locations according to their protein expression levels.

[0148] Additionally, or alternatively, embodiments of the invention may provide improved visualization of WSI images, thereby allowing expert pathologists to accurately produce diagnostic, and pathology-related decisions based on ubiquitously used (e.g., H&E) staining techniques.

[0149] Unless explicitly stated, the method embodiments described herein are not constrained to a particular order or sequence. Furthermore, all formulas described herein are intended as examples only and other or different formulas may be used. Additionally, some of the described method embodiments or elements thereof may occur or be performed at the same point in time.

[0150] While certain features of the invention have been illustrated and described herein, many modifications, substitutions, changes, and equivalents may occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the invention.

[0151] Various embodiments have been presented. Each of these embodiments may of course include features from other embodiments presented, and embodiments not specifically described may include various features described herein.

Claims

CLAIMS1. A method of categorizing protein expression in biological cells by at least one processor, the method comprising: receiving at least one first, multi-channel, Whole Slide Image (WSI) of an Immunohistochemistry (H4C) assay of cells, wherein said assay is stained with two or more color-specific markers of two or more respective proteins; applying a stain separation algorithm on the WSI image, to obtain two or more singlechannel marker images, each representing concentration of a respective marker in said assay; based on the two or more marker images, generating an enhanced color image, representing expression of the two or more proteins in said assay; displaying the enhanced color image to a user via a User Interface (UI); and prompting the user to assign at least one label to at least one respective location in the enhanced color image, wherein said label indicates expression of the at least two proteins in a cell of the assay.

2. The method of claim 1, further comprising using the at least one label as supervisory information, to train a machine-learning (ML)-based classification model, so as to (i) identify at least one cell-representative location in the at least one first WSI, and (ii) classify the at least one cell -representative location according to categories of expression of the at least two proteins.

3. The method of claim 2, wherein said categories of expression are selected from a list consisting of: (i) lack of expression of the two or more proteins at the cell-representative location, (ii) expression of a single protein of the two or more proteins at the cellrepresentative location, and (iii) co-expression of the at least two proteins at the cellrepresentative location.

4. The method according to any one of claims 2-3, further comprising: receiving a target WSI image, depicting cells stained with the color-specific markers of the two or more proteins; andinferring the trained classification model on the target WSI image, to classify at least one cell-representative location in the target WSI image according to the categories of expression.

5. The method according to any one of claims 1-4, wherein the channels of the at least one first WSI span a first color space, and wherein applying a stain separation algorithm comprises: calculating a first marker color vector, representing a first marker of the two or more markers in the first color space; calculating a second marker color vector, representing a second marker of the two or more markers in the first color space; transforming, and normalizing the first marker color vector and second marker color vector in a second color space; calculating a third marker color vector, such that the first, second and third marker color vectors span the second color space; and calculating the two or more marker images based on the first and second marker color vectors, in the second color space.

6. The method of claim 5, further comprising: for each of a plurality of pixels of the at least one first WSI, calculating a first coefficient value, representing a weight of the first marker color vector in that pixel; and generating a first marker image of the two or more marker images, corresponding to the first marker, based on the first coefficient values of the plurality of pixels.

7. The method of claim 6, further comprising: for each of a plurality of pixels of the at least one first WSI, calculating a second coefficient value, representing a weight of the second marker color vector in that pixel; and generating a second marker image of the two or more marker images, corresponding to the second marker, based on the second coefficient values of the plurality of pixels.

8. The method of claim 7, further comprising for each of the two or more marker images:obtaining a distribution data element, representing statistical distribution of the corresponding marker in the marker image; based on the distribution data element, computing at least one histogram parameter value; and based on the at least one histogram parameter value, applying a non-linear histogram adaptation algorithm on the marker image, to produce an enhanced, single-channel image.

9. The method of claim 8, wherein said enhanced single-channel image comprises more representations of the corresponding marker, than those depicted in the marker image.

10. The method of claim 9, wherein said enhanced single-channel image has a dynamic range that is wider than that of the marker image.

11. The method according to any one of claims 8-10, wherein computing the at least one histogram parameter value comprises: receiving one or more reference WSI images of IHC assays of cells, wherein said assays are stained with a specific, single marker of the two or more color-specific markers; applying the stain separation algorithm on the one or more reference WSI images, to obtain respective single-channel, reference marker images, representing concentration of a specific protein corresponding to the specific marker; calculating one or more distribution data elements, representing statistical distribution of the specific marker in the one or more reference marker images; and computing the at least one histogram parameter value as a function of a percentile of said statistical distribution in the one or more reference marker images.

12. The method according to any one of claims 8-11, wherein computing the at least one histogram parameter value comprises applying a pretrained ML-based optimization model on the distribution data element, to predict the at least one histogram parameter value.

13. The method of claim 12, further comprising: receiving a training WSI image depicting an assay stained with at least one marker of the two or more color-specific markers, wherein said training WSI image is associated withone or more annotations, and wherein said annotations indicate expression of proteins corresponding to the at least one marker in the assay; applying the stain separation algorithm on the training WSI image, to obtain a training marker image representing concentration of a respective marker in said training WSI image; calculating a training distribution data element, representing statistical distribution of the marker in the training marker image; and using the annotated WSI image as supervisory information, to train the ML-based optimization model, so as to predict the at least one histogram parameter value based on the training distribution data element.

14. The method according to any one of claims 8-13, further comprising recombining the enhanced, single-channel images pertaining to each of the two or more marker images, to generate the enhanced color image.

15. A system for categorizing protein expression in biological cells, the system comprising: a non-transitory memory device, wherein modules of instruction code are stored, and at least one processor associated with the memory device, and configured to execute the modules of instruction code, whereupon execution of said modules of instruction code, the at least one processor is configured to: receive at least one first, RGB WSI of an H4C assay of cells, wherein said assay is stained with two or more color-specific markers of two or more respective proteins; apply a stain separation algorithm on the WSI image, to obtain two or more singlechannel marker images, representing concentration of two or more respective proteins in said assay; based on the two or more marker images, generate an enhanced color image, representing expression of the two or more proteins in said assay; display the enhanced color image to a user via a UI; and prompt the user to assign at least one label to at least one respective location in the enhanced color image, wherein said label indicates expression of the at least two proteins in a cell of the assay.

16. The system of claim 15, wherein the at least one processor is configured to use the at least one label as supervisory information, to train a machine-learning (ML)-based classification model, so as to (i) identify at least one cell-representative location in the at least one first WSI, and (ii) classify the at least one cell-representative location according to categories of expression of the at least two proteins.

17. The system of claim 16, wherein said categories of expression are selected from a list consisting of: (i) lack of expression of the two or more proteins at the cell-representative location, (ii) expression of a single protein of the two or more proteins at the cellrepresentative location, and (iii) co-expression of the at least two proteins at the cellrepresentative location.

18. The system according to any one of claims 16-17, wherein the at least one processor is further configured to: receive a target WSI image, depicting cells stained with the color-specific markers of the two or more proteins; and infer the trained classification model on the target WSI image, to classify at least one cell-representative location in the target WSI image according to the categories of expression.

19. The system according to any one of claims 15-18, wherein the channels of the at least one first WSI span a first color space, and wherein the at least one processor is configured to apply a stain separation algorithm by: calculating a first marker color vector, representing a first marker of the two or more markers in the first color space; calculating a second marker color vector, representing a second marker of the two or more markers in the first color space; transforming, and normalizing the first marker color vector and second marker color vector in a second color space; calculating a third marker color vector, such that the first, second and third marker color vectors span the second color space; andcalculating the two or more marker images based on the first and second marker color vectors, in the second color space.

20. The system of claim 19, wherein the at least one processor is further configured to: for each of a plurality of pixels of the at least one first WSI, calculate a first coefficient value, representing a weight of the first marker color vector in that pixel; and generate a first marker image of the two or more marker images, corresponding to the first marker, based on the first coefficient values of the plurality of pixels.

21. The system of claim 20, wherein the at least one processor is further configured to: for each of a plurality of pixels of the at least one first WSI, calculate a second coefficient value, representing a weight of the second marker color vector in that pixel; and generate a second marker image of the two or more marker images, corresponding to the second marker, based on the second coefficient values of the plurality of pixels.

22. The system of claim 21, wherein the at least one processor is further configured to, for each of the two or more marker images: obtain a distribution data element, representing statistical distribution of the corresponding marker in the marker image; based on the distribution data element, compute at least one histogram parameter value; and based on the at least one histogram parameter value, apply a non-linear histogram adaptation algorithm on the marker image, to produce an enhanced, single-channel image.

23. The system of claim 22, wherein said enhanced single-channel image comprises more representations of the corresponding marker, than those depicted in the marker image.

24. The system of claim 23, wherein said enhanced single-channel image has a dynamic range that is wider than that of the marker image.

25. The system according to any one of claims 22-24, wherein the at least one processor is configured to compute the at least one histogram parameter value by:receiving one or more reference WSI images of IHC assays of cells, wherein said assays are stained with a specific, single marker of the two or more color-specific markers; applying the stain separation algorithm on the one or more reference WSI images, to obtain respective single-channel, reference marker images, representing concentration of a specific protein corresponding to the specific marker; calculating one or more distribution data elements, representing statistical distribution of the specific marker in the one or more reference marker images; and computing the at least one histogram parameter value as a function of a percentile of said statistical distribution in the one or more reference marker images.

26. The system according to any one of claims 22-25, wherein the at least one processor is configured to compute the at least one histogram parameter value by applying a pretrained ML-based optimization model on the distribution data element, to predict the at least one histogram parameter value.

27. The system of claim 26, wherein the at least one processor is further configured to: receive a training WSI image depicting an assay stained with at least one marker of the two or more color-specific markers, wherein said training WSI image is associated with one or more annotations, and wherein said annotations indicate expression of proteins corresponding to the at least one marker in the assay; apply the stain separation algorithm on the training WSI image, to obtain a training marker image representing concentration of a respective marker in said training WSI image; calculate a training distribution data element, representing statistical distribution of the marker in the training marker image; and use the annotated WSI image as supervisory information, to train the ML-based optimization model, so as to predict the at least one histogram parameter value based on the training distribution data element.

28. The system according to any one of claims 22-27, wherein the at least one processor is further configured to recombine the enhanced, single-channel images pertaining to each of the two or more marker images, to generate the enhanced color image.