Generating quantitative ground truth for ihc stained slides
Patent Information
- Authority / Receiving Office
- IL · IL
- Patent Type
- Applications
- Current Assignee / Owner
- AGILENT TECHNOLOGIES INC
- Filing Date
- 2024-12-02
- Publication Date
- 2026-07-01
AI Technical Summary
Current methods for generating ground truth for training machine learning models to analyze histological slides are inefficient and prone to human error, particularly in achieving high accuracy for digital pathology applications.
A method is developed to generate quantitative ground truth by obtaining images of biological samples stained with a quantitative approach that converts antibody/antigen complexes into dots, and using these images to create annotations for training machine learning models.
This approach enables the creation of accurate and objective ground truth data, improving the training and performance of machine learning models in digital pathology, particularly for tasks like IHC stain quantification.
Smart Images

Figure 00000056_0000 
Figure 00000057_0000 
Figure 00000057_0001
Abstract
Description
GENERATING QUANTITATIVE GROUND TRUTH FOR IHC STAINED SLIDESCross-Reference to Related Applications
[0001] This application claims priority to U.S. Provisional Application Nos. 63 / 605,424, and 63 / 605,435, which were filed on December 1, 2023, the entire contents of which are hereby incorporated by reference.Technical Field
[0002] The present disclosure relates generally to methods and devices for use in detecting targets in biological tissues aided by imaging, and, in particular, to digital microscopy. The present disclosure further relates to methods and devices for generating ground truth for training machine learning models to analyze images of biological samples, training such models, and using the trained models for the analysis.Background
[0003] Histological specimens are frequently disposed upon glass slides as a thin slice of patient tissue fixed to the surface of each of the glass slides. Using a variety of chemical or biochemical processes, one or several colored compounds may be used to stain the tissue to differentiate cellular constituents, which can be further evaluated utilizing microscopy. Bright- field slide scanners are conventionally used to digitally analyze these slides.
[0004] When analyzing tissue samples on a microscope slide, staining the tissue or certain parts of the tissue with a colored or fluorescent dye can aid the analysis. The ability to visualize or differentially identify microscopic structures is frequently enhanced using histological stains. Hematoxylin and eosin (“H&E”) stains are the most commonly used stains in light microscopy for histological samples.
[0005] In addition to H&E stains, other stains or dyes have been applied to provide more specific staining and provide a more detailed view of tissue morphology. Immunohistochemistry (“IHC”) stains have great specificity, as they use a peroxidase substrate or alkaline phosphatase (“AP”) substrate for IHC staining, providing a uniform staining pattern that appears to the viewer as a homogeneous color with intracellular resolution of cellular structures, e.g., membrane, cytoplasm, and nucleus. Formalin Fixed Paraffin Embedded (“FFPE”) tissue samples, metaphase spreads or histological smears are typically analyzed by staining on a glass slide, where a particular biomarker, such as a protein or nucleic acid ofinterest, can be stained with H&E and / or with a colored dye, hereafter “chromogen” or “chromogenic moiety.” IHC staining is a common tool in evaluation of tissue samples for the presence of specific biomarkers. In situ hybridization (“ISH”) may be used to detect target nucleic acids in a tissue sample. ISH may employ nucleic acids labeled with a directly detectable moiety, such as a fluorescent moiety, or an indirectly detectable moiety, such as a moiety recognized by an antibody which can then be utilized to generate a detectable signal. Further approaches such as Fluorescence ISH (“FISH”) have been also applied.
[0006] Compared to other detection techniques, such as radioactivity, chemo-luminescence or fluorescence, chromogens generally suffer from much lower sensitivity, but have the advantage of a permanent, plainly visible color which can be visually observed, such as with bright field microscopy. However, more substrates with additional properties which may be useful in various applications, including multiplexed assays, such as IHC or ISH assays, are needed.
[0007] Additional capabilities for advanced image analysis of histological slides which may be used, for example, in digital pathology may improve detection and assessment of specific molecular markers, tissue features, and organelles, or the like. However, achieving high accuracy has been a challenging issue not only for machine learning based algorithms but also for experienced professionals.Brief Summary
[0008] In recent years, digital pathology has gained more popularity as many stained tissueslides are digitally scanned with high resolution (e.g., 40*) and viewed as whole slide images (“WSIs”) using digital devices (e.g., PCs, tablets, etc.) instead of standard microscopes. Having the information in a digital format enables digital analyses that may be applied to WSI to facilitate diagnoses. Recently, quantitative staining approaches have been developed, which convert antibody / antigen complexes into dots. These dots may then be detected, and provide a quantitative measure for expression of a desired molecules (e.g., proteins).
[0009] Developing a robust automated approach is particularly challenging due to the huge diversity of shape, color, orientation, and density of staining in different tissue and stain types. Hence, there is a need for more robust and scalable solutions for implementing digital microscopy imaging, and, more particularly, to methods, systems, and apparatuses for implementing digital microscopy imaging using deep learning-based segmentation, implementing instance segmentation based on annotations, and / or implementing user interface configured to facilitate user annotation for instance segmentation within biological samples.
[0010] Accordingly, in some aspects the present disclosure provides systems and methods that facilitate detection of dots that result from staining a biologic sample with a quantitative approach using artificial intelligence (“Al”), e.g., machine learning (“ML”) models. Such detection may be used to generate (a basis for) ground truth that may be further used to train the artificial intelligence.
[0011] According to a first aspect, a method is provided of generating training data, the method comprising: i) obtaining a first image of a biological sample; ii) generating a first set of annotations of the first image based on a second image of said biological sample stained by a quantitative approach converting antibody / antigen complexes into dots; and iii) outputting the first image and a ground truth including said first set of annotations as data for training a first ML model, for analyzing a biological sample.
[0012] According to a second aspect, in addition to the first aspect, generating the first set of annotations is based on a statistical measure of the dots in the second image.
[0013] According to a third aspect, in addition to the second aspect, the statistical measure includes a number of dots and / or a density of dots.
[0014] According to a fourth aspect, in addition to the second aspect or the third aspect, generating the first set of annotations comprises generating an annotation for each region out of a plurality of regions of the first image by determining the statistical measure for a matching region in the second image matching said region in the first image.
[0015] According to a fifth aspect, in addition to the second aspect or the third aspect, generating the first set of annotations comprises: i) obtaining a plurality of regions in the second image; ii) for each of the plurality of regions in the second image: a) computing the statistical measure, and b) associating the computed statistical measure as an annotation out of the first set of annotations with a location, in the first image, of a region matching said region in the second image.
[0016] According to a sixth aspect, in addition to the fourth aspect or the fifth aspect, each region is a rectangle, a square, or a circle at a previously configured position in the first image.
[0017] According to a seventh aspect, in addition to any of the fourth to sixth aspect, said first set of annotations includes said statistical measure associated with a location of said respective region within the first image and / or with a location of said respective matching region on the second image.
[0018] According to an eighth aspect, in addition to any of the fourth to seventh aspect, at least one region of the plurality of regions is a closed area selected in the first image.
[0019] According to a ninth aspect, in addition to any of the first to eighth aspect, said generating the first set of annotations includes inputting the second image into a second processing; and obtaining, as the output of the second processing, said first set of annotations.
[0020] According to a tenth aspect, in addition to any of the first to ninth aspect, said first image is an image of the biological sample stained with IHC, H&E, and FISH.
[0021] According to an eleventh aspect, in addition to any of the first to tenth aspect, the method further comprises: i) obtaining the second image stained with a quantitative approach converting antibody / antigen complexes into dots, using a first concentration of the antibody / antigen; ii) obtaining a third image of the biological sample stained with a quantitative approach converting antibody / antigen complexes into dots, with a second concentration different form the first concentration, wherein the second image and the third image are images of respective slides with consecutive sections of the biological sample; iii) generating a second set of annotations of the first image based on the third image; and iv) outputting the first image and a ground truth including said second set of annotations as data for training the first ML model for analyzing a biological sample.
[0022] According to a twelfth aspect, in addition to any of the first to eleventh aspect, the staining approach is an IHC staining to visualize HER2 epitope in human formalin fixed, paraffin embedded tissue.
[0023] According to a 13-th aspect, a computerized method is provided of training an ML model, the method comprising: inputting, to the ML model the first image and the ground truth output by the method of generating training data according to any of the first to twelfth aspect; and modifying at least one parameter of the machine learning model based on the input first image and the ground truth.
[0024] According to a 14-th aspect, in addition to the 13-th aspect, the method further comprises iterations of performing the method of generating training data according to any of the first to twelfth aspect and said steps of inputting and modifying.
[0025] According to a 15-th aspect, a method of analysis is provided, comprising: inputting a fourth image of the biological sample, or a first image of a second biological sample, to an ML model, trained using the method according to the eleventh aspect; and obtaining, from the output of the ML model an analysis of the biological sample or the second biological sample.
[0026] According to a 16-th aspect, in addition to the 15-th aspect, the analysis result indicates a presence, absence, or amount of an antigen in the biological sample or the second biological sample.
[0027] According to a 17-th aspect, in addition to the 15-th or 16-the aspect, the analysis is a classification of the biological sample or the second biological sample.
[0028] According to an 18-th aspect, in addition to any of the first to the 17-the aspect, said ML model is a first ML model and said generating the first set of annotations includes: i) detecting objects in an image of a biological sample, the detecting comprising: a) obtaining an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots, and b) detecting dots in said image of a biological sample using a trained second ML model pre-trained to detect dots; ii) inputting the second image into a second processing; and iii) obtaining, as the output of the second processing, said first set of annotations.
[0029] According to a 19-th aspect, in addition to the 18-the aspect, the detecting further comprises generating one or more annotations associated with the image of a biological sample and with the dots detected in the detecting step.
[0030] According to a 20-th aspect, in addition to the 18-th or 19-the aspect, said one or more annotations indicate location of respective one or more detected dots.
[0031] According to a 21-st aspect, in addition to the 20-the aspect, the location of dots is a location in three spatial dimensions.
[0032] According to a 22-nd aspect, in addition to any of the 19-th to 21-st aspect, the detecting further comprises: a) providing an input interface capable of enabling a human user to obtain an updated one or more annotations by: i) deleting an annotation from the one or more annotations, ii) modifying an annotation among the one or more annotations, and / or iii) adding an annotation to the one or more annotations; and b) storing the updated one or more annotations.
[0033] According to a 23-rd aspect, in addition to the 22-nd aspect, the method further comprises using the updated one or more annotations as a ground truth to train said second ML model or another Al model for detecting dots.
[0034] According to a 24-th aspect, in addition to any of the 19-th to 23 -rd aspect, the method further comprises providing an annotation viewing interface configured to display said image of the biological sample together with said one or more annotations.
[0035] According to a 25-th aspect, in addition to the 24-th aspect, the annotation viewing interface is a side-by-side viewing interface configured to display a first view showing a first- view image based on said image of the biological sample and including said one or more annotations overlaid beside a second view showing a second-view image of the biological sample.
[0036] According to a 26-th aspect, in addition to the 25-th aspect, said second-view is captured with settings different from those of said image of the biological sample and includes annotations corresponding to said one or more annotations of the first-view image.
[0037] According to a 27-th aspect, in addition to any of the 25-th or 26-th aspect, said second-view image of the biological sample is based on or includes a z-stack including two or more mutually different focal plane images of the same field of view, FOV, of the biological sample.
[0038] According to a 28-th aspect, in addition to the 19-th aspect, said second-view image is obtained by flattening or 3D deconvolution of the z-stack including combining the two or more focal plane images into a single image of the biological sample.
[0039] According to a 29-th aspect, in addition to any of the 24-th to 28-th aspect, the annotation viewing interface is configured to view each of said one or more annotations as a graphics of an unfilled contour of a preconfigured shape surrounding one of the detected dots.
[0040] According to a 30-th aspect, in addition to the 29-th aspect, the annotation viewing interface is configured to enable switching between focal planes of the z-stack.
[0041] According to a 31-st aspect, in addition to the 29-th or 30-th aspect, the annotation viewing interface is configured to: i) enable a user to select a focal plane out of the z-stack and to mark a dot in the selected focal plane, and ii) store an identification of said selected focal place in association with the marked dot.
[0042] According to a 32-nd aspect, in addition to any of the 29-th to 31-st aspect, the shape is a rectangle or a square positioned such that the one of the detected dots is in its geometric center.
[0043] According to a 33-rd aspect, in addition to any of the 29-th to 32-nd aspect, the annotation viewing interface is configured to enable switching between: i) viewing each of said annotations as a graphics of a dot co-located with one of the detected dots, and ii) said viewing each of said annotations as a graphics of an unfilled contour of a preconfigured shape surrounding one of the detected dots.
[0044] According to a 34-th aspect, in addition to the 33-rd aspect, said switching is triggered by changing a field of view of said image of the biologic sample.
[0045] According to a 35-th aspect, in addition to any of the 18-th to 34-th aspect, the detection of the dots further includes categorizing or filtering said one or more dots based on additional cellular or non-cellular markers.
[0046] According to a 36-th aspect, in addition to any of the first to 35-th aspect, the detection of the dots further comprises categorizing or filtering said one or more dots based on additionalone or more images of the biologic sample including one or more images of the same slide or of consecutive slides.
[0047] According to a 37-th aspect, in addition to any of the 18-th to 36-th aspect, the method further comprises training said second ML model for detecting objects in an image of a biological sample, the training comprising steps of: a) obtaining an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots; b) obtaining one or more annotations associated with the image of the biological sample and with the dots; and c) adapting one or more parameters of said second ML model according to said image of a biological sample and, as a ground truth, said one or more annotations into the Al model.
[0048] According to a 38-th aspect, in addition to the 37-the aspect, the training further comprises: a) obtaining an image of a negative control biological sample stained with an IHC or fluorescence-based approach; and b) adapting one or more parameters of said Al model according to said image of a negative control biological sample and, as a ground truth, no annotations or an annotation indicating no dots.
[0049] According to a 39-th aspect, a computer program is provided, stored on a non- transitory medium and including instructions which when executed on one or more processors causes the one or more processors to perform the steps of the method according to any of the first to 38-th aspect.
[0050] According to a 40-th aspect, a training data generating device is provided, comprising: a data interface; and processing circuitry that, in operation, i) obtains a first image of a biological sample via the data interface; ii) generates a first set of annotations of the first image based on a second image of said biological sample stained by a quantitative approach converting antibody / antigen complexes into dots; and iii) outputs the first image and a ground truth including said first set of annotations as data for training a first ML model, for analyzing a biological sample.
[0051] According to a 41-st aspect, methods of determining or estimating a level of antibody expression are provided, using any of the analytical or training methods described herein, wherein such methods further comprise determining a number of dots per cell in one or more cells in the biological sample or the second biological sample, using a first ML model or a second ML model, as described herein, and determining or estimating a level of antibody expression in the biological sample or the second biological sample based on the determined number of dots per cell. In some aspects, the level of antibody expression in the biological sample or the second biological sample is determined or estimated based on a mean number ofdots per cell. In some aspects, the biological sample or the second biological sample comprises formalin-fixed and paraffin-embedded cells
[0052] It is noted that the present disclosure also provides devices of which the processing circuitry performs any of the methods described herein. The present disclosure provides an integrated circuit which embodies the processing circuitry as described above.Brief Description of the Drawings
[0053] A further understanding of the nature and advantages of particular embodiments may be realized by reference to the remaining portions of the specification and the drawings, in which like reference numerals are used to refer to similar components. In some instances, a sub-label is associated with a reference numeral to denote one or multiple similar components. When reference is made to a reference numeral without specification to an existing sub-label, it is intended to refer to all such multiple similar components.
[0054] FIG. l is a schematic drawing illustrating obtaining two sections of a biologic sample and staining one of them quantitatively.
[0055] FIG. 2 is a flow diagram of an exemplary method generating training data.
[0056] FIG. 3 is a flow diagram of an exemplary method that makes use of the generated training data.
[0057] FIG. 4 is a schematic drawing of an exemplary implementation illustrating using the generated training data for training.
[0058] FIG. 5 is a flow diagram of an exemplary method that generates region-based annotations.
[0059] FIG. 6 is an exemplary view of an image and a schematic drawing of a region.
[0060] FIG. 7 is a set of views of quantitatively stained images, stained with different concentrations of antigen.
[0061] FIG. 8 is a schematic drawing of an exemplary implementation in which the annotations are expressed by a heat map.
[0062] FIG. 9 is a flow diagram of an exemplary method for analyzing a biologic sample.
[0063] FIG. 10 is a block diagram illustrating a device that may implement the method of ground truth generating, using it for training or an ML model, and analyzing biologic samples using the trained ML model.
[0064] FIG. 11 is a flow diagram exemplifying a method for dot detection and including schematic representation of an image of a biologic sample with and without annotations.
[0065] FIG. 12 is a block diagram illustrating an exemplary device for dot detection.
[0066] FIG. 13 is a flow diagram illustrating a method got annotation generation and updating, as well as a method for training of an artificial intelligence on the fly.
[0067] FIG. 14 is a schematic drawing of graphical user interface allowing displaying of an image of a biological sample and the associated annotations alongside with some exemplary buttons for updating annotations.
[0068] FIG. 15 is an exemplary excerpt of an image of a biologic sample stained with qlHC.
[0069] FIG. 16 is a schematic drawing illustrating a side-by-side view of a graphical user interface.
[0070] FIG. 17 is a schematic drawing illustrating some possible forms of an annotation.
[0071] FIG. 18 is an exemplary flow diagram illustrating phases connected with application of an artificial intelligence.
[0072] FIG. 19 is an exemplary flow chart illustrating a method for training an artificial intelligence model.
[0073] FIG. 20 is an exemplary content of a memory storing functional modules for configuring processing circuitry to perform the modules’ functionalities.
[0074] FIG. 21 is an exemplary block diagram illustrating a system including the dot detection device of Fig. 2.
[0075] FIG. 22 is a schematic drawing illustrating application of qlHC.
[0076] FIG. 23 is a graph illustrating performance of an exemplary implementation of the present disclosure, in particular, qlHC dot count in circular patches versus GE051 detailed score.
[0077] FIG. 24 is a graph illustrating performance of an exemplary implementation of the present disclosure, in particular modeling performances on validation set with region separation without tissue separation: the graph shows predictions versus ground truth of a test set averaged across tissue.
[0078] FIG. 25 is a set of graphs illustrating performance of an exemplary implementation of the present disclosure, in particular it shows in (a) qlHC patch counts distribution of training set and thresholds for bin labeling, (b) confusion matrix of test set patch prediction versus ground-truth, (c) predictions versus ground truth of test set averaged across regions in each tissue, and (d) predictions versus ground truth of test set averaged across tissue.
[0079] FIG. 26 provides a set of images of formalin-fixed paraffin-embedded cell lines with different levels of HER2 expression stained with IHC (top row) and qlHC (bottom row).
[0080] FIG. 27 is a graph showing the measured qlHC dots / cell in an exemplary set of cell lines with different HER2 expression.
[0081] FIG. 28 provides a pair of images of the same biological sample stained with qlHC, with detected qlHC dots annotated using points (left image) and boxes (right image).Detailed Description
[0082] The present disclosure relates, in general, to methods, programs, systems, and apparatuses for facilitating digital microscopy imaging (e.g., digital pathology or live cell imaging, etc.). More specifically, the present disclosure relates to implementing digital microscopy imaging using quantitative staining approaches converting antibody / antigen complexes into dots and using the dots as a quantitative measure to generate ground truth for analyzing otherwise stained or unstained portions.
[0083] The present disclosure relates generally to methods and devices for use, for example, in multiplexed assays or in cases of consecutive slides that may be observed together for detecting target molecules or other parts of biological tissue. Such methods have a wide utility in diagnostic applications, in choosing appropriate therapies for individual patients, or in training neural networks or developing algorithms for use in such diagnostic applications or selection of therapies. The present disclosure may also facilitate implementing annotation data collection and autonomous annotation, and, more particularly, implementing imaging of biological samples for generating training data for developing deep learning based models for image analysis, for cell classification, for feature of interest identification, and / or for virtual staining of biological samples.
[0084] The following detailed description illustrates a few exemplary embodiments in further detail to enable one of skill in the art to practice such embodiments. The described examples are provided for explanatory purposes and are non-limiting and non-exhaustive.
[0085] In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the described embodiments. It will be apparent to one skilled in the art, however, that other embodiments of the present disclosure may be practiced without some of these specific details. In other instances, certain structures and devices are shown in block diagram form. Several embodiments are described herein, and while various features are ascribed to different embodiments, it should be appreciated that the features described with respect to one embodiment may be incorporated with other embodiments as well. By the same token, however, no single feature or features of any described embodiment should be considered essential to every embodiment described and / or claimed herein, as other embodiments may omit such features.
[0086] In clinical pathology and in particular in digital pathology, detection of protein in intact formalin-fixed, paraffin-embedded tissue may be performed using immunohistochemistry, which is semi-quantitative. For example, recent approaches have brought quantitative immunohistochemistry (qlHC) that enables quantification of e.g. a protein directly in formalin-fixed, paraffin-embedded tissue by counting of dots. The qlHC technology can be combined with standard immunohistochemistry, and assessed using standard bright- field microscopy or image analysis. The qlHC method is in principle similar to classic IHC, and like classic immunohistochemistry, the basis for the amplification is enzyme deposition (typically Horse radish peroxidase (“HRP”)). A primary antibody binds a target (protein / receptor) of interest. A secondary, HRP -labeled antibody recognizes the primary antibody and binds. These first two steps in the qlHC reaction are directly comparable to standard IHC. However, in qlHC, only a predetermined fraction of the secondary antibodies is labeled. The secondary-labeled antibody is mixed with non-labeled antibody to increase robustness of the assay. Finally, an amplification reaction generates a dot centered around the labeled antibody, i.e., directly at the single target. The number of dots can be counted and as the ratio between labeled and unlabeled secondary antibody is known, the qlHC assay allows for a direct correlation between number of dots and amount of biomarker present in the tissue.
[0087] However, quantitative approaches are not limited to qlHC and may be also applied with ISH methods against DNA or RNA, imaging mass cytometry or immunofluorescencebased methods, or the like.
[0088] Some embodiments herein provide methods and devices that use qlHC dots (or dots obtained by a similar approach) for generating accurate quantitative ground truth e.g., for IHC stained slides.
[0089] Currently, a still leading solution to generate ground truth for variously stained (e.g., IHC and H&E) tissue slide interpretation and scoring tasks is still manual annotation. However, manual annotations are tedious, time consuming and - more importantly - subjective and thus, in many cases a consensus score by board of pathologists is required. There have been some other existing solutions, such as a method for fluorescence IHC on consecutive sections. Accordingly, while using fluorescence as an objective chemical “annotation” layer to provide quantitative signal, it is mostly comparative signal and requires guessing / calibration for the true expression of the inquired antibody. Other methods use molecular profiling / measurements to get an accurate quantitative estimation, however, these methods may lack spatial resolution and are generally more complex and expensive.
[0090] Generating ground truth based on quantitative staining
[0091] One of the problems addressed in the present disclosure is thus the challenging task of collecting accurate ground-truth data at scale for training a computational statistical model (e.g machine learning or Al model) for instance in whole slide pathology image interpretation. When working with medical data, achieving reliable ground-truth is even more challenging to collect and requires trained personnel. The dots (such as qlHC dots) provide chemically based signal of the protein expression and is not based on human interpretation of staining that is known to be biased and / or noisy in some cases.
[0092] Trained Al model for IHC stain interpretation based on the qlHC ground truth enables IHC stain quantification beyond what is visible by a human eye. With this approach, it may possible to enable generating a massive, quantitative, and objective ground-truth labelling by computationally aligning the qlHC WSI to the inquired WSI and projecting a measure of the dot distribution to the inquired WSI to serve as the ground-truth labels. This, in turn, can lead to the development of a more accurate computational Al model applied to WSIs. For example, a computational model could be implemented as a classification / regression model for predefined regions within the WSI. In other scenarios, the computational model may serve as a segmentation model for desired tissue structure.
[0093] Trained Al model for IHC stain interpretation based on the qlHC ground truth enables IHC stain quantification beyond what is visible by eye. The present approach is not limited to qlHC, but may be adapted for any other quantitative staining approaches.
[0094] In the present disclosure, non-limiting examples of addressing the issued mentioned above include methods and devices for generating ground truth annotations (for instance at scale for whole slide images (WSI) of pathology stained tissue slides) based on quantitative staining as illustrated in Fig. 1. This ground truth enables development and / or improvement of Al-based methods (or in general, ML-based methods) for quantitative estimation of e.g. protein expression in regular IHC or H&E stained tissue or for other applications.
[0095] As can be seen in Fig. 1, a biologic sample 110 is sliced to obtain two or more consecutive sections of the same tissue. In this example, the biologic sample is a breast tissue block. Thus, at first, two or more consecutive tissue sections are cut from the tissue block. Them these tissue blocks may be processed e.g. into formalin fixed, paraffin-embedded slices. Here, a first slice is a clinical tissue section 120 and a second slice 130 is a tissue section to be used for generating one or more annotations.
[0096] One section (e.g. the first slice) may be stained with the required IHC staining and another one (e.g. the second slide) with qlHC for the inquired antibody. In Fig. 1, two sections 120, 130 from a block with a very low expression of human epidermal growth factor receptor2 (HER2) protein have been cut. The first section 120 is stained with membranal HER2 staining (GE001), with nearly undetectable stain expression, and the second section 130 is stained with quantitative IHC (qlHC). The amount of qlHC dots in the second slice is used as a ground truth label for HER2 expression in the matching area in the tissues stained with GE001. On the right hand side of Fig. 1 on the top are shown two examples 140, 150 of a first section of a tissue stained with GE001. There is no visible staining in the captured tumor regions (corresponding to biologic samples) here. On the bottom right hand side of Fig. 1, the corresponding two examples 160, 170 of a second section of the tissue stained with glHC are shown. The number of dots is proportional to the HER2 expression and can be used as a ground truth for the respective images 140, 150.
[0097] It is noted that the qlHC staining and / or the IHC staining may be performed by more than one consecutive tissue sections. This may improve the overall accuracy, if the more than one consecutive tissue sections are used as well either for generating additional grount truth (if qlHC stained) or to generate a further correlated input (if stained by IHC or unstained or otherwise stained).
[0098] According to an embodiment, a method 200 is provided for generating training data as illustrated in Fig. 2. The method 200 comprises the step 210 of obtaining a first image of a biological sample. Such first image may be, for instance, the clinical tissue section 120 illustrated in Fig. 1. It is noted that the first image may be any image of the biologic sample (stained or unstained).
[0099] The method 200 further comprises the step 220 of generating a first set of annotations of the first image based on a second image of said biological sample stained by a quantitative approach converting antibody / antigen complexes into dots. The first set of annotations comprises one or more annotations. As will be described below, there are many possibilities as to the format of the annotations. For example, the second image itself may serve as a ground truth (annotation) of the first image. On the other hand, the dots after a detection applied to the second image, may serve as annotations or a count of dots or any statistics of the dots.
[0100] The second image may be an image of the same biologic sample section as the first image. For example, the biologic sample is captured to produce the first image and then stained with a quantitative approach and captured again to produce a second image. In addition or alternatively, the biologic sample may be stained and captured to produce the first image and then stained again with a quantitative approach and captured again to produce a second image. The term “alternatively or in addition” here indicates that it is possible to have more than oneimages of the same section stained (e.g. sequentially) with different quantitative or non- quantitative approaches.
[0101] However, the first image and the second image may alternatively be images of two respective sections of the same biologic sample as shown in Fig. 1. In particular, one section may be unstained and another one stained quantitatively. Alternatively, or in addition, one section may be stained with a first approach and another one section may be stained with a second approach. The first approach may be non-quantitative and the second approach may be quantitative. The term “alternatively or in addition” here indicates that it is possible to have more than one sections stained with different quantitative or non-quantitative approaches.
[0102] The quantitative approach converting antibody / antigen complexes into dots may be e.g., a qlHC (which is exemplified below with reference to Fig. 22). However, the present disclosure is not limited to qlHC staining and any other quantitative staining approach that is known in the art may be utilized.
[0103] The method 200 further includes the step 230 of outputting the first image and a ground truth including said first set of annotations as data for training a first ML model, for analyzing a biological sample.
[0104] The outputting step 230 thus outputs a training data pair of an input image and the ground truth. The first ML model may be any kind of machine learning model such as Al that may be formed by one or more layers of a neural network (e.g., a deep network). For the purpose of image processing, convolutional neural networks may be suitable as an example. For instance, a ResNet or any other kind of a classification architecture may be employed. However, the present disclosure is not limited to convolutional networks or to networks which include convolutional layers. In general, any neural networks such as multilayer perceptron or the like may be used.
[0105] The above mentioned analyzing of the biologic sample may be, for instance, a classification of the input image. For example, the output of the first ML model may be a class specifying whether or not (or with which probability) the analyzed biologic sample includes certain objects or structures, e.g., tumor cells or the like. However, the present disclosure is not limited thereto and the first ML model is not necessarily a classification model that classifies an input image into some discrete groups. Rather, the first ML model may be a regression model. For example, when using qlHC that it is quantitative, the first ML model may be trained as an estimation model which goes beyond classifying into discreet groups. Rather, it may output a position on a gradient. In this way, for example, different cancer grades that form a continuum rather than discrete groups may be analyzed more suitably.
[0106] It is noted that the present disclosure is not limited to applications in analyzing cancer cells or to digital pathology. For example, any other kind of immune diseases may be analyzed based on corresponding biologic samples.
[0107] According to an example, the step 220 of generating the first set of annotations is based on a statistical measure of the dots in the second image. For instance, the statistical measure includes a number of dots and / or a density of dots. In order to obtain such measure, in some exemplary implementations, the dots may be obtained by an automated processing that recognized dots (e.g. their respective location). Then the recognized dots may be counted to obtain their number or another measure such as density.
[0108] However, the present disclosure is not limited to these measures and does not necessarily involves counting of the dots. For example, it is conceivable to directly use the quantitatively stained (second) image as the ground truth, without previously detecting or recognizing the dots. It is further conceivable, especially in cases in which the dot staining has a distinctive color, different from the prevalent color of the biologic sample, a color histogram may be an indicative of presence and / or amount of dots.
[0109] The results of the method 200 may be further used for training. This is illustrated in Fig. 3. In particular, Fig. 3 shows a computerized method 300 of training the first ML model. The method comprises inputting 310, to the first ML model, the first image and the ground truth output by the method 200 as described above with reference to Fig. 2. The training method 300 further comprises the step 320 of modifying at least one parameter of the first ML model based on the input first image and the ground truth.
[0110] In an exemplary implementation, the method 300 comprises iterations of performing the method 200 of generating training data and said steps of inputting 310 and modifying 320. This is illustrated in Fig. 3 by a dashed arrow leading from step 320 back to step 200. In other words, the generation of the ground truth can be performed ”on the fly“, e.g., during (interleaved with) training. However, it is noted that the training does not need to be performed on the fly and does not even need to be performed on the instance of the first ML model used for inference (analysis of the biologic sample). For example, the first ML model may be trained in a training phase preceding the employment of the ML model for the analysis of biologic samples in inference phase. However, the instance of the first ML model used for inference does not need to be trained in a training phase at all. For example, one or more parameters of the first ML model may be obtained from another source, e.g., from a result of training of another instance of said ML model (e.g., a model with a similar structure / architecture including corresponding adjustable parameters).
[0111] Fig. 4 schematically illustrates the training and possible actions taken before the training in an exemplary implementation. The first ML model is exemplified in Fig. 4 as a neural network 470 having adjustable weights. In this specific example, the first ML model 470 is configured for HER2 expression quantification. Such configuration is obtained by training (on the fly, pre-training of the ML model 470, or copying the trained weights from another ML model similar to the first ML model 470). The first ML model 470 may be used to quantify the HER2 expression within a very low range, which is one of advantages of the present disclosure. The first ML model 470 outputs 480 a range from 0 to 1 indicating the HER2 quantification, e.g., how much HER2 is present in an input image and / or in respective regions of the input image. Application to HER2 in Fig. 4 is only exemplary and not to limit the present disclosure. Moreover, it is noted that the above-mentioned range is merely exemplary, since according to an exemplary clinical protocol HER2 expression may be in the range between 0 up to score 1. However, any other range may be used instead.
[0112] The inputs to the first ML model 470 are annotations 440 that represent the ground truth and the input image 460 (e.g., corresponding to the first image obtained as described with reference to Figs. 1 and 2). Thus, based on these inputs, the weights of the ML model 470 may be adapted. The output 480 may be useful to test the quality of the inference in the training phase. In the inference phase, it represents the result of the analysis of the biologic sample.
[0113] Region based annotations
[0114] It is possible to have a single one annotation 440 for the first image. For example, this may be the number of dots within the first image. On the other hand, it may be advantageous to associate different regions of the first image with respectively separate annotations, i.e., annotations determined separately for the matching regions. It is noted that the size and shape of the regions in the first image and the matching regions in the second image are not necessarily the same.
[0115] For example, the step 220 of generating the first set of annotations comprises generating an annotation for each region out of a plurality of regions of the first image by determining the statistical measure for a matching region in the second image matching said region in the first image. A region is not limited here to in any particular shape or size or way of obtaining the region.
[0116] In order to obtain matching regions, alignment 450 may be applied to the regions. Such alignment may be a global alignment and / or a local alignment. A global alignment refers to an alignment of the entire region or image. A local alignment aligns portions of a region and / or image. The present disclosure is not limited to any particular alignment approach. Aregion matching by looking for a match that minimizes a certain metric (e.g., sum of absolute differences or the like) may be used. In other words, the alignment may be performed automatically. In addition or alternatively, the alignment may be performed or assisted by a human manually using a graphic user interface (“GUI”).
[0117] As illustrated in Fig. 5, for example, the step 220 of generating the first set of annotations comprises the step 510 of obtaining a plurality of regions in the second image.
[0118] The step 220 according to Fig. 5 further comprises for each z-th region of the plurality of regions in the second image: computing 520 the statistical measure; and associating 530 the computed statistical measure as an annotation out of the first set of annotations with a location, in the first image, of a region matching said region in the second image.
[0119] In other words, the method performs a loop and in each z-th loop iteration, z-th annotation (or annotations) of the region is determined and associated with the z-th region. The loop iterations may start with the first one of the plurality of regions, e.g., at iteration in which z=l (or z=0 - this is just a region numbering convention) and end with the last one of the plurality of regions. At the end (or at the beginning) of each loop iteration, i is incremented as illustrated in step 540. In other words, the loop repeats steps 520, 530 and 540 for increasing until all regions are associated with their respective annotations. For example, said first set of annotations includes said statistical measure associated with a location of said respective region within the first image and / or with a location of said respective matching region on the second image.
[0120] For instance, each region is a rectangle, a square, or a circle at a previously configured position in the first image. Fig. 6 illustrates such regions. In Fig. 6, a number of qlHC dots is calculated that are located within a circle 620 with an exemplary radius = 256 pixels. Thus, the label is generated for a region in the first image that corresponds to the region 620 in the second image, for which the dots are counted.
[0121] Fig. 6 shows a first view 610 of a GUI in that the region within which the dots are calculated is a circle 620. Thus, location of the region may be given by a center of the circle. Fig. 6 shows further a second view 650 of the GUI in that the region within which the dots are calculated is a circle 660 or a square 670. As also illustrated in Fig. 4, in an exemplary implementation, each region of the qlHC (second) image is used to generate a label for its corresponding patch in the IHC stained (first) image.
[0122] In general, the region form is not limited to circle, square, or rectangle. These shapes are easy to handle and provide an easy reference point for their location (e.g., center of thecircle, center of the square or rectangle, or top left corner of the square or rectangle, or another corner of the square or rectangle, or like).
[0123] In general, the regions may have any shape or size. In an example, at least one region of the plurality of regions is a closed area selected in the first image. Such closed area may be enclosed e.g., by a polygon or a free-hand drawn contour.
[0124] The step 510 of obtaining the regions in Fig. 5 may include automatic and / or human operated determination of the regions. For example, the first image may be segmented into same-size regions and then, annotations may be determined for the matching regions in the second image. In this example, it may be computationally efficient to have a region with a square or rectangular shape (as in Fig. 6, 660). In this way, it is possible to define nonoverlapping regions that cover the entire first image and the annotations for the respective regions.
[0125] However, it is noted that the present disclosure is not limited to non-overlapping regions. In general, the regions may overlap. In an example, the regions are circles (as shown in Fig. 6, 620, 670). The circles may overlap so that the entire first image is covered. However, it is also possible to define regions that do not overlap.
[0126] In the above examples, the first image was segmented into regions and annotations have been determined in the second image for regions matching the regions in the first image. However, the present disclosure is not limited to this way of operation. It is possible to segment the second image into regions, annotate these regions and then assign annotations to the matching regions in the first image.
[0127] The size of the region may be set by a user (e.g., the GUI may provide user with the corresponding input possibility) or determined automatically, e.g., according to the image or field of view (FOV) resolution or the like.
[0128] The regions to which an image is segmented may have all the same size or may differ in size. For example, it is conceivable to determine region size based on density of dots. For example, the smaller the density, the larger the circle.
[0129] It is noted that the regions do not have to cover the entire image. The regions may be regions of interest, defined by a user or automatically, which are annotated and the remaining portions of the image that are not annotated. For example, a region of interest may have a semantic associated with it such as duct or non-tumor, or the like.
[0130] When referring to an image, as indicated above, the first or the second image may be meant, as the segmentation may be performed in first or in second image. It is further noted that in case an alignment is performed as shown in step 450 of Fig. 4, it may be possible toapply the same kind of segmentation to the first and second image and assume that the collocated regions of the first image and the second image correspond - i.e., match with each other in the first and the second image.
[0131] In general, there may be a different splitting / covering (segmentation) applied to the input image (the first image) and the ground truth (second image). For example, the input may be a rectangular or a square patch (because of how images are represented in memory), but the ground truth may be computed over a circle centered on the square as illustrated in view 650 of Fig. 6. For example, the first image may be segmented into square regions 670. Then, in the second image, matching regions that are circular (620, 670) are determined. These regions may be bounding the square region or be bounded by the square region or be centered on the square region but have the same area (equiarea circle) as the square region. The latter is shown in the view 650. The equiarea circle would be located between the square-bounding and square- bounded circles.
[0132] As mentioned above, the image does not have to be covered by regions / annotations. Alternatively or in addition, the GUI may provide a user the possibility to draw or indicate a region for which then automatically the annotation is obtained. For example, the GUI may provide one or more of the following functions to a user: drawing a free-hand contour (enclosing a region), drawing a polygon contour, drawing a circle / square / rectangle at a user determined position and a user determined size, or the like. After the user draws the polygon (in the view of the first image or the second image), method 200 may be applied, which computes the annotation, e.g., counts the dots located within the user-drawn region and displays the count to the user in the GUI. It is noted that the annotation does not have to be dot count as already mentioned above.
[0133] Returning back to the example of Fig. 4, it is noted that a WSI may be first divided into patches and these may then be processed as described above for the first image and the second image. For example, in Fig. 4, input patches 460 are plural patches of the same WSI and they correspond to the first image. Each of the patches may be processed as the first image described above (independently of the other patches). The generation of the labels is performed in the block 440. Correspondingly, the second image may also be in form of a plurality of patches 410, stained quantitatively (here by way of example, using qlHC).
[0134] Said generating 220 the first set of annotations can include inputting the second image into a second processing; and obtaining, as the output of the second processing, said first set of annotations. The second processing may be any kind of processing. For example, it may include automatic dot detection (recognition). For instance, blocks 420 and 430 represent a specificexemplary way in which dots may be automatically detected in the second image and will be described later in more detail (see for instance section “Dot detection in stained images" below). The detected dots may then be used to compute the annotation(s) in block 440.
[0135] However, the present disclosure is not limited to such MLM based approaches. Rather, the second processing may be any kind of processing including dot detection by an algorithm using e.g., pattern matching or feature extraction or similar approaches.
[0136] It is noted that that the second processing is not to limit the present disclosure. The embodiments and exemplary implementations described above can also work without the second processing, e.g., by inputting directly the dot-stained image 410 or the like. Moreover, the annotation generation may be assisted or performed by a human user; for example, the qlHC images may in general be scored manually.
[0137] In summary, in Fig. 4, the WSI of sections 410 / 430 and 460 are aligned 450 and can be used in an end-to-end pipeline to train the first Al model 470 for scoring the inquired (first) tissue section 460 while its labelling is generated by a metric based on the qlHC dots. Here, the HER2 qlHC stained section 410 is used for ground-truth labelling 440 by detecting 420 / 430 and counting the qlHC dots in a predefined region. The qlHC “labels” together with their corresponding image patches from the inquired (and aligned) tissue section 460 are being used to train HER2 scoring model 470.
[0138] It is noted that the selection of region location and calculating annotation can be done once for all locations in the first image as the start of the training process. Alternatively. The annotations may be computed for each location every time it gets sampled during training (could in principle be more than once).
[0139] In the particular example of Fig. 4, similarly to Fig. 1, the first image is stained with a HercepTest™ mAb pharmDx (Dako Omnis) that is a semi quantitative immunohistochemical assay based on a primary monoclonal rabbit antibody (clone DG44) and an assay-specific visualization reagent. The assay determines HER2 protein overexpression in formalin-fixed, paraffin-embedded (FFPE) breast cancer tissues processed for histological evaluation.
[0140] Said first image may an image of the biological sample stained with immunohistochemistry (IHC), hematoxylin and eosin (H&E), or fluorescence in situ hybridization (FISH) or any other staining approach such as phase contrast, differential interference contrast, darkfield. It may even be unstained.
[0141] As mentioned above, the staining approach is a HER2 IHC based staining in the example of Fig. 1 and 4. However, the present disclosure is not limited to that and any staining to visualize a particular antibody or merely a special molecule in formalin fixed, paraffinembedded tissue may be applied. FFPE is not limiting either, depending on the type of biologic sample, other fixations and embedding approaches may be used instead.
[0142] Annotations for staining with different concentrations
[0143] The performance of the biologic sample analysis may depend on the density of the qlHC dots, which depends on the concentration level of the labelled secondary antibody (see 2230 in Fig. 22). Thus, when generating ground truth (e.g. the annotations), more than one image of the biologic sample may be considered, capturing the biologic sample stained with a different concentrations of the labelled secondary antibody / antigen. For example, the concentration may be selected according to the desired dynamic range of the expression of the target protein.
[0144] Fig. 7 shows three different qlHC dots densities corresponding to different antibody concentrations that have been exemplarily examined: 2 pM, 4 pM, and 8 pM of labelled secondary antibody (2230), respectively, in the respective views (a), (b), and (c). Above in the examples referred to with reference to Fig. 1 and 4, in the ultra-low expression regime, 4 pM concentration has been chosen in order to have reliable ground-truth labelling. In cases where a broader dynamic range is required, several qlHC stained consecutive sections can be used. According to an exemplary implementation, the method for generating ground truth comprises obtaining the second image stained with a quantitative approach converting antibody / antigen complexes into dots, using a first concentration of the labelled antibody / antigen (e.g. qlHC reagents such as the labelled secondary antibody); and obtaining a third image of the biological sample stained with qlHC with a second concentration different form the first concentration, wherein the second image and the third image are images of respective slides with consecutive sections of the biological sample.
[0145] Then a second set of annotations of the first image is generated based on the third image. The first image and a ground truth including said second set of annotations is then output as data for training the first ML model for analyzing a biological sample. The second set of annotations may be used in addition to the first set of annotations to train the MLM.
[0146] The analysis of the biologic sample may be, as also above, measuring antibody expression in tissue images where it is sometimes barley possible. Coverage of the desired range of protein expression may be achieved by tuning the concentration of qlHC reagents. Coverage of the desired range of expression with a robust number of qlHC dots may then be achieved. In addition, when dynamic range of the antibody is too large, a series of consecutive slides can be used with different qlHC dots concentrations to cover the dynamic range with the desired antibody expression resolution as indicated above (second set of annotations). It isnoted that the above first set of annotations and second set of annotations are not to limit the present disclosure. In general, there may be more than two different sets of annotations associated with an input image (the first image) of the biologic sample.
[0147] Visualization
[0148] In practice, it may be desirable to enable visualization of the labeled first image, e.g., to enable displaying of the first picture and an indication of the annotation(s) on a GUI. For example, dot distribution or any of its derived metrics could be overlaid / visualized on top of the inquired WSI (first image) and used as an assistive layer for a clinician pathologist. Alternatively, or in addition, the dot distribution or any of its derived metrics could be overlaid / visualized on top of the second image (e.g. qlHC image).
[0149] One possible way to visualize qlHC based antibody distribution / information is by using a semi-transparent heat map indicating for each pixel its regional qlHC dots count. This is illustrated in Fig. 8. In particular, Fig. 8 shows in view 810 the first image (e.g. IHC stained) and in view 820 the second image (e.g. qlHC stained). In view 830, a heat map is overlaid semi-transparently on the top of the first image. The transparency level and / or smoothness of the heat map may be varied to enable easier visualization of the number or the density of the qlHC dots.
[0150] However, the present disclosure is not limited to such visualization or to providing visualization at all.
[0151] Analyzing a biological sample
[0152] Fig. 9 illustrates a method 900 of analyzing a biological sample. The method 900 comprises inputting 910 a fourth image of an inquired biological sample to an ML model trained using the computerized method as described above with reference to Fig. 3 and 4. The method further comprises obtaining 920, from the output of the ML model an analysis of the inquired biological sample.
[0153] For example, the analysis result indicates presence, absence, or amount of an antigen in said inquired biological sample. In some implementations, the analysis is a classification of the biological sample. However, the present disclosure is not limited to such applications. As mentioned above, rather than outputting a discrete group or class, the MLM may be configured (trained or provided with parameters obtained by training) to output a degree / amount of the desired antigen or, in general, molecule.
[0154] Systems and devices
[0155] The methods described with reference to Figs. 2, 3, 5, and 9 may be performed by a correspondingly configured devices. For example, a training data generating device 1000 isprovided, illustrated in Fig. 10. The device 100 comprises a data interfacel080 and processing circuitry 1010 that, in operation, i) obtains a first image of a biological sample via the data interface; ii) generates a first set of annotations of the first image based on a second image of said biological sample stained by a quantitative approach converting antibody / antigen complexes into dots; and iii ) outputs the first image and a ground truth including said first set of annotations as data for training a first ML model for analyzing a biological sample.
[0156] Moreover, the device 1010 may further comprise a memory 1020, a display control device 1030, a communication interface 1040, and an operation input interface. The memory 1020 may store a program that when executed on the processing circuitry 1010, performs steps of any of the methods described above. The processing circuitry 1010 may include or be one or more processors. However, the processing circuitry is not limited to general purpose processors or special purpose processors. It may include programmable hardware, application specific hardware and / or other kinds of electronics. The memory 1020 may include one or more volatile and / or non-volatile memory modules and may also store training data, the first and / or the second image, and / or the annotations or the like. However, it is noted that these data may be also stored in an external storage. The display control 1030 is a module that in operation controls a display device (which may be but does not have to be built in the device 1000) to display a user interface that is suitable for and / or configured to display for instance the first image and / or the second image and / or the annotations. For example, the display may include views as described with reference to Figs. 6, 7, or 8 or the like. The display may display annotations overlaid on the first and / or the second image.
[0157] The communication interface 1040 may serve for communication with the external devices such as storages or servers or clients via using or more standardized or proprietary, wireless or wired technologies. For instance, the communication interface 1040 may implement a protocol stack for accessing internet on Ethernet or Wireless LAN or cellular communication basis or the like. It may implement a protocol stack that enable communication via Bluetooth or the like. The operation input interface 1050 enables a user to provide input. For example, the operation input interface may be formed by a touch screen (that may be the same as the display device mentioned above) and / or mouse in combination with a user interface displayed on the display and / or a standard or proprietary keyboard and / or specific operation panels / buttons or the like. The device 1000 may be part of the system described in Fig. 21, in addition or alternatively to the device 1200.
[0158] Embodiments and examples of the present disclosure and in particular, the use of qlHC stained tissue for generating ground truth for the WSI includes may provide advantagessuch as an accurate alignment of the IHC stained WSI and qlHC stained consecutive section, calibration of existing scoring methods automatic / manual, development of accurate models of qlHC dots detection, and development of statistical metrics for the quantification of the qlHC dots distributions analysis (e.g., kernel density estimation, spatial analysis), possibly together with cell distribution.
[0159] In general, quantitative staining approaches may be used to train an MLM for analyzing samples stain with other approaches. This enables a more accurate annotations and consequently a better training and inference.
[0160] Dot detection in stained images
[0161] As described above, in order to generate the annotations (220, 440), it may be advantageous to recognize the location of the dots in the second image as briefly indicated with reference to Fig. 4 above, and in particular to blocks 420 and 430.
[0162] The dot detection may be performed in any way, e.g., fully or partially under assistance by a human, which may lack efficiency. Classical computer vision algorithms may be used for detecting dots, for example, based on color difference, morphology, shape and size. However, such methods may suffer from some limitations. For example, such methods may not be sufficiently robust to changes in staining quality, color and variable shape and / or size of dots. Moreover, such methods are not easy to tune in order to prevent or at least reduce false positive detections due to dot-like structures that may be present in a biological samples. For instance, in a biologic tissue, objects or structures may look similar to dots but are not necessarily the dots - such objects may be cell nucleoli, or the like.
[0163] In order to provide a more efficient approach, in an embodiment illustrated in Fig. 11, an image of a biologic sample (i.e., the second image mentioned above) is obtained 1110, stained with a quantitative approach converting antibodies or antigens into dots. Such exemplary image 1160 of a biologic simple is shown in Fig. 11, including dots - one of the dots being marked by an arrow 1150.
[0164] Then, dots are detected in said image 1160 using a second machine learning model (MLM). In order to be able to detect the dots, the second MLM is pre-trained to detect dots. The detection of the dots is thus performed by inputting 1120 the image 1160 into the second MLM and by obtaining 1130 as an output of the second MLM the result of the detection, i.e., a location of the detected dots within the image 1160. These locations are illustrated in an image 1190 (corresponding to image 1160) by rectangular contours, such as rectangle 1170 surrounding position of the dot pointed to by the arrow 1150 in image 1160. In the following,when referring to the MLM, what is meant is the second MLM for detecting dots, not the first MLM for analyzing the biologic sample.
[0165] In this context, “pre-trained” means that the MLM is configured with one or more parameters that have been obtained in a training phase preceding the application of the MLM for the detection of dots (inference phase). It is not necessary to perform the actual training phase for each MLM instance used for detection of the dots. The training may be performed once for one MLM instance and the resulting parameters of such trained MLM instance may then be stored and provided to configure other MLM instances. The present disclosure is not limited to any particular way in which the MLM is parametrized. As will be described below, training phase may be even performed - even on the fly.
[0166] The quantitative approach converting antibodies or antigens into dots includes conversion of antigen-antibody complexes using staining into dots as described above, e.g. for qlHC. The dots are obtained by massive amplification by several steps of enzyme-catalyzed deposition at the sites of primary antibody where secondary antibody-dextran-HRP polymers are bound.
[0167] The method 1100 of Fig. 11 is a computerized method for detecting objects in an image such as 1160 of a biological sample. In some embodiments, the biological sample might include, without limitation, one of a human tissue sample, an animal tissue sample, or a plant tissue sample, and / or the like, where the objects of interest might include, but is not limited to, at least one of normal cells, abnormal cells, damaged cells, cancer cells, tumors, subcellular structures, or organ structures, and / or the like.
[0168] Some advantages of the method 1100 of Fig. 11 over classical image processing methods may include a more accurate and robust detection thanks to using an MLM, e.g., an Al-based model. Fig. 11 shows an image 1160 of a tissue stained with qlHC scanned and shown at 40x magnification. qlHC dots can be seen as DAB (3,3 ’-Diaminobenzidine) brown dots. Cell nuclei are stained with Hematoxylin (in practice blue). However, other approaches and resulting colors are possible.
[0169] The computerized method 1100 of Fig. 11 may be performed by a device 1200 with an exemplary structure shown in Fig. 12. In particular, the device 1200 may be a dot detection device. Processing circuitry 1210 may perform steps 1110, 1120, and 1130 described with reference to Fig. 11. In particular, the obtaining of the image 1160 of a biologic sample may be performed via a processing circuitry interface 1280. For example, the processing circuitry interface may be an interface to a bus 1290 or another communication medium. Via the processing circuitry interface 1280, the processing circuitry 1210 may be connected to variousother modules of the device 1200 such as a memory 1220, a display control 1230, a communication interface 1240, and an operation input interface 1250. It is noted that this structure is only exemplary and schematic. For example, the bus 1290 may in fact be a communication system including one or more buses and / or other communication media. The processing circuitry may be any piece of hardware including one or more processors (such as general-purpose processors) and / or programmable hardware pieces (such as FPGAs or the like) and / or application specific hardware (such as ASICs) and / or electronic circuits or elements. When including processors, the memory 1220 may store software that may be capable of configuring the processing circuitry (when running on it) to perform the method 1100. The memory 1220 may, for example, store the MLM and / or the trained parameters of the MLM.
[0170] The dot detection device 1200 may further include the display control module (display controller) 1230 that, in operation, controls a display device which may be connected to or a part of the device 1200. For example, the display control module 1230 may control the display device to display the image 1160 and / or the results of the detection or the like. The communication interface 1240 is an interface for communication of the device 1200 with other devices. For example, it may include communication interfaces such as Ethernet, WiFi (IEEE 802.11), Bluetooth, cellular communication networks (e.g., LTE, New Radio, or the like), or any other communication interfaces - standardized or proprietary. The operation input interface 1250 is an interface for inputting operation commands. It may be an interface enabling a human user input (e.g. via a keyboard, specialized keys or buttons, mouse, joystick or the like).
[0171] As shown in Fig. 13, the dot detection method 1100 may be used as an input for generating 1300 one or more annotations (such as 1170) associated with the image 1160 of a biological sample and with the dots (such as 1150) detected in the detecting (step 1120).
[0172] The generating 1300 of annotations means in general processing the output of the MLM to obtain an annotation in a form further usable by a human or a machine. The annotation may be generated by the MLM or by another processing following the MLM. For example, said one or more annotations indicate location of respective one or more dots detected in the detecting step. In general, the location may be annotated by providing coordinates of the detected dots within the image 1160, e.g., horizontal and vertical spatial coordinates x and y. In some exemplary implementations, the location of dots is a location in three spatial dimensions. In other words, in addition to the horizontal and vertical spatial coordinates x and y, the location includes depth coordinate z.
[0173] The annotations may indicate the location by way of explicitly providing numbers denoting the coordinates, and / or by providing markers that are superimposed to the image 1160 on the positions or in the proximity of the positions of the detected dots. Such markers may be empty rectangles as the rectangle 1170 in Fig. 11 or may have another shape or size.
[0174] In general, one of the advantages of the dot detection presented herein is that the MLM may be trained to distinguish between the dots resulting from a quantitative approach such as qlHC and other dot-like structures or objects in the biologic sample.
[0175] Updating annotations
[0176] It is noted that the annotations generated in step 1300 of Fig. 13 may be used in various ways. For example, they may be displayed 1350 to assist human users (e.g. pathologists or other clinical practitioners) and / or stored together with or in association with the image 1160.
[0177] However, the annotations may be also used as a ground truth for training MLMs. The annotations may be further refined 1360 - possibly iteratively - and then used as a ground truth for additional training of the MLM used for dot detection or for training of other MLMs.
[0178] Correspondingly, the method described with reference to Fig. 13 may comprise a step of providing an input interface capable of enabling a human user to obtain 1360 an updated one or more annotations. Such interface may correspond to the operation input interface 1250 shown in Fig. 12, the human user may be enabled to update the one or more annotations by: i) deleting an annotation from the one or more annotations, ii) modifying an annotation among the one or more annotations, and / or iii) adding an annotation to the one or more annotations; and the updated one or more annotations may be stored in association with the image 1160. They may be displayed 1350 again. This may be performed by the display controller 1230 of the device 1200 or Fig. 12.
[0179] The annotation process may thus be used in a cycle of improvement, where the user adjusts 1360 the annotation produced 1100, 1300 by the MLM. Then, a new MLM may be trained based on the refined annotations and this cycle can be repeated until a desired model is achieved (e.g., having a desired accuracy). In other words, the method may comprise a step of storing 1310 the updated annotations (e.g., in the memory 1220 or in another storage which may ne external to the device 1200) and using the updated one or more annotations as a ground truth to train 1320 said MLM. Such additional training on the fly may help continuously improving the MLM performance. The MLM additionally trained may then be used again for detecting 1100 dots.
[0180] However, it is noted that the present disclosure is not limited to implementations allowing for training 1320 in the fly. It is not necessary to re-train the MLM model used forinference. The updated annotations may be stored 1310 used for other purposes such as training another MLM for detecting dots or merely for viewing.
[0181] Any of the methods and apparatuses mentioned above may further provide an annotation viewing interface configured to display said image of the biological sample together with said one or more annotations. For example, in Fig. 12, the display control 1230 may control a display to view the image of the biological sample together with said one or more annotations.
[0182] Fig. 14 illustrates, in a schematic manner, an exemplary user interface 1400, with an image and annotation display field 1410 and display buttons 1420 (for adding an annotation), 1430 (for modifying an annotation, e.g., moving it to a different location), 1440 (for deleting an annotation), and 1450 (for saving the updated annotations). The image and annotation display field 1410 is an example for the annotation viewing interface. Within the image and annotation display field 1410, an image 1500 of the biologic sample as shown in Fig. 15 may be shown and annotations may be overlaid on it.
[0183] Fig. 14 shows also an example of an input interface formed by buttons 1420-1450 that enable updating of the annotations as described above. In this example, the input interface is a GUI including active fields represented by images of a respective buttons together with an input device capable of registering a user input. This may be any device including a computer mouse, a touch screen, a touch pad, a pen, arrow keys, speech control device, or the like. Fig. 14 shows merely exemplary user interface buttons for updating annotations. However, the present disclosure is not limited to such interfaces. In general, not all four buttons must be present. For example, a modification button 1430 may not be necessary, as the same effect may be achieved by deleting 1440 and adding 1420 annotations. It is further noted that the GUI buttons are only one example of input possibility. There may be physical keys (isolated or on a standard keyboard) associated with the respective actions of adding, modifying position, deleting, and / or saving the annotations.
[0184] Moreover, the GUI 1400 is not limited to the four annotation updating actions. There may be further buttons such as a button for triggering re-training of the MLM with the currently displayed image with annotations (in the GUI field 1410). Alternatively or in addition, there may be settings enabling selection of different kinds of annotations (e.g., annotations having different sizes and shapes) and / or different views or the like.
[0185] Regarding view configurations, according to an exemplary implementation, the annotation viewing interface is a side-by-side viewing interface configured to display a first view showing a first-view image based on said image of the biological sample and includingsaid one or more annotations overlaid beside a second view showing a second-view image of the biological sample.
[0186] A side-by-side interface 1600 is schematically illustrated in Fig. 16. It comprises a first viewing field 1610 and a second viewing field 1620. Such arrangement may be particularly ergonomic for a human user. The first view filed 1610 may show the biologic sample image 1160 (or an image based thereon) with annotations whereas the second view field 1620 may show the biologic sample image 1160 without annotations. One advantage of such viewing is that the user may see at the same time the annotations and the original image in which the annotation do not overlap with the content. This may be particularly suitable if there are closely located dots.
[0187] When referring to the first-view image based on said image of the biological sample, what is meant is that the first-view image may be directly said image of the biological sample or may be an image based on it. For example, the image based on the image of the biological sample may be the image of the biological sample filtered or otherwise processed, e.g., to produce an image with a lower / higher contrast or brightness or the like.
[0188] The second-view image may also include the image of the biological sample or it may include another image taken of the same biologic sample. For example, said second-view is captured with settings different from those of said image of the biological sample and includes annotations corresponding to said one or more annotations of the first-view image.
[0189] In an exemplary implementation, said second-view image of the biological sample is based on or includes a z-stack including two or more different focal plane images of the same field of view, FOV, of the biological sample. By different focal planes, what is meant is focal planes different from each other (mutually different). In general, the term “z-stack” imaging refers to obtaining a plurality of pictures taken at a set interval between the first and last planes of focus of a sample (e.g., a biological entity or tissue or in general any sample). Thus, each picture corresponds to a respective focus setting.
[0190] In an exemplary implementation, said second-view image is obtained by flattening or 3D deconvolution of the z-stack including combining the two or more focal plane images into a single image of the biological sample. The flattening or 3D deconvolution may be performed by any known approach. For example, for the flattening, extended depth of focus (EDF) may be used. EDF is an algorithm designed to scan through each picture in the set of pictures captured with different focus, and form a single composite image with all the parts that are determined to be in focus.
[0191] Provision of a single image based on several differently focused images has the advantage that such single image carries information combined from the different images and thus provides the user with a possibility to better distinguish the dots. Nevertheless, the present disclosure is not limited to displaying a single combined image. In some exemplary implementations, the annotation viewing interface 600 may provide a function of browsing through the images of the z-stack for a user in the second view 620 and / or a function for selecting a particular one of the different focal-plane images to be viewed.
[0192] In particular, the annotation viewing interface may be configured to enable switching between focal planes of the z-stack. For example, the annotation viewing interface is configured to: i) enable a user to select a focal plane out of the z-stack and to mark a dot in the selected focal plane, and ii) store an identification of said selected focal plane in association with the marked dot. In this way, some information of the depth (z coordinate) is obtained.
[0193] It is noted that in the above describes examples, it was assumed that the first-view image is in the view field 1610 whereas the second-view image is in the view field 1620. However, this was only for the sake of example. In practice, the first-view image may be in the second view field 1620 whereas the second-view image may be in the first view field 1610. In fact, the side-by-side view shown in Fig. 16 is only schematic and may be a part of a larger GUI including further functions such as those described with reference to Fig. 14 or the like. For example, the view 1600 may correspond to the viewing field 1410 in Fig. 14.
[0194] The present disclosure is not limited to displaying images captured at different respective focus planes. In addition or alternatively, images may be shown that were captured from a different angle or by different capturing apparatuses or with different settings or the like. For instance, the second-view image may be a consecutive slice of the same biologic sample or an image of the biologic sample, but differently stained, or the like.
[0195] It is noted that the second-view image may also include the annotations. In an exemplary implementation, annotations are presented in both views 1610 and 1620, and they are linked (associated with each other). The linking means that when annotation is moved in one of the views 1610, 1620 — it moves accordingly in the respective second view 1620, 1610. This may improve accuracy of the annotations (since the signal may not always be sufficiently clear in the original view). For example, the linking may be established by a preceding step of aligning the images of the two views together (globally, or per patch). The GUI may provide a possibility of selecting a function of a “linked view”. If a user selects the linked view function, the two image views (one of them or both including annotations) are linked - if one of theviews is moved (translated, rotated, and / or zoomed in or out) by a user within the viewing field, the other one of the views is moves correspondingly.
[0196] In general, the displaying or not of the annotations may be configurable by a user for the first-view image and / or the second-view image.
[0197] The user GUI (such as those described with reference to Figs. 14 or 16) may further comprise a zooming functionality which enables a user to change a field of view (“FOV”) of the image 1160 e.g., by zooming in or out, by translation or rotation, or the like.
[0198] Displaying annotations
[0199] The present disclosure is not limited to any particular shape, color or size of the annotations. In an example, the annotation viewing interface is configured to view each of said one or more annotations as a graphics of an unfilled contour of a preconfigured shape surrounding one of the detected dots. For instance, the shape is a rectangle or a square positioned such that the one of the detected dots is in its geometric center.
[0200] Examples of such annotations encircling respective dots in an image 1700 are provided in Fig. 17. Annotations 1710 (rectangle or square), 1720 (circle), or 1730 (rotated rectangle or square) are empty contours. It is noted that it is not necessary to completely enclose the dots within the annotation contour. Moreover, annotations that merely point to the dots are also conceivable and exemplifies as an arrow 1740 or even annotations that at least partially overlap with the dot and mark its center such as cross 1750 in Fig. 17. The later may enable a visually more precise localization of the dots.
[0201] In general, it may be advantageous to provide annotations that do not touch the dot and do not overlap the dot they are marking (such as 1710, 1720, 1730, 1740). Moreover, it may be advantageous to provide annotations that do not touch or overlap any dot (not even the dots in the proximity of the marked dot). In special cases, such exemplary displaying rules may be lifted. For example, in case dots are highly diffused (out of focus) the contour (annotation) may be allowed to overlap outer parts of the diffused dot. It is possible to have a displaying rule according to which the central part of the dot is not obscured or overlapped by the annotation or by any annotation. The later may be more difficult to achieve especially in cases of high dot concentrations. It may be more ergonomic to a human viewer, if the annotations do not overlap with each other. On the other hand, some overlap would not cause any problems and such annotations may still enable a human user to clearly see the marked dots.
[0202] In an exemplary implementation, the annotation viewing interface is configured to enable switching between: i) viewing each of said annotations as a graphics co-located withone of the detected dots, and ii) said viewing each of said annotations as a graphics of an unfilled contour of a preconfigured shape surrounding one of the detected dots.
[0203] It is noted that the graphics co-located (and at least partially overlapping) with one of the detected dots may be for instance a graphics of a dot, e.g., corresponding in size to the detected dot or fixed-size, irrespectively of the size of the detected dot. This is not to limit the present disclosure: the graphics co-located with one of the detected dots may have a different form, e.g., the cross 1750 or the like.
[0204] The switching may be automatically triggered or triggered by a user. Regarding automatic triggering, in an exemplary implementation, said switching is triggered by changing a field of view of said image of the biologic sample. For example, in case a zoom level is below a certain threshold (previously configured or fixed), the co-located (at least partially overlapping) displaying of the annotation is applied. In case the zoom level is above (or equal to) the certain threshold, the empty-shape annotation displaying may be applied.
[0205] Alternatively or in addition, the automatic triggering may be based on the contents of the field of view, for example based on the density of the dots and / or the proximity of the dots to each other. For instance, for densities higher than a threshold (pre-configured or fixed), annotations at least partially overlapping the dots and collocated with them are displayed. For the densities lower than or equal to the threshold, the empty contour annotations are displayed.
[0206] Regarding user-triggered annotations, an input interface may provide a possibility for a user to switch between annotation types (either at least partially overlapping and collocated or empty contours surrounding the dots). In addition or alternatively, the input interface may provide a possibility for a user to switch between (to configure) sizes, colors, shapes, and / or filling of the annotations.
[0207] Preprocessing or postprocessing
[0208] The above-mentioned dot detection using a MLM may be enhanced by additional processing steps. For example, the step of detecting dots further includes categorizing or filtering said one or more dots based on additional cellular or non-cellular markers.
[0209] Cellular markers are, for instance, nuclear, membranal or cytoplasmic markers or the like. The categorization can be performed using an MLM such as an Al based on a neural network or an algorithmic approach or the like. In case of using the MLM, it is possible to use the same MLM as the one for detecting dots, but it may be trained with additional labels (e.g., distinguishing in which cellular structure the dots are located). Labels may further include labels corresponding to properties of the dots such as dot sizes dot shapes, dot intensity, focal plane or the like.
[0210] The categorization classes can be, for example one or more classes distinguishing: whether or not a dot is located in a tumor region; whether or not a dot is located in a normal / stroma region; whether or not a dot is located in a cell nuclei; whether or not a dot is located on a membrane; and / or whether or not a dot is located in the vicinity of an immune cell.
[0211] In other words, the dots may be classified according to the surrounding biological structures.
[0212] In addition or alternatively, the step of detecting dots further comprises categorizing or filtering said one or more dots based on additional one or more images of the biologic sample including one or more images of the same slide or of consecutive slides.
[0213] For example, additional stains on consecutive slides (and / or on the same slide) may be used to generate the labels. In addition or alternatively, the dots may be filtered based on their z plane - for example, the dots that are blurred (out-of-focus) may be excluded from the set of detected (annotated) dots.
[0214] Training of the (second) MLM
[0215] Fig. 18 is an exemplary flow diagram that illustrates phases in tasks using an MLM or, more specifically, an Al such an Al based at least partially on a neural network architecture.
[0216] Step 1810 represents training data collection. In this stage, training images of stained biologic samples are collected together with the respective ground truth data. An exemplary ground truth data may be a list of coordinates of dots in a respective training image. However, the ground truth does not need to be list, it may be a 2D image in which merely the detected dot positions are marked or it may be directly the training image but with annotations added. The present disclosure is not limited to any particular format of the ground truth data.
[0217] The training data may be obtained as described above, by using an MLM to detect the dots and enabling a user to update the annotations (positions of the dots within the respective training image). The training data may be alternatively or in addition obtained purely by a human user marking the dots in the respective training image, or by another approach that enables marking the dots (such as any known algorithm specifically developed to detect dots (e.g., feature extraction, pattern matching, or the like).
[0218] Step 1820 represents a training phase. The training phase will be described in more detail below. In the training phase, the training images associated with their respective ground truth data are fed to the MLM so as to enable it to learn (train).
[0219] Step 1830 may be performed, but does not have to be. Step 1830 is a testing phase. In principle, the testing step 1830 corresponds to the inference phase 1840. However, the testing phase is performed for training data that has not been used in the training and enables evaluating the quality of the MLM by testing its output against the ground truth for the training images that were not used in the training phase.
[0220] Step 1840 represents an inference phase, i.e., the phase in which the MLM is applied to new data for which ground truth is unknown. In other words, the inference phase is the dot detection as described above, e.g., with reference to Fig. 11.
[0221] As can be seen in Fig. 18 and as already briefly discussed with reference to Fig. 13, the inference phase may be supplemented by a human user or another algorithm such as filtering or additional classification to modify the result of the detection by the MLM. Such modified result of detection may then be used as a new ground truth data for the training image and may be fed back to the training phase. Similar approach may be performed in the training phase.
[0222] Fig. 19 illustrates a training method 1900. The training method 1900 is a computerized method of training an Al model (or, in general, an MLM) for detecting objects in an image of a biological sample. The method comprises steps of: obtaining 1910 an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots; obtaining 1920 one or more annotations associated with the image of the biological sample and with the dots; and adapting 1930 one or more parameters of said Al model according to said image of a biological sample and, as a ground truth, said one or more annotations into the Al model.
[0223] Adapting the one or more parameters of the Al may include, for instance, modification of some weights and / or biases of a neural network. However, the present disclosure is not limited to this. The adapting may include changing the model itself or switching to a non-neural -network based machine learning method or the like.
[0224] It is noted that annotation here are not limited to the actually displayed annotations, but merely represent data (information) that is associated with the image. For example, such annotations are locations of the dots within the image. These locations may be in two dimensions or in three dimensions.
[0225] In order to further improve the training, the method may further comprise a step of obtaining an image of a negative control biological sample stained with an IHC or a fluorescence based approach; and a step of adapting one or more parameters of said Al modelaccording to said image of a negative control biological sample and, as a ground truth, no annotations or an annotation indicating no dots produced by the approach.
[0226] The negative controls may be stained with a regular qlHC (or a similar approach used also for the positive biologic samples). However, in order to obtain the negative control, the quantitative approach may be run with a standard protocol (e.g., qlHC), but without the dextran-HRP-ab polymer capable of generating a “real” dot - so all generated dots would be ghost dots. Specifically, the negative controls may be produced by replacing the reagent in the labelled secondary antibody (2230 in Fig. 12) step with buffer. Apart from that, everything is run as a positive reaction. That means that the HRP enzyme needed to catalyze the precipitation of the substrate in step 2260 of Fig. 12 is not present.
[0227] It is noted that the training may be performed in a device similar to dot detector 1200. In particular, in order to implement both, training 1820 and inference 1840, the processing circuitry 1210 may, in operation, perform steps described with reference to Figs. 11, 13, 18, and / or 19. This may be facilitating by the memory 1220 storing the corresponding program modules as shown schematically in Fig. 20.
[0228] Fig. 20 shows a memory portion 1222 of the memory 1220 with functional modules dedicated to dot detection (module 2010), user input processing (module 2020), GUI operation module (2030), and training (module 2040). It is note that not all modules must be present.
[0229] The present disclosure may provide merely a dot detector, in which case, only the dot detection module 2010 would be included. There may be no training module needed, as the MLM may be capable of only inference and may have the MLM parameter fixedly stored. The parameters may come from some preceding training, performed on a similar MLM implemented by a different device.
[0230] The device 1200 may implement a user input processing module 2020 that receives a user input via the operation input interface 1250, and determines an action to be taken upon that specific user input. In addition or alternatively, the device 1200 may implement a GUI operation module 2030 for representing a GUI (producing images representing the GUI) that are then provided over the display controller 1230 to a display device for displaying. This representing of the GUI may depend (receive input from) the input processing module 2020.
[0231] In addition or alternatively to the dot detection module 2010, the device 1200 may include a training module 2040 for performing the training as described above.
[0232] Fig. 21 illustrates the device 1200 described already with reference to Fig. 12, now as a part of a system that is capable of interacting with a human user. In particular, the device 1200 is connected over its display control interface to a display device 1260. The display device1260 may be a standalone screen external, and connectable and disconnectable from the device 1200 or it may be a display device permanently connected to and being a part of the device 1200. Such display devise may be any kind of display such as an OLED, LCD or the like.
[0233] Moreover, a user input device 1280 may be connected via the operation input interface 1250 of the device 1200. The user input device 1280 may be a keyboard including one or more keys such as a standard computer keyboard or the like or any kind or keyboard. Alternatively, or in addition, the input device 1280 may include a touch screen or a mouse or another means for pointing a cursor on various position for instance within a graphical user interface displayed on the display device 1260. The communication interface 1240 of the device 1200 may be configured to connect the device 1200 with a network 1295. As illustrated in Fig. 21, an image capturing device 1270 (such as a slide scanner or the like) that is a source of the first image and the second image may be also connected to the network 1295, so that the device 1200 may obtain the image(s) of a biologic sample directly from the image capturing device 1270. However, this way of obtaining the images is only exemplary. The images are not necessarily obtained directly from the image capturing device 1270. They may be obtained from an external storage over the network 1295. The communication interface 1240 may include or be an USB interface, so that the images may be obtained e.g., from an USB storage device or the like.
[0234] The images may be stored and / or obtained together with a set of corresponding annotations (e.g., as training data). Alternatively or in addition, the annotations may be determined and stored together with the image by the device 1200 as described above.
[0235] Exemplary embodiment based on qlHC
[0236] An exemplary embodiment may apply qlHC staining (generating dots). Such qlHC staining is illustrated in Fig. 22, (where an exemplary qlHC application is presented, namely dots that are expression of human epidermal growth factor receptor 2 (HER2). An example of a qlHC application can be seen in K. Jensen et al. “A novel quantitative immunohistochemistry method for precise protein measurements directly in formalin-fixed, paraffin-embedded specimens: analytical performance measuring HER2,” Mod. Pathol., vol. 30, issue 2, p. ISO- 193, Feb. 2017 (“Jensen 2017”).
[0237] Like in classic immunohistochemistry, the basis for the amplification is enzyme deposition —typically Horse radish peroxidase (HRP). In step 1 of Fig. 12, a primary antibody 2210 binds the target (protein / receptor) 2220 of interest. In step 2, a secondary, HRP -labeled antibody recognizes the primary antibody 2230, 2240 and binds. These first two steps in the qlHC reaction are directly comparable to standard immunohistochemistry. However, in qlHC,only a pre-determined fraction of the secondary antibodies (e.g., 2230) is labeled by HRP (2250). The secondary-labeled antibody 2230 is mixed with non-labeled antibody 2240 to increase robustness of the assay. In step 3, enzyme substrate 2260 is added and deposited. In step 4, HRP labeled antibody 2270 binds deposited substrate 2260. Finally, in step 5, an amplification reaction generates a dot 2280 centered around the labeled antibody, i.e., directly at the single target. The number of dots can be counted and as the ratio between labeled and unlabeled secondary antibody is known, the qlHC assay allows for a direct correlation between number of dots and amount of biomarker present in the tissue.
[0238] Classical computer vision algorithms may be disadvantageous for detecting qlHC dots, inter alia due to a highly localized nature of the dots in three dimensions (“3D”) as well as in two dimensions (2D). For example, out-of-focus dots may look much larger and diffused in comparison with in-focus dots, making it nigh-impossible to detect them using classical image processing methods. The classical image processing methods are disadvantaged in dealing with detecting individual dots which are clustered or seem clustered when viewed in a single 2D focal plane, although being separated in the third dimension. In order to overcome these issues, for example, 3D deconvolution might be used, which requires a z-stack 3D scanning of the tissue at inference time. While it is possible to manually mark dot locations to serve as the ground truth for training an Al-based model, without using a side-by-side view of flattened z-stack scans or selectable focal plan scans it is much harder to do so accurately, since qlHC dots are highly localized in 3D and so might look diffused and confusing when out-of- focus.
[0239] Marking of dot locations without the use of see-through annotation mode may be more difficult because the dots may often be quite small and a filled marker (such as a filled circle) or in general an overlapping marker (such as the cross 750) might hide the dots. Reducing false-positive detection may be achieved by manual marking of non-qlHC structures as non-dots. This has the disadvantage of being a very time-consuming process. However, such provision of negative annotations (that mark dot-like structures that are not qlHC dots or dots of the desired staining approach) may speed up the training and / or increase its accuracy when used as ground truth for training data.
[0240] Using z-stack (multiple focal planes) scanning and / or flattened z-stack or 3D deconvolved z-stack for improved generation of ground truth for model training may provide more ergonomy to the human users, including as an assistive layer for manual annotators marking dot locations An iterative annotation process can be provided, where at each iteration the previously trained model is used to pre-annotate qlHC dots on a scan and a human annotatorcorrects the pre-annotations, changing, adding or deleting. The corrected annotations are used to train the next iteration of the detection model. Using negative control stained tissues (stained with qlHC without the part connecting to the marker) may help to efficiently generate ground truth for reducing false positive detection. A trained model is applied to a negative control slide, and any dot detected by the model is assured to be a false-positive, as there are no actual dots in the slide.
[0241] User interface improvements for accurate and efficient marking dot locations may be provided as described above: A see-through annotation mode, enabling easy transition e.g. between annotations presented as (filled) dots and (unfilled) rectangles. The unfilled shape may provide some advantages, as filled markers hide the qlHC dot being marked. It may be useful for both ground truth annotation and for viewing model-inferred dot detections in a synchronized manner. Side-by-Side viewing of a panel showing the image being annotated and a panel showing selected z-stack focal planes or a flattened z-stack image may further improve the features of the user interface. Using the spatial relation of detected qlHC dots to expression of additional markers may be used to categorize or filter detected qlHC dots, including nuclear, membranal or cytoplasmic markers, as well as non-cellular markers.
[0242] Using in-focus analysis of z-stack scans, to localize detected qlHC dots in z-axis as well as in 2D, obtaining 3D localization of dots. Such 3D localization can enable more accurate assignment of detected dots to adjacent cells as indicated by additional markers. Advantages of approached described herein include improving the solution to the problem of detecting qlHC dots and localizing them in 2D and 3D, overcoming the difficulties of identifying dots due to confusing tissue morphology, changes in staining, dots clustering and out-of-focus dots. In addition, the present disclosure addresses the problem of efficiently and accurately generating annotated data for training an Al-based model for detecting qlHC dots, including the generation of annotated data for reducing false-positive detections.
[0243] More efficient and accurate data annotation process may be achieved thanks to an iterative process assisted by user interface improvements and z-stack based additional information making annotations more precise and easier to provide. Reducing false-positive detection using large amounts of efficiently generated false-positive annotations by applying trained model to negative controls.
[0244] The MLM may be an Al of any kind, or an ensemble thereof. For example, a model using deep neural networks such as a convolutional neural network may be used.
[0245] Methods for detection of qlHC dots may include training of an Al-based (e.g., neural network) qlHC dot detection model similar to those described in U.S. Patent No. 11,748,881.Therein, a fully convolutional network (“FCN”) and in particular the so-called U-net architecture is used, as suggested by O. Ronneberger et al. “U-Net: Convolutional Networks for Biomedical Image Segmentation”, 2015, arXiv: 1505.04597. The U-net has a contracting path and an expansive path, which gives it the u-shaped architecture. The contracting path is a typical convolutional network that consists of repeated application of convolutions, each followed by a rectified linear unit (“ReLU”) and a max pooling operation. During the contraction, the spatial information is reduced while feature information is increased. The expansive pathway combines the feature and spatial information through a sequence of up- convolutions and concatenations with high-resolution features from the contracting path. Convolutional neural networks may be applied to patches (of a predetermined and possibly configurable size) of the input image that may be, in some applications, processed in parallel. However, the present disclosure is not limited to such implementations of MLM.
[0246] Summary of embodiments relating to dot detection
[0247] According to a first aspect, a computerized method is provided of detecting objects in an image of a biological sample, the method comprising steps of: obtaining an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots; and detecting dots in said image of a biological sample using a trained Al model pretrained to detect dots.
[0248] According to a second aspect, in addition to the first aspect, the method further comprises a step of generating one or more annotations associated with the image of a biological sample and with the dots detected in the detecting step.
[0249] According to a third aspect, in addition to the first or second aspect, said one or more annotations indicate location of respective one or more dots detected in the detecting step.
[0250] According to a fourth aspect, in addition to the third aspect, the location of dots is a location in three spatial dimensions.
[0251] According to a fifth aspect, in addition to any of the second to fourth aspect, the method further comprises steps of providing an input interface capable of enabling a human user to obtain an updated one or more annotations by: i) deleting an annotation from the one or more annotations, ii) modifying an annotation among the one or more annotations, and / or iii) adding an annotation to the one or more annotations; and storing the updated one or more annotations.
[0252] According to a sixth aspect, in addition to the fifth aspect, the method further comprises a step of using the updated one or more annotations as a ground truth to train said Al model or another Al model for detecting dots.
[0253] According to a seventh aspect, in addition to any of the second to sixth aspect, the method further comprises a step of providing an annotation viewing interface configured to display said image of the biological sample together with said one or more annotations.
[0254] According to an eighth aspect, in addition to the seventh aspect, the annotation viewing interface is a side-by-side viewing interface configured to display a first view showing a first-view image based on said image of the biological sample and including said one or more annotations overlaid beside a second view showing a second-view image of the biological sample.
[0255] According to a ninth aspect, in addition to the eighth aspect, said second-view is captured with settings different from those of said image of the biological sample and includes annotations corresponding to said one or more annotations of the first-view image.
[0256] According to a tenth aspect, in addition to the eighth or ninth aspect, said second- view image of the biological sample is based on or includes a z-stack including two or more mutually different focal plane images of the same field of view, FOV, of the biological sample.
[0257] According to a eleventh aspect, in addition to any of the eight to tenth aspect, said second-view image is obtained by flattening or 3D deconvolution of the z-stack including combining the two or more focal plane images into a single image of the biological sample.
[0258] According to a twelfth aspect, in addition to any of the seventh to eleventh aspect, the annotation viewing interface is configured to view each of said one or more annotations as a graphics of an unfilled contour of a preconfigured shape surrounding one of the detected dots.
[0259] According to a thirteenth aspect, in addition to the twelfth aspect, the preconfigured shape is a rectangle or a square positioned so that the one of the detected dots is in its geometric center.
[0260] According to a fourteenth aspect, in addition to any of the seventh to thirteenth aspect, the annotation viewing interface is configured to enable switching between focal planes of the z-stack.
[0261] According to a fifteenth aspect, in addition to any of the seventh to fourteenth aspect, the annotation viewing interface is configured to: i) enable a user to select a focal plane out of the z-stack and to mark a dot in the selected focal plane, and ii) store an identification of said selected focal place in association with the marked dot.
[0262] According to a sixteenth aspect, in addition to any of the second to seventh to fifteenth aspect, the annotation viewing interface is configured to enable switching between: i) viewing each of said annotations as a graphics of a dot co-located with one of the detected dots, and ii)said viewing each of said annotations as a graphics of an unfilled contour of a preconfigured shape surrounding one of the detected dots.
[0263] According to a seventeenth aspect, in addition to the sixteenth aspect, said switching is triggered by changing a field of view of said image of the biologic sample.
[0264] According to an eighteenth aspect, in addition to any of the first to seventeenth aspect, the step of detecting dots further includes categorizing or filtering said one or more dots based on additional cellular or non-cellular markers.
[0265] According to a nineteenth aspect, in addition to any of the first to eighteenth aspect, the step of detecting dots further comprises categorizing or filtering said one or more dots based on additional one or more images of the biologic sample including one or more images of the same slide or of consecutive slides.
[0266] According to a twentieth aspect, a computerized method is provided of training an Al model for detecting objects in an image of a biological sample, the method comprising steps of: i) obtaining an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots; ii) obtaining one or more annotations associated with the image of the biological sample and with the dots; and iii) adapting one or more parameters of said Al model according to said image of a biological sample and, as a ground truth, said one or more annotations into the Al model.
[0267] According to a 21-st aspect, in addition to the twentieth aspect, the method further comprises obtaining an image of a negative control biological sample stained with an IHC or fluorescence-based approach; and adapting one or more parameters of said Al model according to said image of a negative control biological sample and, as a ground truth, no annotations or an annotation indicating no dots.
[0268] According to a 22-nd aspect, in addition to twentieth or 21 -st aspect, said one or more annotations indicate location of respective one or more dots detected in the detecting step.
[0269] According to a 23-rd aspect, in addition to twentieth or 22-nd aspect, the location of dots is a location in three spatial dimensions.
[0270] According to a 24-th aspect, a computerized method is provided of training a first Al model for detecting objects in an image of a biological sample, the method comprising steps of: i) obtaining an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots; ii) obtaining one or more annotations associated with the image of the biological sample and with the dots, wherein said obtaining includes: a) obtaining an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots, and b) detecting dots in said image of a biologicalsample using a trained second Al model pre-trained to detect dots; and iii) adapting one or more parameters of said first Al model according to said image of a biological sample and, as a ground truth, said one or more annotations into the first Al model.
[0271] According to a 25-th aspect, a computer program is provided, stored on a non- transitory medium and including instructions which when executed on one or more processors causes the one or more processors to perform the steps of the method according to any of the first to 24-th aspect.
[0272] According to a 26-th aspect, a detection device is provided comprising: a data interface; a storage; and processing circuitry that, in operation, i) obtains over said data interface an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots, ii) detects dots in said image of a biological sample using a trained Al model pre-trained to detect dots, and iii) stores indication of locations of said respective detected dots as one or more annotations.
[0273] According to a 27-th aspect, in addition to twentieth or 26-th aspect, the detection device further comprises an input interface capable of enabling a human user to update one or more annotations by: i) deleting an annotation from the one or more annotations, ii) modifying an annotation among the one or more annotations, and / or iii) adding an annotation to the one or more annotations; wherein the processing circuitry, in operation, stores the updated one or more annotations in the storage.
[0274] According to a 28-th aspect, in addition to 26-th or 27-th aspect, the detection device further comprises an annotation viewing interface configured to display said image of the biological sample together with said one or more annotations.
[0275] According to a 29-th aspect, a training device is provided, comprising: a data interface; a storage; and processing circuitry that, in operation: i) obtains an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots via the data interface; ii) obtains one or more annotations associated with the image of the biological sample and with the dots via the data interface; iii) adapts one or more parameters of an Al model according to said image of a biological sample and, as a ground truth, said one or more annotations, thereby training the Al model for detecting objects in an image of a biological sample.
[0276] It is noted that the present disclosure also provides devices of which the processing circuitry performs any of the methods described herein. The present disclosure provides an integrated circuit which embodies the processing circuitry as described above.
[0277] Experimental results
[0278] Fig. 23 shows a comparison of the number of qlHC dots as an estimate of HER2 expression to the manual scoring based on membranal HER2 IHC staining. In particular, Fig. 23 shows correlation between dots' count and manual estimation of the GE051 staining. Specifically, here the tissue is scored only in a selected a rectangular region (out of the entire WSI region).
[0279] The dash line is a linear fit to the data. Points in the figure represent different cases where a manual score was given on one image and the dots were counted on the second (consecutive) slide. The qlHC dots were counted in a many (could be also overlapping) circles with defined radius (here r=256 pixels) within the defined region, so for each rectangular region many counts are available and a standard deviation can thus be provided for each case (as can be seen by the vertical lines accompanying the points). As can be seen from the figure, number of qlHC dots is a suitable measure for the HER2 antibody expression. Correspondingly, Fig 4. presents a computational pipeline where the actual HER2 expression is derived from the qlHC dot count (the GT). For the purpose of Fig. 23, the HER2 expression was scored manually and compared to the dot counts annotated by the MLM.
[0280] Fig. 24 shows a comparison of the Al model-based estimate of HER2 expression as described with reference to Fig. 4 (x axis shows the inference result of the MLM 470) with the actual qlHC dot count (y axis) summarized for the whole tumor region in each slide.
[0281] Figure 25 illustrates results of the qlHC Al trained model prediction. In (a), qlHC dots per patch distribution for model training dataset are shown. Dashed vertical lines represent thresholds for counts of percentiles in 10% binning. Part (b) of Fig. 25 shows a confusion matrix of true label versus predicted label (predicted by the trained Al model for patches of test dataset). As can be seen in the figure, the correlation between the predicted and true values is considerable. Part (c) of Fig. 25 shows true label versus the predicted label (predicted by the trained Al model for patches of test dataset averaged per region in a slide). Part (d) of Fig. 25 shows the true label versus the predicted label (predicted by the trained Al model for patches of test dataset averaged per slide). Labels in (b)-(d) represent percentiles binning of the counts calculated form the distribution in (a) red dashed lines. Dots sizes in (c) and (d) represent the number of nuclei per region and slide, respectively. As can be seen, the correlation coefficient for these graph reaches correlation coefficient values above 0.8, so that the model is closely following the ground truth.Examples
[0282] Example 1. Evaluation of a qlHC detection model on formalin-fixed paraffin- embedded cell lines.
[0283] An experimental study was conducted to evaluate an exemplary qlHC detection model based on the present disclosure, with various formalin-fixed paraffin-embedded cell lines. This study demonstrated a linear relationship between expression level of HER2 and the qlHC dot count with concentration of qlHC chemical substrate. As illustrated by the results of this study, qlHC can be used as a tunable system, where the qlHC dot count is directly proportional to the concentration to the dot-generating component. This allows for quantitative HER2 detection with a high signal / noise ratio at different protein expression levels.
[0284] Methods and Results
[0285] Samples from several formalin-fixed paraffin-embedded cell lines (MDA-MB-468, MDA-MB-231, MDA-MB-175, MDA-MB-453, and SK-BR-3) with different levels of HER2 expression (classified in a range from “0” to 3+”) were stained with IHC using a HercepTest™ mAb (Dako Omnis) or qlHC (FIG. 26). The sample preparation and staining protocol used for this study is described in Jensen 2017.
[0286] The number of qlHC dots per cell line were counted and the mean result for each cell line is shown in the graph provided as FIG. 27, with shading around each mean to show the 95% confidence interval. FIG. 28 provides representative images of qlHC-stained cells examined in this study, with the detected qlHC dots annotated as points (left) or with boxes (right). As illustrated by FIG. 27, a linear relationship exists between the mean dots / cell count and the level of HER2 expression in these cell lines and, as such, the present methods provide reliable and consistent quantitative expression information in FFPE cell lines that have different HER2 expression levels.* * *
[0287] It is noted that the term “first” (e.g., in the terms “first ML model”, “first set of annotations”, or the like) here is used merely for labeling. It does not have any quantitative or qualitative implications, e.g., it does not suggest any kind of order of application or any specific feature of the model itself. The same applies for the term “second”, “third”, “fourth” image, set of annotations etc.
[0288] While certain features and aspects have been described with respect to exemplary embodiments, one skilled in the art will recognize that numerous modifications are possible. For example, the methods and processes described herein may be implemented using hardware components, software components, and / or any combination thereof. Further, while various methods and processes described herein may be described with respect to particular structural and / or functional components for ease of description, methods provided by various embodiments are not limited to any particular structural and / or functional architecture butinstead can be implemented on any suitable hardware, firmware and / or software configuration. Similarly, while certain functionality is ascribed to certain system components, unless the context dictates otherwise, this functionality can be distributed among various other system components in accordance with the several embodiments.
[0289] Moreover, while the procedures of the methods and processes described herein are described in a particular order for ease of description, unless the context dictates otherwise, various procedures may be reordered, added, and / or omitted in accordance with various embodiments. Moreover, the procedures described with respect to one method or process may be incorporated within other described methods or processes; likewise, system components described according to a particular structural architecture and / or with respect to one system may be organized in alternative structural architectures and / or incorporated within other described systems. Hence, while various embodiments are described with or without certain features for ease of description and to illustrate exemplary aspects of those embodiments, the various components and / or features described herein with respect to a particular embodiment can be substituted, added and / or subtracted from among other described embodiments, unless the context dictates otherwise. Consequently, although several exemplary embodiments are described above, it will be appreciated that the disclosure is intended to cover all modifications and equivalents within the scope of the following claims.
[0290] The descriptions of the various embodiments of the present disclosure have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.
[0291] It is expected that during the life of a patent maturing from this application many relevant machine learning models will be developed and the scope of the term machine learning model is intended to include all such new technologies a priori.
[0292] As used herein the term “about” refers to + / - 10 %.
[0293] The terms “comprises”, “comprising”, “includes”, “including”, “having” and their conjugates mean “including but not limited to”. This term encompasses the terms “consisting of’ and “consisting essentially of’.
[0294] The phrase “consisting essentially of’ means that the composition or method may include additional ingredients and / or steps, but only if the additional ingredients and / or steps do not materially alter the basic and novel characteristics of the claimed composition or method.
[0295] As used herein, the singular form “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a compound” or “at least one compound” may include a plurality of compounds, including mixtures thereof.
[0296] The word “exemplary” is used herein to mean “serving as an example, instance or illustration”. Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments and / or to exclude the incorporation of features from other embodiments.
[0297] The word “optionally” is used herein to mean “is provided in some embodiments and not provided in other embodiments”. Any particular embodiment may include a plurality of “optional” features unless such features conflict.
[0298] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the various embodiments described and / or claimed herein. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0299] Whenever a numerical range is indicated herein, it is meant to include any cited numeral (fractional or integral) within the indicated range. The phrases “ranging / ranges between” a first indicate number and a second indicate number and “ranging / ranges from” a first indicate number “to” a second indicate number are used herein interchangeably and are meant to include the first and second indicated numbers and all the fractional and integral numerals therebetween.
[0300] It is appreciated that certain features of the embodiments described and / or claimed herein, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination or as suitable in any other described embodiment. Certainfeatures described in the context of various embodiments are not to be considered essential features of those embodiments, unless the embodiment is inoperative without those elements.
[0301] Although the disclosure has been described with reference to specific embodiments, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations thereof that fall within the spirit and broad scope of the present disclosure and claims.
[0302] It is the intent of the applicant(s) that all publications, patents and patent applications referred to in this specification are to be incorporated in their entirety by reference into the specification, as if each individual publication, patent or patent application was specifically and individually noted when referenced that it is to be incorporated herein by reference. In addition, a citation or identification of any reference in this application shall not be construed as an admission that such reference is available as prior art to the present application or any patent resulting therefrom. To the extent that section headings are used, they should not be construed as necessarily limiting. In addition, any priority document(s) of this application is / are hereby incorporated herein by reference in its / their entirety.
Claims
CLAIMS1. A method of generating training data, the method comprising: obtaining a first image of a biological sample; generating a first set of annotations of the first image based on a second image of said biological sample stained by a quantitative approach converting antibody / antigen complexes into dots; and outputting the first image and a ground truth including said first set of annotations as data for training a first machine learning (“ML”) model for analyzing a biological sample.
2. The method of claim 1, wherein the generation of the first set of annotations is based on a statistical measure of the dots in the second image.
3. The method of claim 2, wherein the statistical measure includes a number of dots and / or a density of dots.
4. The method of claims 2 or 3, wherein generating the first set of annotations comprises generating an annotation for each region out of a plurality of regions of the first image by determining the statistical measure for a matching region in the second image matching said region in the first image.
5. The method of claims 2 or 3, wherein generating the first set of annotations comprises: obtaining a plurality of regions in the second image; for each of the plurality of regions in the second image: computing the statistical measure, and associating the computed statistical measure as an annotation out of the first set of annotations with a location, in the first image, of a region matching said region in the second image.
6. The method of claims 4 or 5, wherein each region is a rectangle, a square, or a circle at a previously configured position in the first image.
7. The method of any one of claims 4 to 6, wherein said first set of annotations includes said statistical measure associated with a location of said respective region within the first image and / or with a location of said respective matching region on the second image.
8. The method of any one of claims 4 to 7, wherein at least one region of the plurality of regions is a closed area selected in the first image.
9. The method of any one of claims 1 to 8, wherein said generating the first set of annotations comprises inputting the second image into a second processing; and obtaining, as the output of the second processing, said first set of annotations.
10. The method of any one of claims 1 to 9, wherein said first image is an image of the biological sample stained with immunohistochemistry (“IHC”), hematoxylin and eosin (“H&E”), and / or fluorescence in situ hybridization (“FISH”).
11. The method of any one of claims 1 to 10, further comprising: obtaining the second image stained with a quantitative approach converting antibody / antigen complexes into dots, using a first concentration of the antibody / antigen; obtaining a third image of the biological sample stained with a quantitative approach converting antibody / antigen complexes into dots, with a second concentration different form the first concentration, wherein the second image and the third image are images of respective slides with consecutive sections of the biological sample; and generating a second set of annotations of the first image based on the third image; and outputting the first image and a ground truth including said second set of annotations as data for training the first ML model for analyzing a biological sample.
12. The method of any one of claims 1 to 11, wherein the staining approach is an IHC staining to visualize HER2 epitope in human formalin fixed, paraffin embedded tissue.
13. A method of training a machine learning (“ML”) model, the method comprising: inputting, to the ML model the first image and the ground truth output by the method of generating training data according to any one of claims 1 to 12; andmodifying at least one parameter of the machine learning model based on the input first image and the ground truth.
14. The method of claim 13, comprising one or more iterations of performing the method of generating training data according to any one of claims 1 to 12 and said steps of inputting and modifying.
15. A method of analysis, comprising: inputting a fourth image of the biological sample, or a first image of a second biological sample, to an ML model trained using the method of claim 11; and obtaining, from the output of the ML model, an analysis of the biological sample or the second biological sample.
16. The method of claim 15, wherein the analysis result indicates a presence, absence, or amount of an antigen in the biological sample or the second biological sample.
17. The method of claims 15 or 16, wherein the analysis is a classification of the biological sample or the second biological sample.
18. The method of any one of claims 1 to 17, wherein said ML model is a first ML model and said generating the first set of annotations includes: detecting objects in an image of a biological sample, the detecting comprising: obtaining an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots, and detecting dots in said image of a biological sample using a trained second ML model pre-trained to detect dots; inputting the second image into a second processing; and obtaining, as the output of the second processing, said first set of annotations.
19. The method of claim 18, wherein the detecting further comprises generating one or more annotations associated with the image of a biological sample and with the detected dots.
20. The method of claims 18 or 19, wherein said one or more annotations indicate a location of respective one or more of the detected dots.
21. The method of claim 20, wherein the location of dots is a location in three spatial dimensions.
22. The method of any one of claims 19 to 21, wherein the detecting further comprises: providing an input interface capable of enabling a human user to obtain an updated one or more annotations by: i) deleting an annotation from the one or more annotations, ii) modifying an annotation among the one or more annotations, and / or iii) adding an annotation to the one or more annotations; and storing the updated one or more annotations.
23. The method of claim 22, further comprising using the updated one or more annotations as a ground truth to train said second ML model or another Al model for detecting dots.
24. The method of any one of claims 19 to 23, further comprising providing an annotation viewing interface configured to display said image of the biological sample together with said one or more annotations.
25. The method of claim 24, wherein the annotation viewing interface is a side-by-side viewing interface configured to display a first view showing a first-view image based on said image of the biological sample and including said one or more annotations overlaid beside a second view showing a second-view image of the biological sample.
26. The method of claim 25, wherein said second-view is captured with settings different from those of said image of the biological sample and includes annotations corresponding to said one or more annotations of the first-view image.
27. The method of claims 25 or 26, wherein said second-view image of the biological sample is based on or includes a z-stack including two or more mutually different focal plane images of the same field of view, FOV, of the biological sample.
28. The method of claim 19, whereinsaid second-view image is obtained by flattening or 3D deconvolution of the z-stack including combining the two or more focal plane images into a single image of the biological sample.
29. The method of any one of claims 24 to 28, wherein the annotation viewing interface is configured to view each of said one or more annotations as a graphics of an unfilled contour of a preconfigured shape surrounding one of the detected dots.
30. The method of claim 29, wherein the annotation viewing interface is configured to enable switching between focal planes of the z-stack.
31. The method of claim 29 or 30, wherein the annotation viewing interface is configured to: enable a user to select a focal plane out of the z-stack and to mark a dot in the selected focal plane, and store an identification of said selected focal place in association with the marked dot.
32. The method of any one of claims 29 to 31, wherein the shape is a rectangle or a square positioned such that the one of the detected dots is in its geometric center.
33. The method of any one of claims 29 to 32, wherein the annotation viewing interface is configured to enable switching between: i) viewing each of said annotations as a graphics of a dot co-located with one of the detected dots, and ii) said viewing each of said annotations as a graphics of an unfilled contour of a preconfigured shape surrounding one of the detected dots.
34. The method of claim 33, wherein said switching is triggered by changing a field of view of said image of the biologic sample.
35. The method of any one of claims 18 to 34, wherein the detection of the dots further includes categorizing or filtering said one or more dots based on additional cellular or non- cellular markers.
36. The method of any one of claims 1 to 35, wherein the detection of the dots further comprises categorizing or filtering said one or more dots based on additional one or more images of the biologic sample including one or more images of the same slide or of consecutive slides.
37. The method of any one of claims 18 to 36, further comprising training said second ML model for detecting objects in an image of a biological sample, the training comprising: obtaining an image of a biological sample stained with a quantitative approach converting antibody / antigen complexes into dots; obtaining one or more annotations associated with the image of the biological sample and with the dots; adapting one or more parameters of said second ML model according to said image of a biological sample and, as a ground truth, said one or more annotations into the Al model.
38. The method of claim 37, wherein the training further comprises: obtaining an image of a negative control biological sample stained with an IHC or fluorescence based approach; adapting one or more parameters of said Al model according to said image of a negative control biological sample and, as a ground truth, no annotations or an annotation indicating no dots.
39. A computer program stored on a non-transitory medium and including instructions which when executed on one or more processors causes the one or more processors to perform the steps of the method of any one of claims 1 to 38.
40. A training data generating device, comprising: a data interface; and processing circuitry that, in operation, obtains a first image of a biological sample via the data interface; generates a first set of annotations of the first image based on a second image of said biological sample stained by a quantitative approach converting antibody / antigen complexes into dots; andoutputs the first image and a ground truth including said first set of annotations as data for training a first machine learning (“ML”) model, for analyzing a biological sample.
41. The method of any one of claims 19 to 21, wherein the method further comprises determining a number of dots per cell in one or more cells in the biological sample or the second biological sample, using the first ML model or the second ML model, and determining or estimating a level of antibody expression in the biological sample or the second biological sample based on the determined number of dots per cell.
42. The method of claim 41, wherein the level of antibody expression in the biological sample or the second biological sample is determined or estimated based on a mean number of dots per cell.
43. The method of claim 41 or claim 42, wherein the biological sample or the second biological sample comprises formalin-fixed and paraffin-embedded cells.