Cell classification using center emphasis of feature maps

The method employs a neural network to generate feature maps and concentric crops from digital pathology images, improving the classification of biological structures by emphasizing central image features and reducing contextual influence.

JP2025514820APending Publication Date: 2025-05-09GENENTECH INC +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024562205
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-04-22
Filing Date
2023-03-31
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Current digital pathology methods struggle to accurately classify biological structures, particularly when the region of interest is not precisely centered within the image, leading to potential misclassification due to surrounding contextual information.

Method used

A computer-implemented method using a trained neural network with convolutional layers generates a feature map, creates multiple concentric crops, and processes these crops to produce output vectors representing the central region's characteristics, thereby enhancing classification accuracy.

Benefits of technology

This approach improves the accuracy of classifying biological structures by effectively weighting central image information more heavily, reducing the influence of surrounding areas, and enhancing the robustness of the classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025514820000001_ABST
    Figure 2025514820000001_ABST
Patent Text Reader

Abstract

Techniques described herein include, for example, generating a feature map for an input image, generating multiple concentric crops of the feature map, and generating an output vector that represents characteristics of structures depicted in a central region of the input image using the multiple concentric crops. Generating the output vector may include, for example, aggregating the set of output features generated from the multiple concentric crops, and several methods of aggregating are described. Applications to classification of structures depicted in central regions of input images are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 333,924, filed April 22, 2022, which is incorporated by reference herein in its entirety for all purposes. [Background technology]

[0002] Digital pathology may involve the interpretation of digitized images to accurately diagnose a subject's disease and guide therapeutic decision-making. In a digital pathology solution, an image analysis workflow may be established to automatically detect or classify biological objects of interest (e.g., tumor cells that are positive or negative for a particular biomarker or other indicator, etc.). A typical digital pathology solution workflow includes obtaining a slide of a tissue sample, scanning a preselected area or the entire tissue slide with a digital image scanner (e.g., a whole slide imaging (WSI) scanner) to obtain a digital image, and performing image analysis on the digital image using one or more image analysis algorithms (e.g., to detect objects of interest). Such a workflow may also include quantifying the objects of interest based on the image analysis (e.g., counting or identifying object-specific or cumulative areas of the objects), and may further include quantitative or semi-quantitative scoring of the sample (e.g., positive, negative, moderate, weak, etc.) based on the results of the quantification. Summary of the Invention

[0003] In various embodiments, a computer-implemented method for classifying an input image is provided that includes generating a feature map for the input image using a trained neural network including at least one convolutional layer, generating a plurality of concentric crops of the feature map, using information from each of the plurality of concentric crops to generate an output vector representative of characteristics of structures depicted in a central region of the input image, and determining a classification result by processing the output vector.

[0004] In some embodiments, generating the output vector using the multiple concentric crops includes, for each of the multiple concentric crops, generating a corresponding one of the multiple feature vectors using at least one pooling operation, and generating the output vector using the multiple feature vectors. In such methods, the multiple feature vectors may be ordered by radial size of the corresponding concentric crop, and generating the output vector may include separately convolving the trained filter over adjacent pairs of the ordered multiple feature vectors.

[0005] In some embodiments, generating an output vector using the plurality of concentric crops includes, for each of the plurality of concentric crops, generating a corresponding one of the plurality of feature vectors using at least one pooling operation, and generating an output vector using the plurality of feature vectors. In such methods, the plurality of feature vectors may be ordered by radial size of the corresponding concentric crop, and generating an output vector may include convolving a filter trained across a first adjacent pair of the ordered plurality of feature vectors to generate a first combined feature vector, and convolving a filter trained across a second adjacent pair of the ordered plurality of feature vectors to generate a second combined feature vector. In such methods, generating an output vector using the plurality of feature vectors may include generating an output vector using the first combined feature vector and the second combined feature vector.

[0006] In various embodiments, a computer-implemented method for classifying an input image is provided that includes generating a feature map for the input image, generating a plurality of feature vectors using information from the feature map, generating a second plurality of feature vectors using a trained shared model applied separately to each of the plurality of feature vectors, generating an output vector representing characteristics of structures depicted in the input image using information from each of the second plurality of feature vectors, and determining a classification result by processing the output vector.

[0007] In various embodiments, a computer-implemented method for training a classification model including a first neural network and a second neural network is provided, the computer-implemented method including generating a plurality of feature maps using the first neural network and information from images of a first dataset, and training a second neural network using the plurality of feature maps, where each image of the first dataset depicts at least one biological cell, and the first neural network is pre-trained on a plurality of images of a second dataset including images that do not depict biological cells.

[0008] In various embodiments, a computer-implemented method for classifying an input image is provided, the computer-implemented method including: generating a feature map for the input image using a first trained neural network of a classification model; generating an output vector representative of properties of structures depicted in a central portion of the input image using a second trained network of the classification model and information from the feature map; and determining a classification result by processing the output vector, wherein the input image depicts at least one biological cell, the first trained neural network is pre-trained on a first plurality of images including images that do not depict a biological cell, and the second trained neural network is trained by providing the classification model with a second plurality of images depicting a biological cell.

[0009] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0010] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0011] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0012] The terms and expressions used are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the features or portions thereof shown and described, recognizing that various modifications are possible within the scope of the claimed invention. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it should be understood that modifications and variations of the concepts disclosed herein may be reclassified by those skilled in the art, and such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims. [Brief description of the drawings]

[0013] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0014] BRIEF DESCRIPTION OF THE DRAWINGS Aspects and features of the various embodiments will become more apparent from the following detailed description of the invention, taken in conjunction with the accompanying drawings.

[0015] [Figure 1A] FIG. 1A illustrates an example of an image processing pipeline 100 according to some embodiments. [Figure 1B] FIG. 1B shows an example of an input image with a candidate (a structure predicted to be of a particular class) at the center. [Figure 2A] FIG. 2A shows another example of an input image with a candidate (denoted by a red circle) at the center. [Figure 2B] FIG. 2B shows an example of an input image where the candidate (denoted by the red circle) is not in the center of the image. [Figure 2C] FIG. 2C shows an example of an input image in which a candidate (denoted by a red circle) is surrounded by several structures of another class. [Diagram 3] FIG. 3 shows six different examples of importance masks. [Figure 4A] FIG. 4A illustrates a flow diagram of an exemplary process for classifying an input image, according to some embodiments. [Figure 4B] FIG. 4B illustrates a block diagram of exemplary architecture components for classifying an input image, according to some embodiments. [Figure 5A] FIG. 5A illustrates an example of an operation in which a feature map is generated for an input image, according to some embodiments. [Figure 5B] FIG. 5B illustrates an example of an operation in which multiple concentric crops of a feature map are generated, according to some embodiments. [Figure 6] FIG. 6 illustrates another example of a feature map for an input image and an operation in which multiple concentric crops of the feature map are generated, according to some embodiments. [Figure 7]FIG. 7 illustrates an example of operations in which a corresponding downsampling operation is performed on each of multiple concentric crops to generate a corresponding downsampled feature map, a combining operation is performed on the downsampled feature maps to generate an output vector, and the output vector is processed to generate a classification result, according to some embodiments. [Figure 8] FIG. 8 shows an example of a four-class structure. [Figure 9] FIG. 9 illustrates an example of a feature map, corresponding concentric crops, and corresponding feature vectors according to some embodiments. [Figure 10] FIG. 10 illustrates the example of FIG. 9 further including an output vector according to some embodiments. [Figure 11A] FIG. 11A illustrates a flowchart of another exemplary process for classifying an input image, according to some embodiments. [Figure 11B] FIG. 11B illustrates a block diagram of another exemplary architecture component for classifying an input image according to some embodiments. [Figure 12] FIG. 12 illustrates the example of FIG. 9 further including a shared model and an output vector according to some embodiments. [Figure 13] FIG. 13 shows the example of FIG. 9 further including a convolution over the radius and output vectors according to some embodiments. [Figure 14A] FIG. 14A illustrates an example operation in which an output vector is generated using multiple concentric crops of a feature map, according to some embodiments. [Figure 14B] FIG. 14B illustrates a block diagram of an implementation of an output vector generation module according to some embodiments. [Figure 15A] FIG. 15A illustrates a block diagram of an implementation of a combination module according to some embodiments. [Figure 15B] FIG. 15B illustrates a block diagram of an implementation of an additional feature vector generation module according to some embodiments. [Figure 16A]FIG. 16A illustrates a block diagram of an implementation of architecture components according to some embodiments. [Figure 16B] FIG. 16B shows a block diagram of an implementation of a feature vector generation module according to some embodiments. [Figure 17A] FIG. 17A shows a flowchart of an exemplary process for training a classification model including a first neural network and a second neural network, according to some embodiments. [Figure 17B] FIG. 17B illustrates a flow diagram of an exemplary process for classifying an input image, according to some embodiments. [Figure 18A] FIG. 18A illustrates a flowchart of an exemplary process for classifying an input image, according to some embodiments. [Figure 18B] FIG. 18B illustrates a flow diagram of an exemplary process for classifying an input image, according to some embodiments. [Figure 19A] FIG. 19A shows a schematic diagram of an encoder-decoder network according to some embodiments. [Figure 19B] FIG. 19B illustrates an example of an operation in which a feature map is generated for an input image using a trained encoder, according to some embodiments. [Figure 20] FIG. 20 shows examples of input images and corresponding reconstructed images generated by an encoder-decoder network (also called an "autoencoder" network) according to some embodiments at different epochs of training. [Figure 21] FIG. 21 shows examples of input images and corresponding reconstructed images produced by various encoder-decoder networks according to some embodiments. [Figure 22A] FIG. 22A illustrates an example of an operation in which a feature map is generated for an input image using a trained encoder and a trained neural network, according to some embodiments. [Figure 22B]FIG. 22B illustrates an example of an operation in which a feature map is generated for an input image using multiple feature maps and a combining operation, according to some embodiments. [Diagram 23] FIG. 23 illustrates an example of an operation in which an output vector is generated using multiple concentric crops of a feature map, according to some embodiments. [Figure 24] FIG. 24 illustrates an example of an operation in which an output vector is generated using a feature map, according to some embodiments. [Figure 25A] FIG. 25A illustrates an example of an operation in which a feature map is generated for an input image using a trained neural network, according to some embodiments. [Figure 25B] FIG. 25B illustrates an example of an operation in which a feature map is generated for an input image using multiple feature maps and a combining operation, according to some embodiments. [Figure 26] FIG. 26 illustrates an example of an operation in which an output vector is generated using multiple concentric crops of a feature map, according to some embodiments. [Figure 27] FIG. 27 illustrates an example of an operation in which a feature map is generated for an input image, according to some embodiments. [Figure 28] FIG. 28 illustrates an example of an operation in which an output vector is generated using multiple concentric crops of a feature map, according to some embodiments. [Figure 29] FIG. 29 illustrates an example of an operation in which a feature map is generated for an input image, according to some embodiments. [Diagram 30] FIG. 30 illustrates an example of an operation in which an output vector is generated using multiple concentric crops of a feature map, according to some embodiments. [Diagram 31] FIG. 31 illustrates an example of a computing system according to some embodiments that may be configured to perform the methods described herein. [Diagram 32] FIG. 32 illustrates an example of a precision-recall curve for a model described herein, according to some embodiments.

[0016] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dash and a second label that distinguishes between the similar components. If only a first reference label is used herein, the description is applicable to any of the similar components having the same first reference label, regardless of the second reference label. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0017] Techniques described herein include, for example, generating a feature map for an input image, generating multiple concentric crops of the feature map, and generating an output vector that represents characteristics of structures depicted in a central region of the input image using the multiple concentric crops. Generating the output vector may include, for example, aggregating a set of output features generated from the multiple concentric crops, and several methods of aggregating are described. For example, aggregating a set of output features generated from the multiple concentric crops may be performed using "radial convolution," which is defined as a convolution and / or convolution over feature maps or vectors derived from concentric crops of feature maps of multiple radii, for each of the multiple concentric crops of the feature map of different radii, along with a corresponding feature map or vector derived from the crop. Applications to classification of structures depicted in a central region of an input image are also described.

[0018] One or more such techniques may be applied, for example, to convolutional neural networks suitable for image classification applications. Technical advantages of such techniques may include, for example, improved image processing (e.g., higher true positives / negatives / detections, lower false positives / negatives), improved prognostic assessment, improved diagnostic enhancement, and / or improved treatment recommendations. Exemplary use cases presented herein include mitosis detection and classification from images of stained tissue samples (e.g., hematoxylin-eosin (H&E) stained tissue samples). 1. Background

[0019] Tissue samples (e.g., tumor samples) may be fixed and / or embedded using a fixative (e.g., a liquid fixative such as a formaldehyde solution) and / or an embedding substance (e.g., a histological wax such as paraffin wax and / or one or more resins such as styrene or polyethylene). Fixed tissue samples may also be dehydrated (e.g., via exposure to an ethanol solution and / or a clearing intermediate agent) prior to embedding. The embedding substance may penetrate the tissue sample in a liquid state (e.g., upon heating). Fixed, dehydrated, and / or embedded tissue samples may be sliced ​​to obtain a series of sections, each section having a thickness of, for example, 4-5 microns. Such sectioning may be performed by first cooling the sample and slicing the sample in a warm water bath. Deparaffinization (e.g., using xylene) and / or rehydration (e.g., using ethanol and water) of the slices may be performed prior to staining of the slices and mounting on glass slides (e.g., as described below).

[0020] Because tissue sections and the cells therein are largely transparent, slide preparation typically involves staining (e.g., automatically staining) the tissue sections to make relevant structures more visible. For example, different sections of tissue may be stained with one or more different stains to represent different characteristics of the tissue. For example, each section may be exposed to a predetermined amount of stain for a predetermined period of time. If the sections are exposed to multiple stains, the sections may be exposed to the multiple stains simultaneously.

[0021] Each section is mounted on a slide and then scanned to create a digital image that can then be examined by digital pathology image analysis and / or interpreted by a pathologist (e.g., using image viewer software). The pathologist may review and manually annotate the digital images of the slide (e.g., tumor areas, necrosis, etc.) to enable the use of image analysis algorithms to extract meaningful quantitative measures (e.g., to detect and classify biological objects of interest). Traditionally, a pathologist may manually annotate each successive image of multiple tissue sections from a tissue sample to identify the same aspect for each successive tissue section.

[0022] One type of tissue staining is histochemical staining, which uses one or more chemical dyes (e.g., acid dyes, basic dyes) to stain tissue structures. Histochemical stains may be used to show general aspects of tissue morphology and / or cell microanatomy (e.g., distinguishing cell nuclei from cytoplasm, showing lipid droplets, etc.). One example of a histochemical stain is hematoxylin and eosin (H&E). Other examples of histochemical stains include trichrome stains (e.g., Masson's trichrome), periodic acid Schiff (PAS), silver stains, and iron stains. Another type of tissue staining is immunohistochemistry (IHC, also called "immunostaining"), which uses a primary antibody that specifically binds to a target antigen of interest (also called a biomarker). IHC may be direct or indirect. In direct IHC, the primary antibody is directly conjugated to a label (e.g., a chromophore or fluorophore). In indirect IHC, a primary antibody is first bound to the target antigen, and then a secondary antibody conjugated with a label (eg, a chromophore or fluorophore) is bound to the primary antibody.

[0023] Mitosis is a stage in the life cycle of a biological cell in which duplicated chromosomes separate to form two new nuclei, and a "mitotic figure" is a cell undergoing mitosis. The prevalence of mitotic figures in tissue samples can be used as an indicator of disease prognosis, especially in oncology. For example, tumors that exhibit high mitotic rates tend to have a poorer prognosis than tumors that exhibit low mitotic rates, and mitotic rate can be defined as the number of mitotic figures per a given number (e.g., 100) of tumor cells. Such an indicator can be applied, for example, to a variety of solid and hematological malignancies. Review by a pathologist of slide images (e.g., of H&E stained sections) to determine mitotic rate can be very labor intensive. 2. Introduction

[0024] An image classification task may be configured as a process of classifying structures depicted in a central region of an image. Such a task may be an element of a larger process, such as a process of quantifying the number of structures of a particular class depicted in a larger image (e.g., a WSI). FIG. 1A shows an example of an image processing pipeline 100 in which a large image expected to contain multiple depictions of structures of a particular class (in this example, a WSI, or tiles of a WSI) is provided as input to a detector 110 (e.g., a trained neural network such as a CNN). The detector 110 processes the large image to locate structures therein that resemble structures of a desired class, called "candidates." The predicted likelihood that a candidate is actually in the desired class may be indicated by a detection score. Candidates with a detection score above a predefined threshold are then processed by a classifier 120 (e.g., including the architectural components described herein) to further predict whether each candidate is actually of the desired class or is a substitute for one or more other classes (e.g., even though it may have a high similarity to structures of the desired class). In such applications, the inputs to the classifier 120 may be image crops, each centered on a different one of the candidates identified by the detector 110.

[0025] The deep learning architecture components described herein (and variations thereof) may be used at the end of a classification model, for example as the final stage of the classifier 120 of the pipeline 100. Such components may be applied, for example, to a classifier on top of a feature extraction backend, although the possibilities and disclosed applications are not limited to such applications. Applying such components may improve, for example, the quality of cell classification (e.g., flattened fully connected end of a neural network based on transfer learning, compared to the standard).

[0026] The architectural components disclosed herein may be particularly advantageous for applications in which the center of an object to be classified (e.g., an individual cell) is exactly or approximately at the center of an input image (e.g., a candidate image as shown in FIG. 1A) provided for classification. Such objects may vary in size over several ranges that can be approximately characterized by different radii.

[0027] The architecture components disclosed herein may also be implemented to aggregate sets of output features generated from various radius crops. Such aggregation may cause the neural network of the architecture component to more strongly weight information from the center of the input image and gradually decrease the influence of surrounding regions as the distance from the center of the image increases. During training, the neural network may learn to aggregate output features according to the most informative distribution of importance influences.

[0028] The techniques described herein may include, for example, applying the same set of convolutional neural network operations across a spectrum of concentric (e.g., concentric) crops of a 2D feature map at various radii from the center of the feature map. By extension, the techniques described herein also include applying the same set of convolutional neural network operations across a spectrum of concentric (e.g., concentric) crops of a 3D feature map (e.g., generated from a volumetric image such as may be generated by a PET or MRI scan) at various radii from the center of the feature map.

[0029] The architecture components described herein may be applied, for example, as part of a mitotic figure detection and classification pipeline (which may be implemented, for example, as an example of pipeline 100 as shown in FIG. 1A). A detector may be configured to detect the location of cells predicted to be mitotic figures ("candidates"). Candidates with a detection score above a predefined threshold may then be processed by a classifier (e.g., including the architecture components described herein) to further predict whether a given cell is in fact a mitotic figure or a substitute for one of one or more other cell classes (e.g., including one or more classes that may share high similarity with mitotic figures). In this application, the input to the classifier may be an image crop centered on a candidate identified by the detector. FIG. 1B shows an example of such an input image, with a candidate (in this case, a mitotic figure) at its center, indicated by a bounding box, and similar structures of different classes also indicated by bounding boxes. FIG. 2A shows another example of an input image (e.g., generated by the mitotic figure detector described above) with a candidate (in this case, a mitotic figure indicated by a red circle) at its center.

[0030] It may be desirable to exploit this property of input images where the region of interest (ROI) is centered within the image by training a neural network to focus more on analyzing cells that are located in the center of the image crop, rather than analyzing neighboring cells.

[0031] Information such as the size of the central cell and / or its relations to its neighbors may strongly influence the classification: solutions in which the input image is obtained by simply extracting the bounding box of the detected cells and rescaling it to a certain size may exclude such information and such solutions may not be good enough to yield optimal results.

[0032] In contrast to focusing on the detected cells by using bounding boxes with sizes that may vary from cell to cell depending on the size of the cells, it may be desirable to take crops around the detected cells with the same resolution and size for each detected cell. Such kind of image crops include the cell's neighborhood that may contain useful contextual information for cell evaluation. Furthermore, the constant resolution of these crops from one cell to another (e.g., without rescaling that may be required for bounding boxes of various sizes) allows for the estimation of size information of the cells.

[0033] Analysis of misclassified test samples can help identify situations that may cause errors. For example, we found that test samples with slightly off-center cell ground truth annotations are more likely to be misclassified, especially if another cell is relatively close to the center of the input image compared to the mitotic ground truth. Figure 2B shows an example of an input image where a mitotic figure (indicated by a red circle) is not in the center of the image.

[0034] This problem can arise, for example, in real-world scenarios where the detected candidates are slightly off-center. One solution to this problem is to apply some proportional random position shift to the training, which can help to extend the model's attention to include cells that are in some neighborhood of the center, rather than being limited only to cells that are exactly in the center.

[0035] However, such a solution may be difficult to calibrate, since random position shifts may give rise to new kinds of misclassifications. One such consequence may be that the model does not always center on the most central cells. For example, if a mitotic figure is surrounded by several tumor cells, it tends to be misclassified as a tumor cell. Such misclassification may occur because the features of neighboring cells have a higher influence on the final output, despite the fact that the target cell to be actually classified is closer to the center. Figure 2C shows an example of an input image in which a mitotic figure (indicated by a red circle) is surrounded by several tumor cells.

[0036] Therefore, a more reliable way of configuring the neural network to assign larger weights to cells closest to the center of the input image may be desirable. One solution is to prepare a hard-coded mask of importance (e.g., with values ​​that are weights ranging from 0 to 1) and multiply (e.g., pixel-wise) the relevant information (a portion of the input image or feature map) by such a mask. The relevant information can be either the input image, or one or more intermediate feature maps generated by the neural network from the input image.

[0037] Figure 3 shows six different examples of importance masks. As a result of multiplying with such a mask, a non-homogeneous convolution output is introduced, forcing the neural network to consider and estimate the distance from the center.

[0038] However, using an importance mask may not be the optimal approach to have the neural network assign a greater importance weight to the center of the input image. The input image may depict cells of various sizes, and the quality of performing center detection may vary between different detector networks and may therefore depend on the particular detector network used. A hard-coded mask selection may not be universally suitable for all situations, and instead a self-calibration mechanism may be desirable.

[0039] Another concern with using importance masks is that directly multiplying the input image with the mask may result in loss of information about the neighborhood and / or loss of uniformity of the inputs of very early convolutional layers. Early convolutional layers tend to extract very basic features of similar image elements in the same way, regardless of their location in the image. Early multiplication with the mask may reduce the similarity between such image elements, which may be detrimental to pattern recognition of neighborhood topologies.

[0040] A baseline approach for transfer learning is to add classical global average pooling followed by one or two fully connected layers and softmax activation after a feature extraction layer from a model pre-trained on a large dataset (e.g., ImageNet). However, global average pooling can have the problem of treating information from all regions of the image as equally important and therefore not leveraging prior knowledge that may be available in this context (e.g., that target objects are mostly close to the image center). The approach described herein involves modifying the classical approach of transfer learning with customizations that leverage the centralized locations of the most important objects.

[0041] The techniques described herein (e.g., techniques using radial convolutions, among others) may be used to implement architecture components in which the component's neural network self-calibrates the distribution of importance between regions at various distances from the image center. In contrast to approaches that multiply the input image with a fixed importance mask, the techniques described herein may avoid the loss of important information about the cell's neighborhood. Furthermore, the techniques described herein may allow low-level convolutions to work in the same way on different (possibly all) regions of the input image, for example, by applying radius-dependent heterogeneity only at later stages of the neural network. Experimental results show that applying the techniques described herein (e.g., techniques involving radial convolutions on feature maps extracted by an EfficientNet backbone) can help neural networks reach higher accuracy in contrast to baselines. 3. Techniques using concentric cropping

[0042] As previously mentioned, the process of cropping the input image to a target object and then scaling the crop to the classifier input size may lead to a loss of information about the target's size and neighborhood. In a technique described in more detail below, constant resolution crops may be generated such that the target object is close to the image center (e.g., at least partially represented in each crop) and the surrounding neighborhood of the object is also included. Such a technique may be implemented to exploit both information about the size of the target object and information about the relationship of the target object to its neighborhood.

[0043] FIG. 4A shows a flow chart of an exemplary process 400 for classifying an input image. The process 400 may be performed using one or more computing systems, models, and networks (e.g., as described herein with respect to FIG. 31). Referring to FIG. 4A, at block 404, a feature map for the input image is generated using a trained neural network including at least one convolutional layer. At block 408, a plurality of concentric crops of the feature map are generated. For each of the plurality of concentric crops, the center of the feature map may coincide with the center of the crop. At block 412, an output vector is generated using information from each of the plurality of concentric crops to represent a characteristic of a structure depicted in a central region of the input image. The structure depicted in the central region of the input image may be, for example, a structure to be classified. The structure depicted in the central region of the input image may be, for example, a biological cell. At block 416, a classification result is determined by processing the output vector. The classification result may, for example, predict that the input image depicts a mitotic figure.

[0044] FIG. 4B illustrates a block diagram of an exemplary architecture component 405 for classifying an input image. The component 405 may be implemented using one or more computing systems, models, and networks (e.g., as described herein with respect to FIG. 28). With reference to FIG. 4B, a trained neural network 420 including at least one convolutional layer receives an input image and generates a feature map for the input image. A cropping module 430 receives the feature map and generates a plurality of concentric crops of the feature map. For each of the plurality of concentric crops, a center of the feature map may coincide with a center of the crop. An output vector generation module 440 is adapted to receive the plurality of concentric crops and generate an output vector that represents a characteristic of a structure depicted in a central region of the input image using the plurality of concentric crops. The structure depicted in the central region of the input image may be, for example, a structure to be classified. The structure depicted in the central region of the input image may be, for example, a biological cell. A classification module 450 determines a classification result by processing the output vector. The classification result may, for example, predict that the input image depicts a mitotic figure.

[0045] FIG. 5A illustrates an example of the operation of block 404 in which a feature map 210 (e.g., a 2D feature map) of an input image 200 is generated (e.g., extracted from the input image 200) using a trained neural network 420 that includes at least one convolutional layer. The trained neural network 420 may be, for example, a pre-trained backbone of a deep convolutional neural network (CNN). Particular examples of such backbones that may be used include, for example, a residual network (ResNet), an implementation of a MobileNet model, or an implementation of an EfficientNet model (e.g., any of a range of models from EfficientNet-B0 to EfficientNet-B7), although any other neural network backbone configured to extract image feature maps may be used (e.g., whether according to another known model or a custom model). The backbone (or backend) of a CNN is defined as the feature extraction part of the network. The backbone may include all of the CNN except for the final layers (e.g., the fully connected layer and the activation layer), or the backbone may also exclude other layers of one or more of the final stages of the network (e.g., global average pooling). In one example, the feature map 210 is the last 2D feature map generated by the network 420 before global average pooling.

[0046] The trained neural network 420 may be trained on a large dataset of images (e.g., more than 10,000, more than 100,000, more than 1 million, or more than 10 million). The large dataset of images may include images depicting non-biological structures (e.g., cars, buildings, manufactured objects). Additionally or alternatively, the large dataset of images may include images that do not depict biological cells (e.g., images depicting animals). The ImageNet project (https: / / www.image-net.org) is an example of a large dataset of images that includes images depicting non-biological structures and images that do not depict biological cells. (A dataset of images that includes images that do not depict biological cells is also referred to herein as a "general dataset.") Semi-supervised learning techniques, such as noisy student training, may be used to increase the size of the dataset for training and / or validation of the network 420 by learning labels for previously unlabeled images and adding them to the dataset. Optional further training (e.g., fine-tuning) of the network 420 in component 405 is also contemplated.

[0047] Training the detector model and / or classifier model (e.g., training the network 420) may include augmenting the set of training images. Such augmentation may include random variations in color (e.g., hue angle rotation), size, shape, contrast, brightness, and / or spatial translation (e.g., to increase robustness to non-centering of the trained model) and / or rotation. In one example of such random augmentation, images in the training set for the classifier network may be slightly proportionally position-shifted along the x-axis and / or y-axis (e.g., up to 10% of the image size along the translation axis). Such an augmentation may provide better adaptation to slight imperfections in the detector network, since these imperfections may slightly shift the detected cell candidates from the center of the input image provided to the classifier. Training a classifier model may involve pre-training the model's backbone neural network (e.g., on a general dataset of images, such as a dataset including images depicting objects that are not visible during inference), and then freezing early layers of the model while continuing to train final layers (e.g., on a specialized dataset of images, such as a dataset of images depicting one or more classes of objects that are to be classified during inference).

[0048] The input image 200 may be generated by a detector stage (e.g., another trained neural network) and may be, for example, a selected portion of the WSI, or a selected portion of a tile of the WSI. The input image 200 may be configured such that the center of the input image is within the region of interest. For example, the input image 200 may depict a biological cell of interest and may be configured such that the representation of the cell of interest is centered within the input image.

[0049] In the example of FIG. 5A, the size of the input image 200 is k channels (ixj) pixels (arranged in rows and columns i and j), and the size of the feature map 210 is p channels (mxn) spatial elements (e.g., features) (arranged in rows and columns m). Typical values ​​of i and j may include, for example, 128 or 256. For example, the input image 200 may be a 128x128 or 256x256 pixel tile of a larger image (e.g., as generated by the detector network described above). Typical values ​​of k may include three (e.g., for images in a three-channel color space such as RGB, YCbCr, etc.). Typical values ​​of m and n may include, for example, any integer from 7 to 33, and the value of m may be, but is not necessarily, equal to the value of n. Typical values ​​of p may include, for example, integers in the range of 10 to 200 (e.g., within the range of 30 to 100).

[0050] 5B illustrates an example of the operation of block 408 in which a cropping operation (performed by cropping module 430 in this example) is used to generate multiple concentric crops 220 of feature map 210. It may be desirable to configure the cropping operation such that the center of each of the multiple concentric crops 220 in a spatial dimension (e.g., dimensions of size 2 and 4 in the example of FIG. 5B) coincides with the center of the feature map 210 in a spatial dimension (e.g., dimensions of size m and n in the example of FIGS. 5A and 5B). For a center-cropped square-shaped N×N feature map, for example, we may obtain crops of any one or more of the following square sizes: N×N, (N-2)×(N-2), (N-4)×(N-4).... down to either 2×2 or 1×1 (depending on whether N is even or odd). 5B, m is equal to n and is even (e.g., divisible by 2), the number of concentric crops 220 is two, and the dimensions of the concentric crops are (2×2)×p and (4×4)×p. In other examples, m and / or n may be odd numbers and / or the number of concentric crops 220 may be three, four, five, or more.

[0051] FIG. 6 illustrates an example of the operation of block 404 applied to a configuration in which the trained neural network 420 of component 405 is implemented as an instance 422 of the efficient net B2 backbone, where the input image is of size (260×260) pixels (potentially rescaled from another size such as 128×128 or 256×256), and the feature map 210 has a size of (9×9) spatial elements xp features. Feature levels P1-P5 and the corresponding feature resolution at each level are also shown. FIG. 6 also illustrates an example of the operation of block 408, where an implementation 432 of the cropping module 430 of component 405 performs a 2D dropout operation on the feature map 210 before cropping the resulting feature map to generate multiple concentric crops 220 (in this example, crops having spatial dimensions (5×5) and (3×3)). It will be appreciated that equivalent results can be obtained by instead performing such a 2D dropout operation within network 422 (e.g., performing a 2D dropout operation on the output of level P5 of network 422 to generate feature map 210).

[0052] 7 illustrates an example of the operations of block 412 applied in a configuration in which an implementation 442 of an output vector generation module 440 of component 405 includes downsampling modules 4420a, 4420b and a combining module 4424. Each downsampling module 4420a, 4420b downsamples a corresponding one of the multiple concentric crops 220 (in this example, the crops shown in FIG. 6) to generate a downsampled feature map, and the combining module 4424 combines the downsampled feature maps to generate the output vector 240. In this particular example, the downsampled feature maps are feature vectors 230a, 230b, and it may be desirable to configure the downsampling modules 230a, 230b such that the dimensions of each feature vector 4420a, 4420b are (1x1)xp. Each of the downsampling modules 4420a, 4420b may downsample the corresponding concentric crop, for example, by performing a global pooling operation (e.g., a global average pooling operation or a global max pooling operation). Examples of the combination module 4424 are described in further detail below.

[0053] The generation of the output vector 240 may be followed by any set of operations, such as, for example, any number of layers, or a single fully connected layer to derive a final prediction score. Figure 7 also illustrates an example of the operation of block 416, where a classification module 450 of component 405 processes the output vector 240 to determine a classification result 250. The classification module 450 may process the output vector 240 using, for example, a sigmoid or softmax activation function. In a particular application for mitosis classification, the following four classes may be used: granulocytes, mitotic figures, resembling cells in appearance (non-mitotic cells resembling mitotic figures), and tumor cells (shown, for example, in Figure 8).

[0054] 9 illustrates a particular example of application of process 400 where the size of feature map 210 is (9×9)×3, the number of concentric crops 220 is 5, the sizes of concentric crops 220 are (9×9)×p, (7×7)×p, (5×5)×p, (3×3)×p, and (1×1)×p, and each of feature vectors 230 is generated by performing global average pooling on corresponding ones of the concentric crops 220 (e.g., each of downsampling modules 4420 performs global average pooling). Alternatively, one or more (possibly all) of feature vectors 230 may be generated by performing global max pooling on corresponding ones of the concentric crops 220 (e.g., one or more (possibly all) of downsampling modules 4420 perform global max pooling). In a further alternative, both types of pooling may be performed, in which case the resulting pairs of feature vectors may be concatenated.

[0055] The combine module 4424 may be implemented to aggregate the set of feature vector instances generated from the concentric crops (e.g., by pooling or other downsampling operations). One example of aggregation is performing either a weighted sum or a weighted average. Such aggregation may be achieved by multiplying each output feature vector by its individual weight and then adding the weighted feature vectors. The result is a weighted sum, which can be further divided by the sum of the weights to obtain the weighted average, although the splitting step is optional since appropriate weights can be learned through a training process anyway and similar practical functionality can be achieved.

[0056] Such an aggregation solution may be implemented, for example, by the assignment of a vector of weights involved in training. The trained weights may learn an optimal distribution of the importance of feature vectors characterizing regions of different radii. In the example of Fig. 10 described below, the vector with values ​​denoted as A, B, C, D, E is a vector of importance weights.

[0057] Figure 10 shows the example of Figure 9 in which the output vector 240 is calculated (e.g., in block 412 or by the combination module 4424) as a weighted sum (or weighted average) of the feature vectors 230. In this example, each of the five feature vectors 230 is weighted by a corresponding one of weights A, B, C, D, and E, which may be trained parameters.

[0058] Aggregation of feature vectors generated from feature maps (e.g., feature vectors generated from multiple concentric crops as described above) may include applying a separately trained model (also called a "shared model," "shared weight model," or "shared vision model") to each of the feature vectors being aggregated. For example, for each of multiple concentric crops of a feature map at different radii, a trained vision model may be applied along the corresponding feature map or vector derived from the crop, as described above as an example of "radial convolution." Here, a "shared model" may be implemented as a solution in which the same set of neural network layers with exactly the same weights are applied to different inputs (e.g., as in the "shared vision model" used to process different input images in Siamese and triplet neural networks). For example, the trained shared weight model may apply the same simultaneous equations across the corresponding feature map or vector derived from the crop, for each of multiple concentric crops. The layers in the shared vision model may differ from each other in their number, shape, and / or other details.

[0059] FIG. 11A shows a flowchart of another exemplary process 1100 for classifying an input image. The process 1100 may be performed using one or more computing systems, models, and networks (e.g., as described herein with respect to FIG. 31). With reference to FIG. 11A, the process 1100 includes an example of block 404 described herein. At block 1108, a plurality of feature vectors are generated using information from the feature map. For example, a plurality of concentric crops of the feature map may be generated at block 1108 (e.g., as described above with reference to block 408), and for each of the plurality of concentric crops generated at block 1108, the center of the feature map may coincide with the center of the crop. At block 1110, a second plurality of feature vectors is generated using the trained model applied separately to each of the plurality of feature vectors generated at block 1108. At block 1112, an output vector is generated that represents characteristics of structures depicted in the input image (e.g., in a central region of the input image) using information from each of the second plurality of feature vectors. The structure shown in the input image may be, for example, a structure to be classified. The structure shown in the input image may be, for example, a biological cell. At block 1116, a classification result is determined by processing the output vector (e.g., as described above with reference to block 416). The classification result may, for example, predict that the input image depicts a mitotic figure.

[0060] FIG. 11B illustrates a block diagram of another exemplary architecture component 1105 for classifying an input image. The component 1105 may be implemented using one or more computing systems, models, and networks (e.g., as described herein with respect to FIG. 31). The component 1105 may be implemented to include a neural network having two sub-networks, a first neural network that is an instance of the trained neural network 420 and a second neural network that is a trained shared model described herein, where the second neural network is configured to process inputs based on the output latent space vectors (feature maps) generated by the first neural network. With reference to FIG. 11B, the component 405 includes an instance (e.g., backbone 422) of the neural network 420 trained as described herein to generate the feature maps. The feature vector generation module 1130 is for generating a plurality of feature vectors using the feature maps. For example, the feature vector generation module 1130 may be adapted to generate a plurality of concentric crops of the feature map (e.g., as an example of the cropping module 430 as described above), and for each of the plurality of concentric crops generated by the feature vector generation module 1130, the center of the feature map may coincide with the center of the crop. The second feature vector generation module 1135 is adapted to generate a second plurality of feature vectors using a trained model (e.g., a second trained neural network) separately applied to each of the plurality of feature vectors generated by the module 1130. The output vector generation module 1140 is adapted to generate an output vector that represents characteristics of a structure depicted in the input image (e.g., in a central region of the input image) using the second plurality of feature vectors. The structure depicted in the input image may be, for example, a structure to be classified. The structure depicted in the input image may be, for example, a biological cell. The classification module 1150 determines a classification result by processing the output vector. The classification result may, for example, predict that the input image depicts a mitotic figure.

[0061] Figure 12 illustrates the example of Figure 9 in which the shared model is applied independently to each of the plurality of feature vectors 235 (e.g., in block 1112 or by the second feature vector generation module 1135) to obtain a second plurality of feature vectors 230. This example also illustrates computing an output vector 240 (e.g., in block 1112 or by the output vector generation module 1140) from the second plurality of feature vectors 235 returned by the shared model by applying self-adjusting aggregation (in this example, by using a weighted sum as described above with reference to Figure 10).

[0062] Another example of aggregation may include concatenating the output feature vectors into a feature table and performing a set of 1D convolutions that exchange information between feature vectors derived from crops of adjacent radii. Such convolutions may be performed, for example, in several layers until a flat vector is reached with information exchanged between all radii. In this way, the training process may learn the relationship between adjacent radii. As another example of "radial convolution", a technique may be described in which a set of operations is convolved over such a changing spectrum (e.g., a spectrum of various radii from the center of the feature map).

[0063] Figure 13 illustrates the example of Figure 9 in which one or more convolutional layers are applied (e.g., in block 1112 or by output vector generation module 1140) to adjacent pairs of feature vectors 230. In a further example, also illustrated in Figure 13, one or more convolutional layers may be applied to adjacent pairs of feature vectors that are the result of applying a shared model to each of the feature vectors 230 independently.

[0064] 14A illustrates an implementation 434 of a cropping module 430 that would generate the multiple concentric crops 220 as five concentric crops having spatial dimensions (1×1), (3×3), (5×5), (7×7), and (9×9). In this example, a feature map 210 of size (9×9) spatial elements×p channels (e.g., features) is included as one of the multiple concentric crops 220.

[0065] 14B illustrates a block diagram of an implementation 446 of the output vector generation module 440, including a pooling module 4460, a shared vision model 4462, and a combination module 4464. The pooling module 4460 generates the feature vector 230 by performing a pooling operation (e.g., global average pooling) on ​​each of the plurality of concentric crops 220. The shared vision model 4462 (e.g., a second trained neural network of the architecture component 405) generates a plurality of 230 second feature vectors by applying the same trained model independently to each of the plurality of 235 feature vectors. In one example, the trained model is implemented as three fully connected layers, each layer having 16 neurons.

[0066] The combining module 4464 generates the output vector 240 using information from each of the plurality of second feature vectors 235. For example, the combining module 4464 may be implemented to calculate a weighted average (or weighted sum) of the plurality 235 of second feature vectors. The combining module 4464 may also be implemented to combine (e.g., to concatenate and / or append) the weighted average (or weighted sum, or a feature vector based on information from such weighted average or weighted sum) with one or more additional feature vectors, and / or to perform additional operations such as dropout and / or application of dense (e.g., fully connected) layers.

[0067] 15A shows a block diagram of such an implementation 4465 of the combine module 4464. The module 4465 calculates a weighted average of multiple second feature vectors 235, combines (e.g., for concatenation) the weighted average with one or more additional feature vectors 250 (e.g., if generated by optional additional feature vector generation module 1160 of the architecture component 405), performs a dropout operation on the combined vector, followed by a dense layer to generate the output vector 240.

[0068] 15B shows a block diagram of an implementation 1162 of the additional feature vector generation module 1160. The module 1162 generates the additional feature vector 250 by applying three layers of 3×3 convolutions (each layer having 16 neurons) to the feature map 210 and flattening the resulting map.

[0069] FIG. 16A illustrates a block diagram of an implementation 1106 of architecture component 1105 that includes an implementation 1132 of feature vector generation module 1130, an implementation 1137 of a second feature vector generation module 1135 (e.g., as an example of shared vision model 4462 described herein), and an implementation 1142 of an output vector generation module 1140. The output vector generation module 1140 may be implemented, for example, as an example of a combining module 1165 described herein. In such a case, the combining module 1165 may be adapted to combine the weighted average (or weighted sum, or a feature vector based on information from such weighted average or weighted sum) with one or more additional feature vectors, for example, as generated by an optional instance of the additional feature vector generation module 1132 of architecture component 1105. FIG. 16B illustrates a block diagram of feature vector generation module 1132 that includes instances of cropping module 434 and pooling module 447 as described above. 4. Techniques for using blended training

[0070] FIG. 17A illustrates a flowchart of an exemplary process 1700 for training a classification model including a first neural network and a second neural network. The process 1700 may be performed using one or more computing systems, models, and networks (e.g., as described herein with respect to FIG. 31). Referring to FIG. 17A, in block 1704, a plurality of feature maps are generated using the first neural network of the classification model and information from images of a first dataset. Each image of the first dataset may depict at least one biological cell. The first neural network may be pre-trained on a plurality of images of a second dataset including images that do not depict biological cells (e.g., ImageNet or another general dataset). The first neural network may be an implementation of the network 420 (e.g., backbone 422) described herein, which may be, for example, to generate each of the plurality of feature maps (as an instance of feature map 210) from a corresponding image of the first dataset (as input image 200). At block 1708, a second neural network of the classification model is trained using information from each of the plurality of feature maps. The second neural network may be, for example, an implementation of a shared vision model described herein (e.g., shared vision model 4462).

[0071] FIG. 17B illustrates a flowchart of another exemplary process 1710 for classifying an input image. The process 1710 may be performed using one or more computing systems, models, and networks (e.g., as described herein with respect to FIG. 31). Referring to FIG. 17B, at block 1712, a feature map is generated using a first trained neural network of a classification model. The input image may depict at least one biological cell, and the first trained neural network may be pre-trained on a first plurality of images including images that do not depict biological cells (e.g., ImageNet or another general dataset). The first neural network may be, for example, an implementation of the network 420 (e.g., backbone 422) described herein, and the feature map may be an instance of the feature map 210 described herein.

[0072] In block 1716, an output vector is generated that represents a characteristic of the structure depicted in the central region of the input image using information from the second trained neural network and the feature map of the classification model. The second trained neural network may be, for example, an implementation of a shared vision model (e.g., shared vision model 4462) described herein. The second trained neural network may be trained by providing the classification model with a second plurality of images depicting biological cells. Block 1716 may be performed, for example, by a module of architecture component 405 or 1105 as described herein. The structure depicted in the central region of the input image may be, for example, a structure to be classified. The structure depicted in the central region of the input image may be, for example, a biological cell. In block 1720, a classification result is determined by processing the output vector. The classification result may, for example, predict that the input image depicts a mitotic figure. 5. Backend Variation: Encoder

[0073] 18A-30 show further examples of methods and architecture components that extend the above examples that may use central emphasis of feature maps (e.g., radial convolutions). The end components of these methods and components are similar to those described above, but the feature maps processed by the end components have more features generated by additional components. The additional components may include, for example, a concatenation of feature maps from two or more models (e.g., feature maps from EfficientNet-B2 and a regional variational autoencoder) and / or feature maps from one or more other CNN backends (e.g., a regional autoencoder; U-Net; one or more other networks pre-trained on either a dataset of images from the ImageNet project or the NoisyStudent dataset, which may be further fine-tuned with the concatenation of specialized datasets (e.g., a dataset of images depicting biological cells).

[0074] 18A-24 relate to implementations of the aforementioned techniques that include a trained encoder configured to generate a feature map. Such an encoder (e.g., an encoder of an encoder-decoder or "autoencoder" network, e.g., a variational autoencoder) may be trained to extract features that are independent of the staining type, and an optimal strategy for transforming learning using a specialized dataset of training images resulted in an improvement in the quality of data from labs, scanners, and stains that was not seen in the training process.

[0075] FIG. 18A illustrates a flowchart of an exemplary process 1800 for classifying an input image. The process 1800 may be performed using one or more computing systems, models, and networks (e.g., as described herein with respect to FIG. 31). With reference to FIG. 18A, in block 1804, a feature map for the input image is generated using a trained encoder including at least one convolutional layer, and the trained encoder generates a latent embedding of at least a portion of the input image. In block 1808, a plurality of concentric crops of the feature map are generated (e.g., as described above with reference to block 408). For each of the plurality of concentric crops, the center of the feature map may coincide with the center of the crop. In block 1812, an output vector is generated that represents a characteristic of a structure depicted in a central region of the input image using the plurality of concentric crops (e.g., as described above with reference to block 412). The structure depicted in the central region of the input image may be, for example, a structure to be classified. The structure depicted in the central region of the input image may be, for example, a biological cell. At block 1816, a classification result is determined by processing the output vector (e.g., as described above with reference to block 416). The classification result may, for example, predict that the input image depicts a mitotic figure.

[0076] FIG. 18B illustrates a flow chart of an exemplary process 1400a for classifying an input image. The process 400 may be performed using one or more computing systems, models, and networks (e.g., as described herein with respect to FIG. 31). With reference to FIG. 18B, in block 1804a, a feature map for the input image is generated using a trained neural network including at least one convolutional layer and a trained encoder including at least one convolutional layer, the trained encoder generating a latent embedding of at least a portion of the input image. The trained neural network may be, for example, an implementation of the network 420 (e.g., 422) described herein, and the feature map may be based on an instance of the feature map 210 described herein. Blocks 1808, 1812, and 1816 are as described above with reference to FIG. 18A. Implementations of the processes 1800 and 1800a are described in further detail below.

[0077] Figure 19A shows a schematic diagram of an encoder-decoder (or "autoencoder") network configured to receive an input image of size (ixj) spatial elements xk channels, encode the image into a feature map in a latent embedding space of size qxrxs, and decode the feature map to generate a reconstructed image of size (ixj) spatial elements xk channels. Figure 19B shows an example of a block 1804 in which the feature map 212 is generated by a trained encoder 150 of such an encoder-decoder network.

[0078] Figure 20 shows example input images and corresponding reconstructed images produced by an encoder-decoder network featuring a latent embedding space of size 4x4x128 at various epochs during training. Figure 21 shows example input images (labeled "REFERENCE") and corresponding reconstructed images produced by an encoder-decoder network for latent embedding spaces of different sizes (4x4x128, 1x1x1024, and 1x1x128).

[0079] 22A and 22B show an example of block 1804a in which a feature map 212 is generated from an input image 200 using a trained encoder 150, a feature map 210 is generated from an input image 200 using a trained neural network 420 (FIG. 22A), and a feature map 214 is generated using the feature maps 210, 212 and a combining operation 160 (FIG. 22B). In this particular example, the size of the feature map 212 is (q×r) spatial elements×s channels, and the size of the feature map 210 is (m×n) spatial elements×p channels. The combining operation 160 may include resizing one or more of the feature maps 210, 212 along a spatial dimension to a common size and then concatenating the common size feature maps to generate the feature map 214.

[0080] Modules 430, 440, and 450 of architecture component 405 as described herein may be used to perform blocks 1808, 1812, and 1816. Figure 23 shows a block diagram of such an example in which a feature map 214 of size (9x9) spatial elements x (p+s) channels (e.g., features) is processed using instances of cropping module 434, pooling module 4460, shared vision model 4462, combining module 4465, and additional feature vector generation module 1162 as described herein to generate output vector 240. Similarly, modules 1130, 1135, 1140, and 1150 of architecture component 405 as described herein may be used to perform blocks 1808, 1812, and 1816. FIG. 24 shows a block diagram of such an example in which a feature map 214 of size (9×9)×(p+s) is processed using instances of the feature vector generation module 1132, the second feature vector generation module 1137, the output vector generation module 1142, and the additional feature vector generation module 1162 described herein to generate an output vector 240. 6. Backend variations: multiple feature maps (e.g., "feature pyramids")

[0081] It may be desirable to implement the architecture described herein to combine multiple feature maps. For example, it may be desirable to combine feature maps that represent an input image at different scales (e.g., a pyramid of feature maps). Figures 25A and 25B show an example of an operation in which multiple feature maps 210 are generated from an input image 200 by an implementation 424 of a trained neural network 420 (Figure 25A), and the feature maps 216 and a combining operation 164 (Figure 25B) are used to generate the feature maps 210. In this particular example, the number of multiple feature maps 210 is two, and the size of the feature maps 210 is (m x n) spatial elements xp channels and (t x u) spatial elements xv channels. The combining operation 164 may include resizing (e.g., rescaling) one or more of the feature maps 210 along a spatial dimension to a common size (e.g., m x n in the example of Figure 25B), and then concatenating the common size feature maps to generate the feature map 216. FIG. 26 shows a block diagram of an example in which a feature map 216 of size (9×9) spatial elements x (p+v) channels (e.g., features) is processed using instances of the cropping module 434, pooling module 4460, shared vision model 4462, combining module 4465, and additional feature vector generation module 1162 as described herein to generate an output vector 240.

[0082] Each of the plurality of feature maps 210 may be the output of a different corresponding layer of the backbone. FIG. 27 shows a block diagram of a particular example in which each of the plurality of feature maps 216 is obtained as the output of a corresponding final layer of an EfficientNet-B2 implementation 426 of a trained neural network 420. In this example, two of the feature maps 210 are rescaled to a common size of (9×9) spatial elements, the common-sized feature maps are combined (concatenated), and a 2D dropout operation is applied to the concatenated maps to generate the feature map 216. FIG. 28 shows a block diagram of an example in which a feature map 216 of size (9×9) spatial elements×(L+M+N) channels (e.g., features) is processed using instances of the cropping module 434, pooling module 4460, shared vision model 4462, combining module 4465, and additional feature vector generation module 1162 as described herein to generate an output vector 240.

[0083] 29 and 30 show examples of combining feature maps from a pyramid as described above with feature maps from a trained encoder as described above. FIG. 29 shows a block diagram of a particular example in which a feature map of dimension (8×8)×H generated using a trained encoder 152 is rescaled and combined (concatenated) with multiple common-sized feature maps as shown in FIG. 27, and a 2D dropout operation is applied to the concatenated map to generate a feature map 218. FIG. 30 shows a block diagram of an example in which a feature map 218 of size (9×9) spatial elements×(H+L+M+N) channels (e.g., features) is processed using instances of the cropping module 434, pooling module 4460, shared vision model 4462, combining module 4465, and additional feature vector generation module 1162 described herein to generate an output vector 240.

[0084] FIG. 31 illustrates an example of a computing system 3100 that can be configured to perform the methods described herein. 7. Experiment

[0085] Four experiments as described below were performed and the results were evaluated using the parameters VACC (validation accuracy) and VMAUC (validation mitosis area under the curve) on the same validation set. The parameter VACC was calculated by measuring the accuracy across all four classes and averaging the results. In this case, accuracy is measured with reference to the top-1 classification (e.g., for each prediction, only the class with the highest classification score is "yes", the other three are all "no"). The parameter VMAUC was calculated as the area under the precision-recall (PR) curve. The PR curve is a plot of the precision (on the Y axis) and recall (on the X axis) of a single classifier (in this case only the mitosis class) at various binary thresholds. Each point on the curve corresponds to a different binary threshold and shows the resulting precision and recall when the threshold is used. If the score of the mitosis class is equal to or greater than the threshold, the answer is "yes", and if the score is below the threshold, the answer is "no". A higher value of VMAUC (e.g., a larger area under the PR curve) indicates that the overall model performs better than many different binary thresholds, and so one can choose among these thresholds. Among these checkpoints, the one with the best achieved VMAUC and the one with the best VACC in the four classes (the two different optimization criteria are achieved at different training times on the same validation set) are shown in Table 1 below. Figure 32 shows an example of a PR curve for the model described herein with reference to Figures 29 and 30.

[0086] Experiment 1: We trained EfficientNet-B2 with "flat-ended" (global average pooling and one fully connected layer) in the final layer. The checkpoint with the best area under the precision-recall curve for the mitosis class: VACC=0.69833, VMAUC=0.74141; the checkpoint with the best validation accuracy (for all four classes): VACC=0.72168, VMAUC=0.72249.

[0087] Experiment 2: Training was performed with exactly the same parameterization of image augmentation and the same dataset, but with a flat termination (this name denotes one of the implementations of the "radial convolution" model that terminates as described herein with reference to, for example, Figs. 12 and 13). The checkpoint with the best area under the precision-recall curve for the mitosis class: VACC = 0.70843, VMAUC = 0.74549; the checkpoint with the best validation accuracy (for all four classes): VACC = 0.73178, VMAUC = 0.73948. In each optimization metric, this improved architecture (with "radial convolution" termination) was able to achieve higher (better) scores on the same validation set and keep both metrics better than the architecture with "flat termination".

[0088] Experiment 3: It was possible to achieve better model performance using the same architecture (with the "EfficientNet-B2" backend and the "Radial Convolution" backend) with the following improvements: Checkpoint with best area under the precision-recall curve for the mitosis class: VACC=0.81161, VMAUC=0.83776; Checkpoint with best validation accuracy (for all four classes): VACC=0.82203, VMAUC=0.83076. Improvements include:

[0089] Improved training techniques (sequential training with selective freezing / unfreezing of some neural network layers in a specific order, manipulation of the learning rate in successive training runs). The best working strategy was to select the set of model weights with the best accuracy achieved so far and perform the following actions:

[0090] 1) First, we train only the model using ADAM gradient descent with a high starting learning rate and then freeze the remaining weights to terminate.

[0091] 2) Second, we unfreeze the selected number of final layers in the backend core part of the model and continue training from the best checkpoint using ADAM with a starting learning rate that is at least 10 times smaller.

[0092] 3) Third, we freeze the entire backend again and train only the end of the neural network (similar to the first step, but with an even lower initial learning rate).

[0093] Improved data augmentation (in-house extension library mixed with built-in extensions from Keras and FastAI); manipulation of these extensions and their amplitudes and occurrence probabilities had a substantial impact on improving the results, making the model more robust to scanner, tissue and / or staining variations. Augmenting the training set with data scanned by the same scanner(s) as the validation set further improved the robustness of the validation data.

[0094] This strategy also included transfer learning from the best model trained on the first dataset (obtained from slides of tissue samples by at least one different scanner from the validation dataset), further training only on the second dataset (obtained from slides of tissue samples by the same scanner as the validation dataset) until the model stops improving, and then training again on the mixed training set. The slides used to obtain the second dataset may be from a different laboratory, a different organism (e.g., canine tissue vs. human tissue), a different type of tissue (e.g., skin tissue vs. various different breast tissues), and / or a different type of tumor (e.g., lymphoma vs. breast cancer) than the slides used to obtain the first dataset. Additionally or alternatively, the slides used to obtain the second dataset may be stained with chemical staining substances and / or in different proportions than the slides used to obtain the first dataset. The scanner used to obtain the first dataset may have different color, brightness, and / or contrast characteristics than the scanner used to obtain the second dataset. Several strategies of mixing different proportions of the first and second datasets were tried and the best one was selected.

[0095] Experiment 4: In further experiments, the strategy remained the same, but an additional feature extraction backend model was added to provide a mixture of feature maps with pre-trained knowledge. A "pyramid of features" strategy was used (selected earlier layers of the backend model were scaled down to match the size of the last layer, and those layers were concatenated to create one feature map with more features corresponding to each local region in the input image, these features representing different levels of abstraction). In this strategy, more terminal layers of the model backbone were unfrozen during model fine-tuning. In some cases, the entire model backbone was unfrozen for fine-tuning, and in other cases, only the last pyramid level and the model termination were unfrozen. In one particular example, an EfficientNet-B0 model backbone was used, and the last 20 layers were unfrozen for fine-tuning (e.g., 18 layers of the last pyramid level of the backbone, and 2 layers of the model ending above the last pyramid level). In another particular example, we used an EfficientNet-B2 model backbone and unfrozen the last 35 layers for fine-tuning (e.g., 33 layers in the last pyramid level of the backbone, and two layers in the model ending above the last pyramid level). (Note that the end of the model applying the radial convolution described herein may have more than two layers.) In other such examples, we unfrozen the entire model backbone for fine-tuning. Additionally, feature maps from the bottleneck layer of one of the regional variational autoencoders were added to further extend the feature maps. We trained the autoencoder on a variety of cells from H&E tissue images, with the same assumption that each cell is at the center of the image crop.

[0096] Such experiments yielded small further improvements: best achieved area under the precision-recall curve for the mitosis class: VMAUC=0.83966; best achieved validation accuracy (for all four classes): VACC=0.82550. In this approach, the end of the "radial convolution" model was only modified to accommodate an increase in the number of input features per region of the feature map.

[0097] [Table 1]

[0098] 8. Typical Systems for Center Emphasis 31 is a block diagram of an exemplary computing environment having an exemplary computing device suitable for use in some exemplary implementations, such as for executing methods 400, 1100, 1700, 1710, 1800, and / or 1800a. The computing device 3105 in the computing environment 3100 may include one or more processing units, cores, or processors 3110, memory 3115 (e.g., RAM, ROM, etc.), internal storage 3120 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or I / O interfaces 3125, any of which may be coupled over a communication mechanism or bus 3130 for communicating information or incorporated in the computing device 3105.

[0099] The computing device 3105 may be communicatively coupled to an input / user interface 3135 and an output device / interface 3140. Either or both of the input / user interface 3135 and the output device / interface 3140 may be a wired or wireless interface and may be detachable. The input / user interface 3135 may include any device, component, sensor, or interface, physical or virtual, that may be used to provide input (e.g., buttons, touch screen interface, keyboard, pointing / cursor control, microphone, camera, Braille, motion sensor, optical reader, etc.). The output device / interface 3140 may include a display, television, monitor, printer, speaker, Braille, etc. In some exemplary implementations, the input / user interface 3135 and the output device / interface 3140 may be incorporated or physically coupled with the computing device 3105. In other exemplary implementations, other computing devices may function as or provide the functionality of the input / user interface 3135 and the output device / interface 3140 for the computing device 3105.

[0100] The computing device 3105 may be communicatively coupled (e.g., via I / O interface 3125) to an external storage device 3145 and to a network 3150 for communicating with any number of networked components, devices, and systems, including one or more computing devices of the same or different configurations. The computing device 3105 or connected computing devices may function as, provide services for, or be referred to as a server, client, thin server, general purpose machine, special purpose machine, or another label.

[0101] I / O interface 3125 may include, but is not limited to, wired and / or wireless interfaces using any communication or I / O protocol or standard (e.g., Ethernet, 802.11x, Universal System Bus, WiMax, modem, cellular network protocols, etc.) for communicating information between at least all connected components, devices, and networks in computing environment 3100. Network 3150 may be any network or combination of networks (e.g., the Internet, a local area network, a wide area network, a telephone network, a cellular network, a satellite network, etc.).

[0102] The computing device 3105 can use and / or communicate using computer usable or computer readable media, including transitory and non-transitory media. Transitory media include transmission media (e.g., metallic cables, optical fibers), signals, carrier waves, etc. Non-transitory media include magnetic media (e.g., disks and tapes), optical media (e.g., CD-ROMs, digital video disks, Blu-ray® disks), solid-state media (e.g., RAM, ROM, flash memory, solid-state storage), and other non-volatile storage or memory.

[0103] The computing device 3105 may be used to implement techniques, methods, applications, processes, or computer-executable instructions in some exemplary computing environments. The computer-executable instructions may be retrieved from a transitory medium, stored on a non-transitory medium and retrieved from the non-transitory medium. The executable instructions may be from one or more of any programming language, scripting language, and machine language (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).

[0104] The processor 3110 may execute under any operating system (OS) (not shown) in a native or virtual environment. One or more applications may be deployed, including a logic unit 3160, an application programming interface (API) unit 3165, an input unit 3170, an output unit 3175, a boundary mapping unit 3180, a control point determination unit 3185, a transformation calculation and application unit 3190, and an inter-unit communication mechanism 3195 for different units to communicate with each other, with the OS, and with other applications (not shown). For example, the trained neural network 420, the cropping module 430, the output vector generation module 440, and the classification module 450 may implement one or more processes described and / or shown in Figures 4A, 11A, 17A, 17B, 18A, and / or 18B. The described units and elements may vary in design, function, configuration, or implementation and are not limited to the provided description.

[0105] In some example implementations, once information or execution instructions are received by the API unit 3165, the information or execution instructions may be communicated to one or more other units (e.g., logic unit 3160, input unit 3170, output unit 3175, trained neural network 420, cropping module 430, output vector generation module 440, and / or classification module 450). For example, after the input unit 3170 detects a user input, the input unit may use the API unit 3165 to communicate the user input to the implementation 112 of the detector 100 to generate an input image (e.g., from the WSI or a tile of the WSI). The trained neural network 420 may interact with the detector 112 via the API unit 3165 to receive the input image and generate a feature map. Using the API unit 3165, the cropping module 430 may interact with the trained neural network 420 to receive the feature map and generate multiple concentric crops. Using the API unit 3165, the output vector generation module 440 may interact with the cropping module 430 to receive the concentric crops and generate an output vector that represents characteristics of the structure depicted in the central region of the input image using information from each of the multiple concentric crops. Using the API unit 3165, the classification module 450 may interact with the output vector generation module 440 to receive the output vectors and determine a classification result by processing the output vectors. Further example implementations of applications that may be deployed may include a second feature vector generation module 1135 as described herein (see, e.g., FIG. 11B).

[0106] In some cases, logic unit 3160 may be configured to control information flow between units and direct services provided by API unit 3165, input unit 3170, output unit 3175, trained neural network 420, cropping module 430, output vector generation module 440, and classification module 450 in some example implementations described above. For example, the flow of one or more processes or implementations may be controlled solely by logic unit 3160 or in conjunction with API unit 3165. 9. Further Considerations

[0107] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0108] The terms and expressions used are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions to exclude any equivalents of the features or portions thereof shown and described, recognizing that various modifications are possible within the scope of the claimed invention. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it should be understood that modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, and such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0109] The description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the description of the preferred exemplary embodiments provides those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.

[0110] In the following description, specific details are given to provide a comprehensive understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other examples, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

Claims

1. 1. A computer-implemented method for classifying an input image, comprising the steps of: generating a feature map for the input image using a trained neural network including at least one convolutional layer; generating a plurality of concentric crops of the feature map; generating an output vector representative of a characteristic of a structure depicted in a central region of the input image using information from each of the plurality of concentric crops; determining a classification result by processing the output vector; 23. A computer-implemented method comprising:

2. The computer-implemented method of claim 1 , wherein the structure depicted in the central region of the input image is a structure to be classified.

3. The computer-implemented method of claim 1 or 2, wherein the structures depicted in the central region of the input image are biological cells.

4. The computer-implemented method of claim 1 , wherein for each of the plurality of concentric crops, a center of the feature map coincides with a center of the crop.

5. The computer-implemented method of claim 1 , wherein the classification result predicts that the input image depicts a mitotic figure.

6. 6. The computer-implemented method of claim 1 , further comprising generating a latent embedding of at least a portion of the input image using a trained encoder including at least one convolutional layer, wherein generating the feature map also uses the latent embedding.

7. 7. The computer-implemented method of claim 1, further comprising: generating a respective one of a plurality of feature maps using a respective one of a plurality of final layers of the trained neural network, wherein generating the feature maps uses a concatenation of the plurality of feature maps.

8. generating the output vector using information from each of the plurality of concentric crops, generating, for each of the plurality of concentric crops, a corresponding one of a plurality of feature vectors using at least one pooling operation; generating the output vector using information from each of the plurality of feature vectors; The computer-implemented method of claim 1 , comprising:

9. 9. The computer-implemented method of claim 8, wherein generating the output vector using information from each of the plurality of feature vectors comprises generating the output vector using a weighted sum of the plurality of feature vectors.

10. ordering the plurality of feature vectors by radial sizes of the corresponding concentric crops; 10. The computer-implemented method of claim 8, wherein generating the output vector comprises separately convolving a trained filter across adjacent pairs of the ordered plurality of feature vectors.

11. the plurality of feature vectors are ordered by radial size of the corresponding concentric crops; Generating the output vector comprises: convolving the trained filter across a first adjacent pair of the ordered plurality of feature vectors to generate a first combined feature vector; convolving the trained filter over a second adjacent pair of the ordered plurality of feature vectors to generate a second combined feature vector; Including, 9. The computer-implemented method of claim 8, wherein generating the output vector using the multiple feature vectors comprises generating the output vector using the first combined feature vector and the second combined feature vector.

12. generating the output vector using the plurality of feature vectors, generating a second plurality of feature vectors using the trained model, the second plurality of feature vectors including separately applying the trained model to each of the plurality of feature vectors; generating the output vector using information from each of the second plurality of feature vectors; The computer-implemented method of claim 8 , comprising:

13. 13. The computer-implemented method of claim 12, wherein generating the output vector using information from each of the second plurality of feature vectors comprises generating the output vector using a weighted sum of the second plurality of feature vectors.

14. 14. The computer-implemented method of claim 1, further comprising using a second trained neural network to select the input image as a patch of a larger image, a central region of the input image depicting a biological cell.

15. The computer-implemented method of claim 1 , wherein processing the output vector comprises applying a sigmoid function to the output vector.

16. 1. A computer-implemented method for classifying an input image, comprising the steps of: generating a feature map for the input image; generating a plurality of feature vectors using the feature map; generating a second plurality of feature vectors using the trained model, the second plurality of feature vectors including applying the trained model separately to each of the plurality of feature vectors; generating an output vector representative of a characteristic of a structure depicted in the input image using information from each of the second plurality of feature vectors; determining a classification result by processing the output vector; 23. A computer-implemented method comprising:

17. 20. The computer-implemented method of claim 16, wherein using the feature map to generate the plurality of feature vectors comprises using at least one pooling operation.

18. 18. A computer-implemented method according to claim 16 or 17, wherein the structure is depicted in a central portion of the input image.

19. 19. The computer-implemented method of claim 16, wherein generating the output vector using the second plurality of feature vectors comprises generating the output vector using a weighted sum of the second plurality of feature vectors.

20. 1. A computer-implemented method for training a classification model comprising a first neural network and a second neural network, the method comprising: generating a plurality of feature maps using a first neural network and information from the images of a first dataset; training the second neural network using information from each of the plurality of feature maps; Including, each image of the first dataset depicts at least one biological cell; A computer-implemented method, wherein the first neural network is pre-trained on a plurality of images of a second dataset that includes images that do not depict biological cells.

21. 21. The computer-implemented method of claim 20, wherein the plurality of images of the second data set includes images depicting non-biological structures.

22. 22. The computer-implemented method of claim 20 or 21, wherein for each image of the first data set, a central region of the image depicts at least one biological cell.

23. 1. A computer-implemented method for classifying an input image, comprising the steps of: generating a feature map for the input image using a first trained neural network of a classification model; generating an output vector representative of characteristics of structures depicted in a central portion of the input image using a second trained neural network of the classification model and information from the feature map; determining a classification result by processing the output vector; Including, the input image depicts at least one biological cell; the first trained neural network is pre-trained on a first plurality of images including images that do not depict biological cells; The computer-implemented method, wherein the second trained neural network is trained by providing the classification model with a second plurality of images depicting biological cells.

24. 24. The computer-implemented method of claim 23, wherein the first plurality of images includes images depicting non-biological structures.

25. 25. The computer-implemented method of claim 23 or 24, wherein for each image of the second plurality of images, a central region of the image depicts at least one biological cell.

26. one or more data processors; a non-transitory computer readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein; A system comprising:

27. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product comprising instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.

Citation Information

Patent Citations

  • Augmented reality microscopy for pathology with quantitative biomarker data overlay

    JP2021515240A