Processing histology images with convolutional neural networks to identify tumors

CN122841925APending Publication Date: 2026-09-29LEICA BIOSYSTEMS IMAGING INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611146814.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-12-29
Filing Date
2018-12-21
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,在实践中,多年来已证明该问题极难使计算机自动化达到所需的准确度水平,即达到与病理学家一样好或比病理学家更好的水平

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841925A_ABST
    Figure CN122841925A_ABST
Patent Text Reader

Abstract

A convolutional neural network (CNN) is applied to identify tumors in a histology image. The CNN is assigned one channel for each of a plurality of tissue classes to be identified, at least one class for each of non-tumor and tumor tissue types. A multi-stage convolution is performed on an image patch extracted from the histology image, followed by a multi-stage transposed convolution to restore the layer to a size matching the input image patch. The output image patch thus has a one-to-one pixel-to-pixel correspondence with the input image patch, such that each pixel in the output image patch is assigned one of the plurality of available classes. The output image patch is then assembled into a probability map, which can be presented alongside or superimposed with the histology image. The probability map can then be linked to the histology image for storage.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Leica Biological Systems Imaging Inc., filed on December 21, 2018, with application number 201880084713.7, entitled "Using Convolutional Neural Networks to Process Histological Images to Identify Tumors".

[0002] Cross-reference to related applications

[0003] This application claims priority to U.S. Provisional Patent Application No. 62 / 611,915, filed December 29, 2017, which is incorporated herein by reference as if fully set forth herein. Technical Field

[0004] This disclosure relates to using convolutional neural networks (CNNs) to process histological images in order to identify tumors. Background Technology

[0005] Cancer is the second leading cause of death among women in North America. Of all types of cancer affecting women, breast cancer is the most common and the second leading cause of cancer death. Therefore, the accuracy of breast cancer treatment has a significant impact on the lifespan and quality of life of a substantial proportion of women who will be affected by breast cancer at some point in their lives.

[0006] Based on the expression of specific genes, breast cancer can be divided into different molecular subtypes. Commonly used classification schemes are as follows:

[0007] 1. Type A lumen: ER+, PR+, HER2-

[0008] 2. Lumen type B: ER+, PR-, HER2+

[0009] 3. Triple-negative breast cancer (TNBC): ER-, PR-, HER2-

[0010] 4. HER2-enriched type: HER2+, ER-, PR-

[0011] 5. Normal mammary gland-like type.

[0012] ER stands for estrogen receptor. PR stands for progesterone receptor. HER2 stands for human epidermal growth factor receptor 2.

[0013] Some of these subtypes (such as luminal type A) are hormone receptor-positive and can be treated with hormone therapy. Other breast cancers are HER2-positive and can be treated with drugs that target HER2, such as trastuzumab (trademarked as Herceptin, a registered trademark of F. Hoffmann-La Roche AG). It is important to determine whether breast cancer is ER, PR, or HER2-positive because non-positive cancers will not respond to this type of treatment and require alternative treatment. Traditionally, positivity is determined by a pathologist examining histological samples on slides under a microscope. On the slide, hormone-specific antibodies are applied to formalin-fixed paraffin-embedded breast tissue sections taken from the patient using immunohistochemistry (IHC). After IHC staining, the immunoreactive tissue appears as a brown precipitate under a microscope. The College of American Pathologists (CAP) recommends evaluating all stained, tumor-containing areas to obtain a positivity percentage. To improve inter- and intra-observer reproducibility, CAP guidelines recommend using image analysis to calculate the positivity percentage. Currently, pathologists performing image analysis must manually delineate tumor regions because they must calculate the positive percentage only on areas with tumor cells. The delineation process itself is tedious and error-prone, and can lead to reduced reproducibility.

[0014] In principle, this problem is suitable for computer automation, using conventional image processing methods such as segmentation and boundary finding, or artificial intelligence methods, especially neural networks and deep learning. However, in practice, it has proven over the years that this problem is extremely difficult to automate to the required level of accuracy—that is, to be as good as or better than a pathologist.

[0015] We have already attempted to use CNNs for image processing of histological images of breast cancer and other cancers.

[0016] Wang et al. [1] described a CNN method for detecting breast cancer metastasis to lymph nodes in 2016[1].

[0017] US2015213302A1[2] describes how to detect cell mitosis in regions of cancerous tissue. After training a CNN, classification is performed based on an automated cell nucleus detection system that performs mitosis counting, and then used for tumor grading.

[0018] Hou et al. 2016[3] processed brain and lung cancer images. Image patches from whole slice images (WSI) were used for block-level predictions given by a block-level CNN.

[0019] Liu et al. [4] used CNN to process image patches extracted from billion-pixel breast cancer histological images to detect and locate tumors.

[0020] Bejnordi et al.

[2017] [5] applied two stacked CNNs to classify tumors in image patches extracted from WSI of breast tissue stained with hematoxylin and eosin (H&E) staining. We also note that Bejnordi et al. also provide an overview of other CNN-based tumor classification methods applied to breast cancer samples (see references 10-13). Summary of the Invention

[0021] According to one aspect of this disclosure, a method for identifying tumors in a histological image or a collection of histological images is provided, the method comprising:

[0022] Receive histological images or collections of histological images from records stored in a data repository;

[0023] Image patches are extracted from the histological image or collection of histological images, wherein the image patch is a region of the histological image or collection of histological images whose size is defined by the number of pixels in its width and height;

[0024] A convolutional neural network is provided with a set of weights and multiple channels, each channel corresponding to one of multiple tissue categories to be identified, wherein at least one of the tissue categories represents non-tumor tissue and at least one of the tissue categories represents tumor tissue;

[0025] Each image patch is input as an input image patch into the convolutional neural network;

[0026] Perform multiple levels of convolution to generate convolutional layers whose size decreases progressively until a final convolutional layer of minimum size is included. Then perform multiple levels of transposed convolution to invert the convolutions by generating deconvolutional layers of progressively increasing size, until the layers are restored to a size that matches the input image patch. Each pixel in the restored layer contains a probability of belonging to each of the tissue categories.

[0027] Based on the probability, tissue categories are assigned to each pixel in the recovery layer to obtain output image patches.

[0028] Following the allocation step, the method may further include the step of assembling the output image blocks into a probability map of the histological image or a set of histological images.

[0029] Following the assembly step, the method may further include the following steps: storing the probability map into the record in the data repository such that the probability map is linked to the histological image or a collection of histological images.

[0030] In our current implementation, in each successive convolutional stage, the depth increases as the size decreases, resulting in convolutional layers with increasing depth and decreasing size. Similarly, in each successive transposed convolutional stage, the depth decreases as the size increases, resulting in deconvolutional layers with decreasing depth and increasing size. Finally, the convolutional layer has the maximum depth and minimum size. Instead of continuously increasing and decreasing the depth through convolutional and deconvolutional stages respectively, an alternative is to design a neural network where every layer except the input and output layers has the same depth.

[0031] The method may further include displaying the histological image or set of histological images on a display together with the probability map, for example, overlaying it on the probability map or displaying it side by side. The probability map can be used to determine which regions should be scored by any IHC scoring algorithm to be used. The probability map can also be used to generate a set of contours around tumor cells that can be presented on a display, for example, to allow pathologists to evaluate the results generated by a CNN.

[0032] In some implementations, the convolutional neural network has one or more skip connections. Each skip connection obtains intermediate results from at least one convolutional layer with a larger size than the final convolutional layer, and subjectes these results to as many transposed convolutions as possible to obtain at least one additional recovery layer that matches the size of the input image patch, wherein the number of transposed convolutions may be zero, one, or more. This at least one additional recovery layer is then combined with the aforementioned recovery layer before the step of assigning tissue categories to each pixel. Further processing steps combine the recovery layer with each of the additional recovery layers to recalculate probabilities, thereby taking into account the results obtained from the skip connections.

[0033] In some implementations, the softmax operation is used to generate probabilities.

[0034] Image patches extracted from one or more histological images can cover an entire region of one or more images. These patches can be non-overlapping image tiles or image tiles that overlap at their edges to help stitch together a probability map. While each image patch should have a fixed number of pixels in width and height to match the CNN, since the CNN will be designed to accept only fixed-size pixel arrays, this does not mean that each image patch must correspond to the same physical region on the histological image. Pixels in the histological image can be combined into lower-resolution patches that cover a larger region. For example, each 2×2 array of adjacent pixels can be combined into a “super” pixel to form a patch with a physical region four times larger than the patch extracted at the original resolution of the histological image.

[0035] Once the CNN is trained, the method can be used to make predictions. The purpose of training is to assign appropriate weights to the inter-layer connections. For training, the records used will include ground truth data, which assigns each pixel in a histological image or a collection of histological images to one of the tissue categories. The ground truth data will be based on annotations of a sufficiently large number of images used by expert clinicians. Training is performed by iteratively applying the CNN, where each iteration involves adjusting the weights based on comparing the ground truth data with output image patches. In our current implementation, the weights are adjusted during training using gradient descent.

[0036] Various options exist for setting tissue categories, but most (if not all) implementations will have in common a distinction between categories of non-tumor tissue and tumor tissue. Non-tumor tissue categories can include one, two, or more subcategories. Tumor tissue categories can also include one, two, or more subcategories. For example, in our current implementation, there are three tissue categories: one representing non-tumor tissue and two representing tumor tissue, wherein the two tumor tissue categories represent invasive tumors and in situ tumors.

[0037] In some implementations, the CNN is applied to one histological image at a time. In other implementations, the CNN can be applied to a synthetic histological image formed by combining a set of histological images obtained from adjacent slices of a tissue region that are stained differently. In yet another implementation, the CNN can be applied in parallel to each image in a set of images obtained from adjacent slices of a tissue region that are stained differently.

[0038] Utilizing results from CNNs, this method can be extended to include a scoring process for tumors defined from pixel-based classification and a reference probability map. For example, the method may further include: defining regions in a histological image corresponding to tumors according to the probability map; scoring each tumor according to a scoring algorithm to assign a score to each tumor; and storing the scores in records in a data repository. Thus, scoring is performed on the histological image, but only within regions identified by the probability map as containing tumor tissue.

[0039] The results can be displayed to the clinician on a monitor. That is, histological images can be displayed together with their associated probability maps, for example, overlaid on the probability maps or displayed side by side. Tumor scores can also be displayed in some convenient way, for example, using text labels on or pointing to the tumor, or displayed next to the image.

[0040] A convolutional neural network can be a fully convolutional neural network.

[0041] Another aspect of the present invention relates to a computer program product for identifying tumors in a histological image or a collection of histological images, the computer program product carrying machine-readable instructions for performing the methods described above.

[0042] Another aspect of the present invention relates to a computer device for identifying tumors in a histological image or a collection of histological images, said device comprising:

[0043] Input, operable to receive histological images or sets of histological images from records stored in a data repository;

[0044] A preprocessing module configured to extract image patches from the histological image or set of histological images, the image patches being regions of the histological image or set of histological images whose size is defined by the number of pixels in their width and height; and

[0045] A convolutional neural network having a set of weights and multiple channels, each channel corresponding to one of multiple tissue categories to be identified, wherein at least one of the tissue categories represents non-tumor tissue and at least one of the tissue categories represents tumor tissue, the convolutional neural network being operable to:

[0046] Each image patch is received as input image patch;

[0047] Perform multiple levels of convolution to generate convolutional layers whose size decreases progressively until a final convolutional layer of minimum size is included. Then perform multiple levels of transposed convolution to invert the convolutions by generating deconvolutional layers of progressively increasing size, until the layers are restored to a size that matches the input image patch. Each pixel in the restored layer contains a probability of belonging to each of the tissue categories.

[0048] Based on the probability, tissue categories are assigned to each pixel in the recovery layer to obtain output image patches.

[0049] The computer device may further include a post-processing module configured to assemble the output image blocks into a probability map of the histological image or set of histological images. Furthermore, the computer device may also include an output operable to store the probability map in a record in a data repository, such that the probability map is linked to the histological image or set of histological images. The device may also include a display and a display output operable to transmit the histological image or set of histological images and the probability map to the display, such that the histological image is displayed together with the probability map, for example, overlaid on or side-by-side with the probability map.

[0050] Another aspect of the invention is a clinical network comprising: a computer device as described above; a data repository configured to store patient data records including histological images or sets of histological images; and a network connection enabling the transfer of patient data records or portions thereof between the computer device and the data repository. The clinical network may further include image acquisition devices, such as microscopes, operable to acquire histological images or sets of histological images and store them in records within the data repository.

[0051] It should be understood that, in at least some embodiments, one or more histological images are digital representations of two-dimensional images obtained from tissue sections using a microscope (particularly an optical microscope), which can be a conventional optical microscope, a confocal microscope, or any other type of microscope suitable for obtaining histological images of unstained or stained tissue samples. In the case of a set of histological images, these histological images can be a series of microscopic images obtained from adjacent sections (i.e., slices) of a tissue region, where each section can be stained differently. Attached Figure Description

[0052] In the following description, the invention will be further described by way of example only with reference to the exemplary embodiments shown in the accompanying drawings.

[0053] Figure 1A This is a schematic diagram of a neural network architecture according to one embodiment of the present invention.

[0054] Figure 1B It shows how Figure 1A In the neural network architecture, global and local feature maps are combined to generate a feature map for predicting the individual category of each pixel in the input image patch.

[0055] Figure 2 This is a color map showing the differences between block-level predictions generated by the following: Liu et al. (Patch A on the left is the original image, and patch B on the right is the CNN prediction, where dark red represents the predicted tumor region); and we used Figure 1A and Figure 1B The CNN predictions (the left tile C is the original image with pathologist's handwritten annotations (red) and CNN predictions (pink and yellow), the right tile D is our CNN prediction, where green is non-tumor, red (= pink in tile C) is invasive tumor, and blue (= yellow in tile C) is non-invasive tumor).

[0056] Figure 3This is a color illustration showing an example of an input RGB image patch (patch A on the left) and the final output tumor probability heatmap (patch B on the right). Patch A also shows an invasive tumor manually outlined by a pathologist (red outline), and further shows an overlay of our neural network's predictions (pink and yellow shaded areas), shown separately in patch B (reddish-brown and blue, respectively).

[0057] Figure 4 This is a flowchart illustrating the steps involved in training a CNN.

[0058] Figure 5 This is a flowchart illustrating the steps involved in making predictions using a CNN.

[0059] Figure 6 This is a block diagram of a TPU, which can be used to perform implementations. Figure 1A and Figure 1B The computations involved in the neural network architecture.

[0060] Figure 7 An exemplary computer network that can be used in conjunction with embodiments of the present invention is shown.

[0061] Figure 8 It can be used, for example, as a Figure 6 A block diagram of a host computer's computing device using a TPU. Detailed Implementation

[0062] In the following detailed description, specific details are set forth for purposes of explanation and not limitation in order to provide a better understanding of this disclosure. It will be apparent to those skilled in the art that this disclosure may be practiced in other embodiments that depart from these specific details.

[0063] We describe a computer-automated tumor detection method that automatically detects and delineates the nuclei of invasive and in situ breast cancer cells. This method is applied to a single input image (such as a WSI) or a set of input images (e.g., a WSI set). Each input image is a digitized histological image, such as a WSI. In the case of a set of input images, these images can be differently stained images of adjacent tissue sections. We use the term "staining" broadly to include staining with biomarkers as well as staining with conventional contrast-enhancing stains.

[0064] Because automated tumor delineation is much faster than manual delineation, it enables the processing of the entire image, rather than simply manually annotating selected patches from the image. Therefore, the proposed automated tumor delineation should allow pathologists to calculate the percentage of positive (or negative) tumor cells in an image, resulting in more accurate and reproducible results.

[0065] The proposed computer-automated method for tumor detection, delineation, and classification uses a convolutional neural network (CNN) to discover each nucleus pixel on the WSI, and then classifies each such pixel into one of the non-tumor categories and one of the multiple tumor categories (breast tumor category in our current implementation).

[0066] The neural network in our implementation is designed to resemble the VGG-16 architecture, which is available at http: / / www.robots.ox.ac.uk / ~vgg / research / very_deep / and described in Simonyan and Zisserman 2014[6], the entire contents of which are incorporated herein by reference.

[0067] The input image is a pathological image stained with any of several common staining agents, as discussed in more detail elsewhere in this document. For CNNs, image patches of a specific pixel size are extracted, such as 128×128, 256×256, 512×512, or 1024×1024 pixels. It should be understood that the image patches can be of any size and do not have to be square, but the number of pixels in the rows and columns of the patch must conform to 2^35. n , where n is a positive integer, because such a number is generally more suitable for direct digital processing by a suitable single CPU (Central Processing Unit), GPU (Graphics Processing Unit), or TPU (Tensor Processing Unit) or array thereof.

[0068] We note that "block" is a technical term used to refer to a portion of an image obtained from a WSI, which typically has a square or rectangular shape. In this regard, we note that a WSI may contain billions or more pixels (a billion-pixel image), so image processing will generally be applied to blocks of manageable size (e.g., approximately 500×500 pixels) for CNN processing. Therefore, the WSI will be processed based on segmenting it into blocks, analyzing the blocks with a CNN, and then reassembling the output (image) blocks into a probabilistic map of the same size as the WSI. The probabilistic map can then be overlaid, for example, semi-transparently overlaid on the WSI or a portion thereof, allowing both the pathological image and the probabilistic map to be viewed together. In this sense, the probabilistic map is used as a superimposed image on the pathological image. The blocks analyzed by the CNN may all have the same magnification, or may be a mixture of different magnifications, such as 5x, 20x, 50x, etc., and thus correspond to different sized physical regions of the sample tissue. These different magnifications can correspond to the physical magnification of the acquired WSI, or the effective magnification obtained by digitally downscaling a physical image at a higher magnification (i.e., higher resolution).

[0069] Figure 1AThis is a schematic diagram of our neural network architecture. Layers C1, C2...C10 are convolutional layers. Layers D1, D2, D3, D4, D5, and D6 are transposed convolutional (i.e., deconvolutional) layers. The lines connecting specific layers indicate skip connections between convolutional layer C and deconvolutional layer D. Skip connections allow for the combination of local features from larger, shallower layers (where "larger" and "shallow" refer to convolutional layers with lower indices) with global features from the last (i.e., smallest, deepest) convolutional layer. These skip connections provide a more accurate profile. Max-pooling layers (each used to reduce the width and height of a block to half) exist after layers C2, C4, and C7, but are not directly shown in the schematic; however, they are implicitly indicated by the correspondingly reduced size of the blocks. In some implementations of our neural network, max-pooling layers are replaced by 1×1 convolutions, resulting in a fully convolutional network.

[0070] The convolutional part of the neural network consists of the following layers in sequence: an input layer (RGB input image patch); two convolutional layers C1 and C2; a first max-pooling layer (not shown); two convolutional layers C3 and C4; a second max-pooling layer (not shown); three convolutional layers C5, C6, and C7; and a third max-pooling layer (not shown). In addition to normal connections to layers C5 and C8, skip connections are used to directly connect the outputs from the second and third max-pooling layers to the deconvolutional layers.

[0071] The outputs of the final convolutional layer C10, the second max-pooling layer (the layer after C4), and the third max-pooling layer (the layer after C7) are then concatenated to separate sequences of "deconvolutional layers." These deconvolutional layers enlarge them back to the same size as the input (image) patch, transforming the convolutional feature maps into feature maps with the same width and height as the input image patch and a number of channels (i.e., the number of feature maps) equal to the number of tissue categories to be detected (i.e., non-tumor types and one or more tumor types). For the second max-pooling layer, since only one stage of deconvolution is needed, it can be seen that it is directly linked to layer D6. For the third max-pooling layer, two stages of deconvolution are needed, reaching layer D5 via the intermediate deconvolutional layer D4. For the deepest convolutional layer C10, three stages of deconvolution are needed, reaching layer D3 via D1 and D2. The result is three arrays D3, D5, and D6 of the same size as the input patch.

[0072] Figure 1A The simplified version of the content shown (though it may not perform well) can omit the skip connections. In this case, layers D4, D5, and D6 will not exist, and the output block will be computed only from layer D3.

[0073] Figure 1B Showing more details Figure 1AThe final step in the neural network architecture is performed as follows: The global feature layer D3 and local feature layers D5 and D6 are combined to generate feature maps for each pixel of the predicted input image patch, representing a separate category. Specifically, Figure 1B This demonstrates how the last three transposed convolutional layers, D3, D5, and D6, are processed into tumor category output blocks.

[0074] Now, we discuss how the above method differs from known CNNs currently used in digital pathology. Such known CNNs assign a category selected from multiple available categories to each image patch. Examples of this type of CNN are given in the papers of Wang et al. 2016[1], Liu et al. 2017[4], Cruz-Roa et al. 2017[8] and Vandenberghe et al. 2017[9]. However, what we have just described is that, in a given image patch, a category selected from multiple available categories is assigned to each pixel. Therefore, instead of generating a single category label for each image patch, our neural network outputs a category label for each individual pixel of a given patch. Our output patch has a one-to-one pixel-to-pixel correspondence with the input patch, such that each pixel in the output patch is assigned one of multiple available categories (non-tumor, tumor 1, tumor 2, tumor 3, etc.).

[0075] In such known CNNs [1,4,8,9], a series of convolutional layers are used to assign a single class to each block, followed by one or more fully connected layers, and then an output vector with as many values ​​as the class to be detected. The predicted class is determined by the position of the maximum value in the output vector.

[0076] To predict the category of each pixel, our CNN uses a different architecture after the convolutional layers. Instead of a series of fully connected layers, we follow the convolutional layers with a series of transposed convolutional layers. The fully connected layers are removed in this architecture. Each transposed layer doubles the width and height of the feature map while halving the number of channels. In this way, the feature map is magnified back to the size of the input block.

[0077] In addition, to improve prediction, we used skip connections as described in Long et al. 2015

[10] , the entire contents of which are incorporated herein by reference.

[0078] Skip connections use shallower features to improve the coarse predictions obtained by amplification from the final convolutional layer C10. Figure 1A The local features from skip connections contained in layers D5 and D6 are compared with those contained in the final convolutional layer. Figure 1A The global features in layer D3 are amplified, and the resulting features are concatenated. Then, as... Figure 1B As shown, the global feature layer D3 and the local feature layers D5 and D6 are connected into a combined layer.

[0079] from Figure 1B The number of channels in the combined layer (or alternatively, directly from the final deconvolutional layer D3 without skip connections) is reduced to match the number of classes through 1×1 convolutions in the combined layer. A softmax operation on this classification layer then converts the values ​​in the combined layer into probabilities. The output block layer has a size of N×N×K, where N is the width and height of the input block in pixels, and K is the number of detected classes. Therefore, for any pixel P in the image block, there exists an output vector V of size K. A unique class can then be assigned to each pixel P by the position of its maximum value in its corresponding vector V.

[0080] Therefore, CNN labels each pixel as either non-cancerous or belonging to one or more of several different cancer (tumor) types. Breast cancer is of particular interest, but the method is also applicable to histological images of other cancers, such as bladder cancer, colon cancer, rectal cancer, kidney cancer, blood cancer (leukemia), endometrial cancer, lung cancer, liver cancer, skin cancer, pancreatic cancer, prostate cancer, brain cancer, spinal cancer, and thyroid cancer.

[0081] To our knowledge, this is the first time this method has been used in CNNs to delineate and classify breast cancer. We have determined that this method improves performance compared to previous computer-automated methods we know of because the contours more closely resemble those drawn by pathologists and considered ground truth.

[0082] Our specific neural network implementation is configured to operate on input images with a specific fixed pixel size. Therefore, as a preprocessing step, for both training and prediction, blocks with the desired pixel size (e.g., N×N×n pixels) are extracted from the WSI, where n = 3 when the WSI is a color image acquired by a conventional visible light microscope, with three pixels associated with the three primary colors (typically RGB) at each physical location (as further described below, "n" can be 3 times the number of synthesized WSIs when two or more color WSIs are combined). Moreover, in the case of a single monochrome WSI, the value of "n" is 1. To accelerate training, the input blocks are also centered and normalized at this stage.

[0083] Our preferred approach is to process the entire WSI, or at least the entire tissue-containing region of the WSI; therefore, in our case, a block is a tile that at least covers the entire tissue region of the WSI. These tiles can be adjacent without overlap, or have overlapping edge regions, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 pixels wide, so that the output blocks of the CNN can be stitched together taking into account any differences. However, if desired, our method can also be applied to random samples of blocks with the same or different magnifications on the WSI, as in the prior art, or can be performed by a pathologist.

[0084] Our neural network is similar in design to the VGG-16 architecture of Simonyan and Zisserman 2014[6]. It uses very small 3×3 convolutional kernels in all convolutional filters. Max pooling is performed using a small 2×2 window and a stride of 2. In contrast to the VGG-16 architecture which has a series of fully connected layers after the convolutional layers, we follow the convolutional layers with a series of “deconvolutions” (more precisely, transposed convolutions) to generate a segmentation mask. This type of upsampling for semantic segmentation has previously been used by Long et al. 2015

[10] for natural image processing, the entire contents of which are incorporated herein by reference.

[0085] Each deconvolutional layer magnifies the input feature map by a factor of two in both width and height. This counteracts the shrinking effect of max pooling layers and produces a class feature map of the same size as the input image. The outputs from each convolutional and deconvolutional layer are transformed by non-linear activation layers. Currently, the non-linear activation layers use the rectified function RoLU(x) = max(0, x). Different activation functions, such as ReLU, leaky ReLU, ELU, etc., can be used as needed.

[0086] The proposed method can be applied without modifying any desired number of tissue categories. The only constraint is the availability of appropriate training data, which has already been classified in a manner expected to be replicated in the neural network. More examples of breast pathology are invasive lobular carcinoma or invasive ductal carcinoma, where a single invasive tumor category in previous examples can be replaced by multiple invasive tumor categories.

[0087] A softmax regression layer (i.e., a multinomial logistic regression layer) is applied to each of the channel blocks to convert the values ​​in the feature map into probabilities.

[0088] After the final transformation by softmax regression, the value at position (x,y) of channel C in the final feature map contains the probability P(x,y), that is, the pixel at position (x,y) in the input image patch belongs to the tumor type detected by channel C.

[0089] It should be understood that the number of convolutional and deconvolutional layers can be increased or decreased as needed, and is limited by the memory of the hardware running the neural network.

[0090] We use mini-batch gradient descent to train the neural network. The learning rate is reduced from an initial rate of 0.1 using exponential decay. We prevent overfitting by using a "dropout" procedure described by Srivastava et al. 2014

[2017] , the entire contents of which are incorporated herein by reference. The network can be trained on a GPU, CPU, or FPGA using any of several available deep learning frameworks. For our current implementation, we use Google Tensorflow, but the same neural network can be implemented in another deep learning framework, such as Microsoft CNTK.

[0091] The neural network outputs a probability map of size N×N×K, where N is the width and height of the input block (in pixels) and K is the number of detected classes. These output blocks are concatenated to form a probability map of size W×H×K, where W and H are the width and height of the original WSI before being divided into blocks.

[0092] Then, by recording the category index with the highest probability at each position (x,y) in the label image, the probability map can be folded into a W×H label image.

[0093] In its current implementation, our neural network assigns each pixel to one of three categories: non-tumor, invasive tumor, and in situ tumor.

[0094] When using multiple tumor categories, the output image can be post-processed into a simpler binary classification of non-tumor and tumor, combining multiple tumor categories. Binary classification can be used as an option when creating images from underlying data, while preserving the multi-tumor classification within the saved data.

[0095] Although the above description of a particular implementation of the invention focuses on a specific method using CNNs, it should be understood that our method can be implemented in a variety of different types of convolutional neural networks. Generally speaking, any neural network that uses convolution to detect increasingly complex features and then uses transposed convolution (“deconvolution”) to magnify the feature maps back to the width and height of the input image should be suitable.

[0096] Example 1

[0097] Figure 2It is in color and shows the difference between block-level predictions (such as those generated by Google’s CNN scheme for the Camelyon competition (Liu et al. 2017[4]) and pixel-level predictions generated by our CNN).

[0098] Figure 2 Plots A and B in the image are from Liu et al. 2017[4]. Figure 7 The images are copied from the originals, while tiles C and D are comparable tiles according to an example of the invention.

[0099] Patch A is a patch from WSI stained with H&E, where the larger dark purple cell clusters in the lower right quadrant are tumors, while the smaller dark purple cells are lymphocytes.

[0100] Plot B is a tumor probability heatmap generated by CNN by Liu et al. in 2017[4]. The authors pointed out that the map accurately identified tumor cells while ignoring connective tissue and lymphocytes.

[0101] Tile C is the original image tile from an example WSI that applies the CNN method embodying the present invention. In addition to the original image, tile C also shows a contour (solid red border) hand-drawn by a pathologist. Furthermore, referring to tile D, tile C also shows the results from our CNN method (the first pink shaded area with a pink border corresponds to the first tumor type, i.e., the tumor type shown in red in tile D; the second yellow shaded area with a pink border corresponds to the second tumor type, i.e., the tumor type shaded in blue in tile D).

[0102] Tile D is a tumor probability heatmap generated by our CNN. It demonstrates how our pixel-level prediction method produces regions with smooth peripheral contours. For our heatmap, different (arbitrarily chosen) colors indicate different categories: green represents non-tumor, red represents the first tumor type, and blue represents the second tumor type.

[0103] Example 2

[0104] Figure 3 It is in color, and an example is shown of the input RGB image patch (patch A on the left) and the final output tumor probability heatmap (patch B on the right).

[0105] Tile A also shows an invasive tumor manually outlined by a pathologist (red outline), and also shows an overlay of our neural network's predictions (pink and yellow shaded areas) shown separately in Tile B.

[0106] Tile B is a tumor probability heatmap generated by our CNN. For our heatmap, different (arbitrarily chosen) colors indicate different categories: green represents non-tumor, reddish-brown represents invasive tumors (shown as pink in tile A), and blue represents in situ tumors (shown as yellow in tile A). Again, it can be seen how our pixel-level prediction method produces regions with smooth peripheral contours. Furthermore, it can be seen how CNN predictions are compatible with manual labeling by pathologists. Additionally, the CNN provides further differentiation between invasive and non-invasive (in situ) tissues, which is not performed by pathologists and is an inherent part of our multi-channel CNN design, which can be programmed and trained to classify tissues into any number of different types as needed and clinically relevant.

[0107] Acquisition and Image Processing

[0108] This method begins with the tissue sample having been sliced, i.e., cut into thin sections, and adjacent sections having been stained with different staining agents. Because the sections are thin, adjacent sections will have very similar tissue structures, but they are not exactly the same because they have different layers.

[0109] For example, there might be five adjacent sections, each stained with a different agent, such as ER, PR, p53, HER2, H&E, and Ki-67. Microscopic images of each section are then acquired. Although adjacent sections will have very similar tissue shapes, the staining will highlight distinct features, such as cell nuclei, cytoplasm, and all features revealed through general contrast enhancement.

[0110] The different images are then preprocessed by aligning, warping, or otherwise modifying them to map the coordinates of any given feature on one image to the same feature on other images. This mapping will handle any differences between the images caused by factors such as slightly different magnifications, orientation differences due to differences in slide alignment in the microscope, or differences in mounting tissue sections onto slides.

[0111] It should be noted that by using the coordinate mapping between different WSIs in a set of adjacent slices that are stained differently, these WSIs can be merged into a single synthetic WSI, from which synthetic blocks can be extracted for CNN processing, where such synthetic blocks will have a size of N×N×3m, where "m" is the number of synthetic WSIs that form the set.

[0112] Then, some standard processing is performed on the image. These image processing steps can be performed at the WSI level or at the level of individual image patches. If the CNN is configured to operate on monochrome images instead of color images, the image can be converted from a color image to a grayscale image. The image can be modified by applying a contrast enhancement filter. Segmentation can then be performed to identify common organizing regions in the image set or simply to reject backgrounds unrelated to organizing regions. Segmentation can involve any or all of the following image processing techniques:

[0113] Analysis based on variance to identify seed tissue regions

[0114] Adaptive thresholding

[0115] Morphological operations (e.g., blob analysis)

[0116] Contour recognition

[0117] Contour merging based on proximity heuristics

[0118] Calculation of invariant image moments

[0119] Edge extraction (e.g., Sobel edge detection)

[0120] Curvature Flow Filtering

[0121] Histogram matching to eliminate intensity variations between consecutive slices

[0122] Multi-resolution rigid body / affine image registration (gradient descent optimizer)

[0123] Non-rigid body deformation / transformation

[0124] Superpixel clustering

[0125] It should also be understood that the image processing steps described above can be performed on the WSI or on individual blocks after block extraction. In some cases, it may be useful to perform the same type of image processing before and after block extraction. That is, some image processing can be performed on the WSI before block extraction, and other image processing can be performed on the blocks after they are extracted from the WSI.

[0126] These image processing steps are described by way of example and should not be construed as limiting the scope of the invention in any way. For example, if there is sufficient processing power, a CNN can directly process color images.

[0127] Training and prediction

[0128] Figure 4 This is a flowchart illustrating the steps involved in training a CNN.

[0129] In step S40, training data containing WSIs for processing is retrieved. These WSIs have been annotated by clinicians to detect, delineate, and classify tumors. The clinicians' annotations represent ground truth data.

[0130] In step S41, the WSI is decomposed into image patches, which are the input image patches for the CNN. That is, image patches are extracted from the WSI.

[0131] In step S42, the image patch is preprocessed as described above. (Alternatively, the WSI can be preprocessed as described above before step S41.)

[0132] In step S43, initial values ​​are set for the CNN weights (i.e., inter-layer weights).

[0133] In step S44, as referenced above... Figure 1A and Figure 1B To be further described, each of the batch of input image patches is fed into the CNN and processed to discover, delineate, and classify patches on a pixel-by-pixel basis. The term "delineate" here is not necessarily the correct term in a strictly technical sense, because our method identifies each tumor (or tumor type) pixel, so it might be more accurate to say that the CNN determines the tumor region for each tumor type.

[0134] In step S45, the CNN output image patches are compared with ground truth data. This can be done on a patch-by-patch basis. Alternatively, if patches covering the entire WSI have been extracted, this can be done at the WSI level or within a sub-region of the WSI consisting of a batch of consecutive patches (e.g., a quadrant of the WSI). In such variations, the output image patches can be reassembled into a probability map of the entire WSI or a consecutive portion thereof, and the probability map can be compared with ground truth data by a computer and visually by a user (if the probability map is presented on a display as a semi-transparent overlay of the WSI).

[0135] In step S46, the CNN then learns from this comparison and updates the CNN weights, for example, using gradient descent. Thus, the learning is fed back into the repeated processing of the training data via a return loop in the process flow, such as... Figure 4 As shown, this allows CNN weights to be optimized.

[0136] After training, the CNN can be applied to WSI independently of any ground truth data, i.e., for prediction in real time.

[0137] Figure 5 This is a flowchart illustrating the steps involved in making predictions using a CNN.

[0138] In step S50, one or more WSIs are retrieved, for example, from a Laboratory Information System (LIS) or other histological data repository for processing. For example, the WSIs are preprocessed as described above.

[0139] In step S51, image blocks are extracted from the WSI or each WSI. The blocks may cover the entire WSI, or they may be randomly or non-randomly selected.

[0140] In step S52, the image patch is preprocessed, for example, as described above.

[0141] In step S53, as referenced above... Figure 1A and Figure 1B Further described, each of a batch of input image patches is fed into the CNN and processed to discover, delineate, and classify patches on a pixel-by-pixel basis. The output patches can then be reassembled into a probability map from which the WSI of the input image patches are extracted. For example, the probability map can be compared to the WSI by a computer device in digital processing and by a user visually (if the probability map is presented on a display as a semi-transparent overlay on or alongside the WSI).

[0142] In step S54, tumor regions are screened to exclude tumors that may be false positives, such as regions that are too small or regions that may be edge artifacts.

[0143] In step S55, the scoring algorithm is run. The scores are cell-specific, and the scores can be aggregated for each tumor and / or further aggregated for the WSI (or a subregion of the WSI).

[0144] In step S56, the results are presented to a pathologist or other relevant skilled clinician for diagnosis, for example by displaying the annotated WSI on a suitable high-resolution monitor.

[0145] In step S57, the CNN results (i.e., probability plot data and optionally metadata related to the CNN parameters, along with any additional diagnostic information added by the pathologist) are saved in a manner linked to a patient data file containing a WSI or set of WSIs that has been processed by the CNN. Thus, the patient data file in the LIS or other histological data repository will be supplemented with the CNN results.

[0146] Computing platform

[0147] The proposed image processing can be performed on a variety of computing architectures, particularly those optimized for neural networks, which can be based on CPUs, GPUs, TPUs, FPGAs, and / or ASICs. In some implementations, the neural network is implemented using the Google Tensorflow software library running on Nvidia GPUs (such as the Tesla K80 GPU) from Nvidia Corporation in Santa Clara, California. In other implementations, the neural network can run on a general-purpose CPU. Faster processing is possible with processors specifically designed to perform CNN computations, such as the TPU disclosed in Jouppi et al. 2017

[11] , the entire contents of which are incorporated herein by reference.

[0148] Figure 6 The TPU of Jouppi et al. 2017

[11] is shown, which is a simplified reproduction of Jouppi’s Figure 1. The TPU 100 has a systolic matrix multiplication unit (MMU) 102, which contains a 256×256 MAC capable of performing 8-bit multiplication and addition on signed or unsigned integers. Weights of the MMU are provided through a weight FIFO buffer 104, which in turn reads weights from memory 106 in the form of an off-chip 8GB DRAM via a suitable memory interface 108. A uniform buffer (UB) 110 is provided to store intermediate results. The MMU 102 is connected to receive input from the weight FIFO interface 104 and the UB 110 (via a systolic data setting unit 112) and outputs the 16-bit product processed by the MMU to an accumulator unit 114. An activation unit 116 performs a nonlinear function on the data stored in the accumulator unit 114. After further processing by normalization unit 118 and pooling unit 120, the intermediate results are sent to UB 110 for re-providalization to MMU 102 via data setting unit 112. Pooling unit 120 can perform max pooling or average pooling as needed. Programmable DMA controller 122 transfers data to or from the TPU's host computer and UB 110. TPU instructions are sent from the host computer to controller 122 via host interface 124 and instruction buffer 126.

[0149] It should be understood that the computing power used to run neural networks (whether based on CPU, GPU or TPU) can be hosted locally in a clinical network, such as the network described below, or remotely hosted in a data center.

[0150] The proposed computer-automated method operates within the context of a Laboratory Information System (LIS), which is typically part of a larger clinical network environment, such as a Hospital Information System (HIS) or a Picture Archiving and Communication System (PACS). In the LIS, the WSI (Waste Slip Intake) is stored in a database, typically a patient information database containing the electronic medical records of each patient. The WSI is taken from stained tissue samples mounted on slides bearing printed barcode labels. These barcode labels tagged the WSI with appropriate metadata, as the microscope used to acquire the WSI is equipped with a barcode reader. From a hardware perspective, the LIS will be a conventional computer network, such as a Local Area Network (LAN) with the required wired and wireless connections.

[0151] Figure 7 An exemplary computer network that can be used in conjunction with embodiments of the present invention is illustrated. Network 150 includes a LAN in hospital 152. Hospital 152 is equipped with a plurality of workstations 154, each of which can access a hospital computer server 156 with an associated storage device 158 via the LAN. LIS, HIS, or PACS archives are stored on storage device 158, making the data in the archives accessible from any workstation 154. One or more of the workstations 154 can access a graphics card and software for computer implementation of the image generation method as described above. The software may be stored locally at workstation 154 or at each workstation, or it may be stored remotely and downloaded to workstation 154 via network 150 when needed. In other instances, the methods embodying the present invention may be executed on the computer server with workstation 154 operating as a terminal. For example, a workstation may be configured to receive user input defining a desired dataset of histological images and to display the resulting images while CNN analysis is performed elsewhere in the system. Furthermore, a plurality of histological and other medical imaging devices 160, 162, 164, 166 are connected to hospital computer server 156. Image data collected using devices 160, 162, 164, and 166 can be directly stored in LIS, HIS, or PACS archives on storage device 156. Therefore, histological images can be viewed and processed immediately after the corresponding histological image data is recorded. The local area network is connected to the Internet 168 via a hospital Internet server 170, which allows remote access to LIS, HIS, or PACS archives. This can be used for remote data access and data transfer between hospitals (e.g., if a patient is transferred), or to allow for external research.

[0152] Figure 8This is a block diagram illustrating an exemplary computing device 500 that can be used in conjunction with the various embodiments described herein. For example, computing device 500 can be used as a computing node in the aforementioned LIS or PACS system, for example, in conjunction with a suitable GPU or Figure 6 The TPU shown is the host computer from which CNN processing is performed.

[0153] The computing device 500 may be a server or any conventional personal computer, or any other processor-enabled device capable of wired or wireless data communication. As will be apparent to those skilled in the art, other computing devices, systems, and / or architectures, including those not capable of wired or wireless data communication, may also be used.

[0154] Computing device 500 preferably includes one or more processors, such as processor 510. Processor 510 may be, for example, a CPU, GPU, TPU, or an array or combination thereof, such as a CPU and TPU combination or a CPU and GPU combination. Additional processors may be provided, such as auxiliary processors for managing input / output, auxiliary processors for performing floating-point mathematical operations (e.g., TPU), dedicated microprocessors (e.g., digital signal processors, image processors) with an architecture suitable for rapidly executing signal processing algorithms, slave processors subordinate to the main processing system (e.g., back-end processors), additional microprocessors or controllers for dual-processor or multi-processor systems, or coprocessors. Such auxiliary processors may be discrete processors or may be integrated with processor 510. Examples of CPUs that may be used with computing device 500 are Pentium processors, Core i7 processors, and Xeon processors, all of which are available from Intel Corporation of Santa Clara, California. An example GPU that may be used with computing device 500 is the Tesla K80 GPU from NVIDIA Corporation of Santa Clara, California.

[0155] Processor 510 is connected to communication bus 505. Communication bus 505 may include a data channel for facilitating information transfer between storage devices and other peripheral components of computing device 500. Communication bus 505 may also provide a set of signals for communicating with processor 510, including a data bus, an address bus, and a control bus (not shown). Communication bus 505 may include any standard or non-standard bus architecture, such as Industry Standard Architecture (ISA), Extended Industry Standard Architecture (EISA), Microchannel Architecture (MCA), Peripheral Component Interconnect (PCI) local bus, or bus architectures of standards issued by the Institute of Electrical and Electronics Engineers (IEEE) (including IEEE 488 Universal Interface Bus (GPIB), IEEE 696 / S-100, etc.).

[0156] The computing device 500 preferably includes a main memory 515, and may also include a secondary memory 520. The main memory 515 provides storage for instructions and data for programs executed on the processor 510, such as one or more of the functions and / or modules discussed above. It should be understood that the computer-readable program instructions stored in the memory and executed by the processor 510 may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written and / or compiled in any combination of one or more programming languages, including but not limited to Smalltalk, C / C++, Java, JavaScript, Perl, Visual Basic, .NET, etc. The main memory 515 is typically a semiconductor-based memory, such as dynamic random access memory (DRAM) and / or static random access memory (SRAM). Other semiconductor-based memory types include, for example, synchronous dynamic random access memory (SDRAM), Rambus dynamic random access memory (RDRAM), ferroelectric random access memory (FRAM), etc., including read-only memory (ROM).

[0157] Computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of execution on a remote computer or server, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet through an Internet service provider).

[0158] Secondary storage 520 may optionally include internal storage 525 and / or removable media 530. Removable media 530 can be read and / or written in any known manner. Removable storage media 530 may be, for example, a magnetic tape drive, an optical disc (CD) drive, a digital versatile optical disc (DVD) drive, other optical drives, flash memory drives, etc.

[0159] Removable storage medium 530 is a non-transitory computer-readable medium on which computer-executable code (i.e., software) and / or data is stored. The computer software or data stored on removable storage medium 530 is read into computing device 500 for execution by processor 510.

[0160] Secondary memory 520 may include other similar elements for allowing computer programs or other data or instructions to be loaded into computing device 500. Such means may include, for example, external storage medium 545 and a communication interface 540 that allows software and data to be transferred from external storage medium 545 to computing device 500. Examples of external storage medium 545 may include external hard disk drives, external optical drives, external magneto-optical drives, etc. Other examples of secondary memory 520 may include semiconductor-based memories such as programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), or flash memory (a block-oriented memory similar to EEPROM).

[0161] As described above, computing device 500 may include communication interface 540. Communication interface 540 allows software and data to be transferred between computing device 500 and external devices (e.g., printers), networks, or other information sources. For example, computer software or executable code can be transferred from a network server to computing device 500 via communication interface 540. Examples of communication interface 540 include a built-in network adapter, network interface card (NIC), PCMCIA network card, card bus network adapter, wireless network adapter, Universal Serial Bus (USB) network adapter, modem, network interface card (NIC), wireless data card, communication port, infrared interface, IEEE 1394 FireWire, or any other device capable of interfacing system 550 with a network or another computing device. The communication interface 540 preferably implements industrially promulgated protocol standards, such as Ethernet IEEE 802 standard, Fibre Channel, Digital Subscriber Line (DSL), Asynchronous Digital Subscriber Line (ADSL), Frame Relay, Asynchronous Transfer Mode (ATM), Integrated Digital Services Network (ISDN), Personal Communication Services (PCS), Transmission Control Protocol / Internet Protocol (TCP / IP), Serial Line Internet Protocol / Point-to-Point Protocol (SLIP / PPP), etc., but it can also implement customized or non-standard interface protocols.

[0162] Software and data transmitted via communication interface 540 are typically in the form of electrical communication signals 555. These signals 555 can be provided to communication interface 540 via communication channel 550. In one embodiment, communication channel 550 can be a wired or wireless network, or any other type of communication link. Communication channel 550 carries signals 555 and can be implemented using a variety of wired or wireless communication devices, including wires or cables, optical fibers, conventional telephone lines, cellular telephone links, wireless data communication links, radio frequency (“RF”) links, or infrared links, to name just a few.

[0163] Computer-executable code (i.e., computer programs or software) is stored in main memory 515 and / or secondary memory 520. Computer programs can also be received via communication interface 540 and stored in main memory 515 and / or secondary memory 520. When executed, such computer programs enable computing device 500 to perform various functions of the disclosed embodiments as described elsewhere herein.

[0164] In this document, the term “computer-readable medium” is used to refer to any non-transitory computer-readable storage medium used to provide computer-executable code (e.g., software and computer programs) to computing device 500. Examples of such media include main memory 515, secondary memory 520 (including internal memory 525, removable medium 530, and external storage medium 545), and any peripheral device (including network information server or other network device) communicatively coupled to communication interface 540. These non-transitory computer-readable media are means for providing executable code, programming instructions, and software to computing device 500.

[90] In embodiments implemented using software, the software may be stored on a computer-readable medium and loaded into computing device 500 via removable medium 530, I / O interface 535, or communication interface 540. In such embodiments, the software is loaded into computing device 500 in the form of an electrical communication signal 555. When executed by processor 510, the software preferably causes processor 510 to perform the features and functions described elsewhere herein.

[0165] I / O interface 535 provides an interface between one or more components of computing device 500 and one or more input and / or output devices. Examples of input devices include, but are not limited to, keyboards, touchscreens or other touch-sensitive devices, biometric sensors, computer mice, trackballs, pen-based pointer devices, etc. Examples of output devices include, but are not limited to, cathode ray tube (CRT), plasma displays, light-emitting diode (LED) displays, liquid crystal displays (LCDs), printers, vacuum fluorescent displays (VFDs), surface-conducting electron emission displays (SEDs), field emission displays (FEDs), etc.

[0166] The computing device 500 also includes optional wireless communication components that facilitate wireless communication via voice and / or data networks. The wireless communication components include an antenna system 570, a radio system 565, and a baseband system 560. In the computing device 500, radio frequency (RF) signals are transmitted and received over the air via the antenna system 570, under the control of the radio system 565.

[0167] Antenna system 570 may include one or more antennas and one or more multiplexers (not shown), which perform switching functions to provide transmission and reception signal paths to antenna system 570. In the reception path, received RF signals may be coupled from the multiplexer to a low-noise amplifier (not shown), which amplifies the received RF signals and transmits the amplified signals to radio system 565.

[0168] Radio system 565 may include one or more radio devices configured to communicate on various frequencies. In one embodiment, radio system 565 may combine a demodulator (not shown) and a modulator (not shown) in an integrated circuit (IC). The demodulator and modulator may also be separate components. In the input path, the demodulator strips the RF carrier signal, leaving a baseband receive audio signal, which is transmitted from radio system 565 to baseband system 560.

[0169] If the received signal contains audio information, the baseband system 560 decodes the signal and converts it into an analog signal. The signal is then amplified and sent to a speaker. The baseband system 560 also receives analog audio signals from a microphone. These analog audio signals are converted into digital signals and encoded by the baseband system 560. The baseband system 560 further encodes the digital signals for transmission and generates a baseband transmission audio signal that is routed to the modulator section of the radio system 565. The modulator mixes the baseband transmission audio signal with an RF carrier signal to generate an RF transmission signal, which is routed to the antenna system 570 and can pass through a power amplifier (not shown). The power amplifier amplifies the RF transmission signal and routes it to the antenna system 570, where the signal is switched to the antenna port for transmission.

[0170] The baseband system 560 is also communicatively coupled to a processor 510, which may be a central processing unit (CPU). The processor 510 can access data storage areas 515 and 520. The processor 510 is preferably configured to execute instructions (i.e., computer programs or software) that can be stored in main memory 515 or secondary memory 520. Computer programs can also be received from the baseband processor 560 and stored in main memory 510 or secondary memory 520, or executed upon receipt. Such computer programs, when executed, enable the computing device 500 to perform various functions of the disclosed embodiments. For example, data storage areas 515 or 520 may include various software modules.

[0171] The computing device also includes a display 575 directly attached to the communication bus 505, which can be provided to replace or be added to any display connected to the aforementioned I / O interface 535.

[0172] For example, various implementation schemes can also be implemented primarily in hardware using components such as application-specific integrated circuits (ASICs), programmable logic arrays (PLAs), or field-programmable gate arrays (FPGAs). The implementation methods of the hardware state machines capable of performing the functions described herein are also readily apparent to those skilled in the art. Various implementation schemes can also be implemented using a combination of hardware and software.

[0173] Furthermore, those skilled in the art will understand that the various illustrative logic blocks, modules, circuits, and method steps described in conjunction with the foregoing figures and embodiments disclosed herein can generally be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above in general terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in various ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention. Additionally, the functional grouping within modules, blocks, circuits, or steps is for ease of description. A particular function or step may be moved from one module, block, or circuit to another without departing from the invention.

[0174] Furthermore, the various illustrative logic blocks, modules, functions, and methods described in conjunction with the embodiments disclosed herein can be implemented or performed using a general-purpose processor, digital signal processor (DSP), ASIC, FPGA, or other programmable logic device designed to perform the functions described herein, discrete gate or transistor logic, discrete hardware components, or any combination thereof. The general-purpose processor may be a microprocessor, but alternatively, the processor may be any processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration.

[0175] Furthermore, the steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium, including network storage media. An exemplary storage medium can be coupled to the processor, allowing the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be integrated with the processor. The processor and storage medium can also reside in an ASIC.

[0176] As mentioned herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0177] Any software component described herein may take many forms. For example, a component may be a standalone software package, or it may be a package incorporated as a "tool" into a larger software product. A component may be downloaded from a network (e.g., a website) as a standalone product or add-on to be installed in an existing software application. A component may also be a client-server software application, a web-enabled software application, and / or a mobile application.

[0178] Embodiments of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0179] Computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other apparatus to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions that implement aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.

[0180] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device, thereby producing a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0181] The flowcharts and block diagrams shown illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing one or more specified logical functions. In some alternative implementations, the functions indicated in the boxes may occur in a different order than those shown in the diagrams. For example, depending on the functions involved, two boxes shown consecutively may actually be executed substantially simultaneously, or these boxes may sometimes be executed in reverse order. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions or executes a combination of dedicated hardware and computer instructions.

[0182] The devices and methods embodying this invention can be hosted and delivered in a cloud computing environment. Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.

[0183] The features are as follows:

[0184] On-demand self-service: Cloud consumers can unilaterally and automatically provision computing power such as server time and network storage on demand without having to interact with service providers.

[0185] Extensive network access: Capabilities can be accessed via the network and through standard mechanisms that facilitate use across various thin-client or thick-client platforms (e.g., mobile phones, laptops, and personal digital assistants (PDAs)).

[0186] Resource pooling: A provider's computing resources are pooled to serve multiple consumers through a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated on demand. There are situations where consumers typically cannot control or are unaware of the exact location of the resources provided, but may be able to specify location independence at a higher level of abstraction (e.g., country, state, or data center).

[0187] Rapid flexibility: Capacity can be rapidly and flexibly (sometimes automatically) provisioned to expand quickly and can be quickly released to shrink quickly. For consumers, the capacity available for provisioning often appears unlimited, and any quantity can be purchased at any time.

[0188] Measuring services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a certain level of abstraction appropriate to service types (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.

[0189] The service model is as follows:

[0190] Software as a Service (SaaS): The capability offered to consumers is the ability to use applications running on cloud infrastructure by the provider. These applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., web-based email). Aside from potentially limited user-specific application configuration settings, consumers neither manage nor control the underlying cloud infrastructure, including the network, servers, operating system, storage, and even the individual application capabilities.

[0191] Platform as a Service (PaaS): This provides consumers with the ability to deploy consumer-created or acquired applications on cloud infrastructure, using programming languages ​​and tools supported by the provider. Consumers neither manage nor control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the deployed applications and may also have control over the configuration of the application hosting environment.

[0192] Infrastructure as a Service (IaaS): This provides consumers with provisioned processing, storage, networking, and other basic computing resources that they can deploy and run, including operating systems and applications. Consumers neither manage nor control the underlying cloud infrastructure, but they have control over the operating system, storage, and deployed applications, and may have limited control over selected network components (e.g., host firewalls).

[0193] The deployment model is as follows:

[0194] Private cloud: A cloud infrastructure that operates exclusively for a single organization. It can be managed by that organization or a third party, and can exist either on-premise or off-premise.

[0195] Community cloud: Cloud infrastructure shared by several organizations to support a specific community with common concerns (e.g., mission, security requirements, policy and compliance considerations). It can be managed by the organization or a third party and can exist internally or externally.

[0196] Public cloud: Cloud infrastructure that is available to the public or large industrial groups and is owned by organizations that sell cloud services.

[0197] Hybrid cloud: A cloud infrastructure consisting of two or more clouds (private, community, or public) that remain distinct entities but are bound together by standardized or proprietary technologies that enable data and applications to be ported together (e.g., cloud bursts for load balancing between clouds).

[0198] Cloud computing environments are service-oriented, emphasizing statelessness, loose coupling, modularity, and semantic interoperability. At its core is the infrastructure comprising a network of interconnected nodes.

[0199] Those skilled in the art will recognize that many improvements and modifications can be made to the foregoing exemplary embodiments without departing from the scope of this disclosure.

[0200] References

[0201] Wang, Dayong, Aditya Khosla, Rishab Gargeya, Humayun Irshad, and Andrew HBeck. 2016. “Deep Learning for Identifying Metastatic Breast Cancer.” ArXivPreprint ArXiv: 1606.05718.

[0202] US2015213302A1(Case Western Reserve University)

[0203] Le Hou, Dimitris Samaras, Tahsin M. Kurc, Yi Gao, James E. Davis, and Joel H. Saltz "Patch-based Convolutional Neural Network for Whole Slide TissueImageClassification", Proc. IEEE Comput Soc Conf Comput Vis PatternRecognit. 2016, pages 2424-2433. doi: 10.1109 / CVPR

[0204] Liu,Yun,Krishna Gadepalli,Mohammad Norouzi,George E Dahl,TimoKohlberger,Aleksey Boyko,Subhashini Venugopalan,et al.2017.“DetectingCancerMetastases on Gigapixel Pathology Images.”ArXiv Preprint ArXiv:1703.02442.

[0205] Babak Ehteshami Bejnordi,Guido Zuidhof,Maschenka Balkenhol,MeykeHermsen,Peter Bult,Bram van Ginneken,Nico Karssemeijer,Geert Litjens,Jeroenvan der Laak″Context-aware stacked convolutional neural networksforclassification of breast carcinomas in whole-slide histopathology images″10May 2017,ArXiv:1705.03678v1

[0206] Simonyan and Zisserman.2014.“Very Deep Convolutional NetworksforLarge-Scale Image Recognition.”ArXiv Preprint ArXiv:1409.1556

[0207] Srivastava,Nitish,Geoffrey E Hinton,Alex Krizhevsky,Ilya Sutskever,and Ruslan Salakhutdinov.2014.“Dropout:A Simple Way to Prevent NeuralNetworksfrom Overfitting.”Journal of Machine Learning Research vol.15(1):pages 1929-58.

[0208] Angel Cruz-Roa,Hannah Gilmore,Ajay Basavanhally,Michael Feldman,Shridar Ganesan,Natalie N.C.Shih,John Tomaszewski,Fabio A.González&AnantMadabhushi.″A Deep Learning approach for quantifying tumor extent″ScientificReports 7,Article number:46450(2017),doi:10.1038 / srep46450

[0209] Michel E.Vandenberghe,Marietta L.J.Scott,Paul W.Scorer,MagnusDenisBalcerzak&Craig Barker.″Relevance of deep learning tofacilitate thediagnosis of HER2 status in breast cancer″.Scientific Reports7,Articlenumber:45938(2017),doi:10.1038 / srep45938

[0210] Long,Jonathan,Evan Shelhamer,and Trevor Darrell.2015.“FullyConvolutional Networks for Semantic Segmentation.”In Proceedings of theIEEEConference on Computer Vision and Pattern Recognition,pages 3431-40ComputerVision and Pattern Recognition(cs.CV),arXiv:1411.4038[cs.CV]

[0211] Jouppi,Young,Patil et al″ln-Datacenter Performance Analysis ofaTensor Processing Unit″44th International Symposium on Computer Architecture(ISCA),Toronto,Canada,24-28 June 2017(submitted 16 April 2017),arXiv:1704.04760[cs.AR]

Claims

1. A method for identifying tumors in histological images or a collection of histological images, the method comprising: Receive histological images or collections of histological images from records stored in a data repository; Image patches are extracted from the histological image or collection of histological images, wherein the image patch is a region of the histological image or collection of histological images whose size is defined by the number of pixels in its width and height; A convolutional neural network is provided with a set of weights and multiple channels, each channel corresponding to one of multiple tissue categories to be identified, wherein at least one of the tissue categories represents non-tumor tissue and at least one of the tissue categories represents tumor tissue; Each image patch is input as an input image patch into the convolutional neural network; Perform multi-level convolutions to generate convolutional layers whose size decreases progressively until a final convolutional layer with a minimum size is included. Then perform multi-level transposed convolutions to invert the convolutions by generating deconvolutional layers whose size increases progressively, until the layers are restored to a size that matches the input image patch. Each pixel in the restored layer contains a probability of belonging to each of the tissue categories. as well as Based on the probability, tissue categories are assigned to each pixel in the recovery layer to obtain output image patches.

2. The method of claim 1, further comprising: A probability map that assembles the output image blocks into the histological image or set of histological images.

3. The method of claim 2, further comprising: The probability map is stored in the record in the data repository such that the probability map is linked to the histological image or collection of histological images.

4. The method of claim 2, further comprising: The histological image or set of histological images and the probability map are displayed on the monitor.

5. The method of claim 1, further comprising: The convolutional neural network is provided with at least one skip connection, each of which obtains an intermediate result from at least one convolutional layer having a larger size than the final convolutional layer, and subjectes these results to as many transposed convolutions as possible to obtain at least one additional recovery layer that matches the size of the input image patch, wherein the transposed convolutions subjected to may be zero, one, or more than one; and Prior to the step of assigning an organization category to each pixel, the recovery layer is further processed to combine the recovery layer with the at least one additional recovery layer in order to recalculate the probability to take the at least one skip connection into account.

6. The method of claim 1, wherein the method is performed to train, wherein the recording includes ground truth data that assigns each pixel in the histological image or set of histological images to one of the tissue categories, the method of claim 1 being performed iteratively, wherein each iteration involves adjusting the weight values ​​of the convolutional neural network based on a comparison of the ground truth data with the output image patch.

7. The method of claim 1, wherein the histological images or set of histological images are a set of histological images obtained from adjacent sections stained differently from a tissue region.

8. The method of claim 2, further comprising: The region corresponding to the tumor in the histological image is defined according to the probability map; Each tumor is scored according to a scoring algorithm, and a score is assigned to each tumor. and The score is stored in the record in the data repository.

9. A computer device for identifying tumors in a histological image or a collection of histological images, the device comprising: Input, operable to receive histological images or sets of histological images from records stored in a data repository; A preprocessing module configured to extract image patches from the histological image or set of histological images, the image patches being a region of the histological image or set of histological images whose size is defined by the number of pixels in width and height; as well as A convolutional neural network having a set of weights and multiple channels, each channel corresponding to one of multiple tissue categories to be identified, wherein at least one of the tissue categories represents non-tumor tissue and at least one of the tissue categories represents tumor tissue, the convolutional neural network being operable to: a) Receive each image patch as input; b) Perform multi-level convolution to generate convolutional layers with progressively decreasing size until a final convolutional layer with the smallest size is included, followed by multi-level transposed convolution to invert the convolution by generating deconvolutional layers with progressively increasing size, until the layers are restored to a size that matches the input image patch, wherein each pixel in the restored layer contains a probability of belonging to each of the tissue categories. as well as c) Assign tissue categories to each pixel in the recovery layer based on the probability to obtain an output image patch.

10. The device as claimed in claim 9, The convolutional neural network described therein has at least one skip connection, each of which obtains intermediate results from at least one convolutional layer having a larger size than the final convolutional layer, and subjectes these results to as many transposed convolutions as possible to obtain at least one additional recovery layer that matches the size of the input image patch, wherein the transposed convolutions subjected to may be zero, one, or more than one; and The convolutional neural network is configured to further process the recovery layer prior to the step of assigning tissue categories to each pixel, such that the recovery layer is combined with the at least one additional recovery layer to recalculate the probability to take into account the at least one skip connection.

Citation Information

Patent Citations

  • Automatic Detection Of Mitosis Using Handcrafted And Convolutional Neural Network Features

    US20150213302A1