Pathological diagnosis assistance method and assistance device using ai
By using machine learning and neural network technologies to assist in pathological diagnosis, the problem of difficulty in identifying uneven differentiation within tumors has been solved, improving the efficiency and accuracy of pathological diagnosis and supporting better endoscopic treatment decisions.
Patent Information
- Application Number
- CN202080096757.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-26
- Filing Date
- 2020-12-25
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2040-12-25
AI Technical Summary
Existing technologies struggle to efficiently and accurately identify and classify the unevenness of differentiation within tumors in pathological diagnosis, impacting endoscopic treatment decisions, especially in early-stage gastric cancer where it is difficult to determine whether additional surgical resection is necessary.
By employing machine learning and neural network technologies, and through the segmentation and integration of image data observed under a microscope, pathologists can be assisted in identifying changes and heterogeneity in the degree of differentiation within tumors, providing a comprehensive assessment of the overall lesion.
It has achieved higher efficiency and accuracy in pathological diagnosis, reduced judgment time, decreased reliance on the experience of pathologists, and improved the accuracy of endoscopic treatment decisions.
Smart Images

Figure HDA0003800045810000011 
Figure HDA0003800045810000012 
Figure HDA0003800045810000021
Abstract
Description
Technical Field
[0001] This invention relates to a pathological diagnosis assistance method and apparatus using AI. Additionally, it relates to a method for classifying tumors from pathological sections in order to determine treatment strategies. Background Technology
[0002] Pathological diagnosis, which involves examining tissues or cells collected from patients using an optical microscope, plays a significant role as the final diagnosis. This is particularly true in oncology, where it becomes crucial for determining treatment strategies, evaluating treatment effectiveness, and predicting prognosis.
[0003] For example, in the case of gastric cancer, if pathological diagnosis indicates an early stage, endoscopic treatments such as endoscopic mucosal resection (EMR) and endoscopic submucosal dissection (ESD) are recommended. Compared to surgery, endoscopic treatment of early gastric cancer offers advantages such as less burden on the patient and better postoperative recovery, resulting in better quality of life (QOL). Therefore, endoscopic treatment is recommended for lesions with a high probability of radical cure.
[0004] Endoscopic resection is a recommended treatment for early gastric cancer, but its application is subject to the following two conditions: (1) the risk of lymph node metastasis is extremely low, and (2) the tumor is in a size and location that can be resected together.
[0005] In the guidelines (Non-Patent Literature 1, 2), for lesions with a presumed risk of lymph node metastasis of less than 1%, endoscopic removal and surgical resection are considered to yield equivalent results, and endoscopic treatment is considered more beneficial in such cases. Furthermore, evidence suggests that the size of intramucosal carcinomas with a lymph node metastasis risk of less than 1% varies depending on the histological type of gastric cancer, especially its degree of differentiation; therefore, pathological diagnosis for determining histological type is important.
[0006] Regarding the histological types of gastric cancer, in Japan, it can be broadly classified into two types according to the Japanese Gastric Cancer Society classification: differentiated carcinoma and undifferentiated carcinoma. This roughly corresponds to the Lauren classification used internationally: Intestinal (intestinal type) and Diffuse (spreading type). Besides the Japanese Gastric Cancer Society classification and the Lauren classification, there are also WHO classifications for gastric cancer histological types, but unless otherwise specified, the Japanese Gastric Cancer Society classification will be used here. In differentiated carcinoma, there is no risk of lymph node metastasis up to 3 cm in size, but in undifferentiated carcinoma, there is a risk of lymph node metastasis if it exceeds 2 cm (Non-Patent Literature 1, 2). Based on this evidence, guidelines for endoscopic treatment of early gastric cancer have been developed. However, the heterogeneity of differentiated and undifferentiated carcinomas often exists within the same lesion, significantly impacting the choice of treatment strategy.
[0007] Therefore, a preoperative diagnosis is performed before endoscopic resection. This diagnosis involves collecting biopsy tissue to obtain information on the histological type and differentiation degree of the cancer, and determining its size based on endoscopic observations. The preoperative diagnosis determines whether endoscopic treatment is suitable and, if appropriate, the extent of resection. Surgery is performed when treatment is deemed appropriate. However, it is difficult to accurately measure the size and depth of invasion from preoperative endoscopic findings; therefore, pathological histological examination of the endoscopically resected specimen is used to determine the radicality of the treatment.
[0008] If the histological observation of the endoscopically resected specimen determines that it is not a radical resection, then it becomes a candidate for additional surgical resection. Specifically, in cases where the tissue type of the resected specimen includes undifferentiated carcinoma larger than 2 cm mixed within a predominantly differentiated carcinoma, it is treated as a non-curative resection and is suitable for additional surgical resection, including lymph node dissection. Furthermore, even in cases where a predominantly differentiated lesion contains undifferentiated lesions infiltrating the submucosal layer (SM), it is also treated as a non-curative resection and is suitable for additional surgical resection.
[0009] To make such a radical diagnosis, the pathologist assesses the distribution and depth of invasion of differentiated and undifferentiated components within the lesion, integrates the overall lesion image, and makes a comprehensive judgment. Specifically, the following process is used for comprehensive judgment and diagnosis. Figure 17 ).
[0010] (1) Preparation of tissue specimens
[0011] After the endoscopic resection specimen is fixed in formalin, the pathologist cuts out the tissue specimen. The resected tissue sections are cut parallel to each other at 2-3 mm intervals, and macroscopic photographs of the fixed specimen with the cut lines are taken.
[0012] After specimen processing (tissue dehydration, clarification, and paraffin impregnation), paraffin blocks are prepared. These blocks are then thinly sliced, stained, and sealed to create slides for examination. The prepared slides contain parallel arrangement of the cut tissue sections (see reference). Figure 17 (Supported for inspection).
[0013] (2) Microscopic observation and diagnosis of tissue specimens
[0014] Perform microscopic observation and record the tissue type. In cases where multiple tissue types coexist, record them all in order of dominance, for example, tub1 (well-differentiated tubular adenocarcinoma) > por (poorly differentiated adenocarcinoma). Examine all sections of the pathological tissue specimen under a microscope to determine the degree of cancer differentiation (differentiated / undifferentiated).
[0015] Regarding the results of microscopic observation, in the macroscopic photograph with added secants, the distribution range of differentiated types was divided into red lines, and the undifferentiated types were divided into blue lines, etc., and recorded accordingly (see reference). Figure 17 (integration).
[0016] By recording the degree of differentiation for each slice, an overall lesion map is created. Using this integrated map, a comprehensive assessment is made of the distribution of undifferentiated components, whether their size exceeds 2 cm, and whether undifferentiated components are present in the submucosal layer (SM) infiltration area.
[0017] That is, after preparing the specimen slide from the paraffin section, the tissue type is determined and classified under a microscope. The results are then integrated into a lesion image, and a comprehensive judgment is made based on the overall specimen. The comprehensive judgment is made through the above three stages.
[0018] Diagnosing based on microscopic examination of specimens requires considerable time and experience. For example, in the case of early-stage gastric cancer resected endoscopically, the process from classification to diagnosis typically takes about 30 minutes, even for a skilled pathologist. Furthermore, it is said that becoming a skilled pathologist requires approximately 12 to 13 years of experience. Therefore, there is a need for methods and systems that can either assist skilled pathologists in making diagnoses more efficient and shortening the time required for judgment, or assist less experienced pathologists in making accurate diagnoses.
[0019] In recent years, attempts have begun to explore the use of artificial intelligence (AI), which has dramatically improved the accuracy of image recognition based on machine learning, to assist in image diagnosis. In endoscopic images, fundus examinations, and other fields, the practical application of auxiliary devices for detecting lesions is progressing.
[0020] In the field of diagnostic assistance for pathological tissues, devices using computer systems for assisted diagnosis have also been disclosed (Patent Documents 1-3). These devices or methods primarily detect characteristic areas of lesions from local tissue images and assist in pathological diagnosis.
[0021] Existing technical documents
[0022] Patent documents
[0023] Patent Document 1: Japanese Patent Application Publication No. 10-197522
[0024] Patent Document 2: Japanese Patent Application Publication No. 2012-8027
[0025] Patent Document 3: Japanese Patent Application Publication No. 2012-73179
[0026] Non-patent literature
[0027] Non-Patent Literature 1: Guidelines for the Treatment of Gastric Cancer (Physician's Edition, Revised 5th Edition, January 2018), edited by the Japanese Gastric Cancer Society, published by Kanehara Publishing Co., Ltd.
[0028] Non-patent literature 2: Hiroyuki Ono et al., ESD / EMR Guidelines for Gastric Cancer (2nd Edition), February 2020, Journal of the Japanese Society of Gastrointestinal Endoscopy, Vol. 62, pp. 275-290.
[0029] Non-patent literature 3: Schlegl, T. et al., IPMI 2017: Information Processing in Medical Imaging pp. 146-157
[0030] Non-patent literature 4: The Cancer Genome Atlas Research Network, 2014, Nature, Vol. 513, pp. 203-209
[0031] Non-patent literature 5: Kather, JN, et al., 2019, Nature Medicine, Vol. 25, pp. 1054-1056.
[0032] Non-patent literature 6: Gatys, LA et al., Journal of Vision, September 2016, Vol. 16, 326. doi: https: / / doi.org / 10.1167 / 16.12.326 Summary of the Invention
[0033] The problem that the invention aims to solve
[0034] However, all existing literature focuses on detecting the presence of lesions by searching for characteristic locations. This involves determining whether a lesion is "cancer (malignant tumor)" or "non-cancer (non-malignant tumor)." In other words, it maps the AI-based judgment of whether a lesion is cancerous or not. In this invention, unlike previous methods, the heterogeneity within cancer is visualized by mapping the judgment results obtained from further classification of cancer based on clinical evidence using AI. Specifically, a system is constructed that not only determines whether a lesion is cancerous or not, but also continuously performs detailed classification (classification) within lesions such as tubular adenocarcinoma and poorly differentiated adenocarcinoma, and then performs a comprehensive judgment after integrating the entire lesion. As mentioned above, in determining the likelihood of radical cure in endoscopic treatment and whether additional surgical resection is necessary, it is necessary not only to determine whether the lesion is cancerous or not at the site of the lesion, but also to identify the fine classifications within the lesion and their distribution and proportions while considering the continuous and heterogeneous differentiation process within the lesion, and to perform a comprehensive evaluation.
[0035] The objective of this invention is to provide a system and method for assisting diagnosis, which, while recording the location within a lesion, performs precise analysis of the lesion site, maps it to the lesion area, and integrates the entire lesion, thereby visualizing changes in the degree of differentiation and the heterogeneity of the tumor. According to the system and method of this invention, a system and method can be provided that visualizes the heterogeneity within a lesion and efficiently performs diagnosis, independent of the pathologist's skill level.
[0036] Methods for solving problems
[0037] This invention relates to the following auxiliary methods, systems, procedures, and learned models for pathological diagnosis.
[0038] (1) A pathological diagnosis auxiliary method, comprising: an image acquisition step of continuously acquiring microscopic observation image data of tissue specimens; a step of dividing the image data into a predetermined size to obtain image blocks in a manner that preserves positional information; a judgment step of determining the tissue type for each image block based on machine learning learning data; and an integration step of displaying the determined tissue type at each location.
[0039] (2) The pathological diagnosis auxiliary method according to (1) is characterized in that, in the above integration step of displaying the tissue type determined by machine learning at each location, the changes in the degree of differentiation and the heterogeneity of the tumor are displayed.
[0040] (3) The pathological diagnosis auxiliary method according to (1) or (2) is characterized in that the above-mentioned machine learning is performed using a neural network.
[0041] (4) The pathological diagnosis auxiliary method according to (3) is characterized in that the above neural network uses images with a resolution of 0.1 μm / pixel to 4 μm / pixel as learning images.
[0042] (5) The pathological diagnosis auxiliary method according to any one of (1) to (4), characterized in that the tissue specimen is derived from cancer tissue and the pathological diagnosis auxiliary method is used to assist in the pathological diagnosis of cancer.
[0043] (6) The pathological diagnostic auxiliary method according to any one of (1) to (5) is characterized in that the above-mentioned tissue specimen is a HE-stained (hematoxylin-eosin staining) specimen or an immunohistochemical staining specimen.
[0044] (7) The pathological diagnosis auxiliary method according to (6) is characterized in that, in the case that the above-mentioned HE staining specimen is an unstandardized specimen, the neural style transfer method is used to generate a simulated immunohistochemical staining specimen, and the tissue type is determined by using a learning model that has learned the immunohistochemical specimen image.
[0045] (8) A pathological diagnosis assistance system, comprising: an image acquisition unit for continuously acquiring microscopic observation image data of tissue specimens; an image processing unit for segmenting the image data into image blocks of a specified size in a manner that preserves positional information; a classification unit for determining the tissue type of each segmented image block based on machine learning learning data; and an integration unit for displaying the classified tissue type at each location.
[0046] (9) The pathological diagnostic auxiliary system according to (8) is characterized in that the image processing unit includes a unit for standardizing tissue staining specimens.
[0047] (10) The pathological diagnosis assistance system according to (8) or (9) is characterized in that the above-mentioned machine learning is a neural network.
[0048] (11) The pathological diagnosis assistance system according to (10) is characterized in that the neural network uses parameter values of tissue type obtained in advance in transfer learning.
[0049] (12) The pathological diagnosis assistance system according to any one of (8) to (11) is characterized in that the disease for which pathological diagnosis assistance is performed is cancer.
[0050] (13) A pathological diagnosis assistance program that enables a computer to perform the following processing: segmenting acquired image data into a specified size in a manner that preserves positional information, using a learning model learned from training data, determining the tissue type for each segmented image through machine learning, and displaying and integrating the determined tissue types at each location.
[0051] (14) The pathological diagnostic aid procedure according to (13) includes the processing of microscopic images of continuously acquired tissue specimens.
[0052] (15) A learned model for determining the tissue type of microscopic image data of a tissue specimen, the learned model enabling a computer to perform the following functions: the learning model includes an input layer that inputs microscopic image data of the tissue specimen to be the object after segmentation into a specified size; and an output layer that displays the tissue type determination results of the segmented image data at various locations of the tissue specimen. The learning model includes a classifier that classifies the tissue types based on data learned from training data, the training data being images after a physician has determined the tissue type, determining the tissue type of the input microscopic image data segmented into a specified size, and displaying it at various locations of the tissue specimen.
[0053] (16) According to the learned model described in (15), the training data is data that has been preprocessed for the images.
[0054] (17) A method for diagnosing a patient using any one of (1) to (7) of the pathological diagnostic aid method, based on an integrated graph showing the analysis results.
[0055] (18) A treatment method for a patient, characterized in that the treatment strategy is determined based on the obtained integrated map using any one of (8) to (12) of the pathological diagnostic auxiliary system. Attached Figure Description
[0056] Figure 1 This is a diagram illustrating one implementation of a pathological diagnostic support system.
[0057] Figure 2 This is a diagram illustrating another implementation of a pathological diagnostic aid system.
[0058] Figure 3 This is a flowchart illustrating the implementation of a method to aid in pathological diagnosis.
[0059] Figure 4 This is a diagram illustrating an example of how image patches are generated.
[0060] Figure 5 This is a diagram illustrating an example of an anomaly detection system based on a deep convolutional generative adversarial network (DCGAN) of normal tissue.
[0061] Figure 6 This is a diagram representing an example of an image dataset prepared for learning purposes.
[0062] Figure 7 This is a graph representing the accuracy of Scratch learning, fine-tuning-based training, and validation using differentiated / undifferentiated image datasets.
[0063] Figure 8 This is a graph representing the accuracy of Scratch learning, fine-tuning-based training, and validation using differentiated and undifferentiated image datasets with / with mixed elements.
[0064] Figure 9 It is a schematic representation of the two-dimensional distribution of morphological features from image patches in diagnostic images, visualized in the form of a scatter plot. The points in the scatter plot represent the local presence of the image patch from which it originates within the overall lesion.
[0065] Figure 10 This is a diagram illustrating the process of using a pathological diagnostic aid device.
[0066] Figure 11 This is a graph showing the construction and results of a two-class classifier for gastric cancer using HE-stained specimens. (A) shows an example of learning using image data. (B) to (D) show examples of learning curves, with (B) representing a baseline study, and (C) and (D) representing adjusted results.
[0067] Figure 12This diagram illustrates the construction and results of a 3-class classifier for gastric cancer using HE-stained specimens. (A) shows an example of learning using image data. (B) shows results using Resenet 18, (C) shows results using Inceptionv3, (D) shows results using Vgg-11 with batch normalization (Vgg-11_BN), (E) shows results using Densenet121, (F) shows results using Squeezenet1 (Squeezenet version 1.0), and (G) shows results using AlexNet.
[0068] Figure 13 This is a graph showing the construction and results of a 3-class classifier for thyroid cancer using HE-stained specimens. (A) shows an example of learning using image data. (B) shows the results using Densenet121, (C) shows the results using Resenet18, (D) shows the results using Squeezenet1, (E) shows the results using Inceptionv3, and (F) shows the results using Vgg-11_BN.
[0069] Figure 14 This is a graph representing a classifier constructed using characteristic pathological tissue images related to gene changes and expression, along with its results.
[0070] Figure 15 This diagram illustrates a method for classifying tissue types after taking mucosal thickness into account. (A) shows an example of obtaining three image patches longitudinally based on mucosal thickness. (B) shows the process of generating a mapping map based on the classification results of the image patches.
[0071] Figure 16 This is a diagram representing the processing of images using neural style transfer.
[0072] Figure 17 This is a diagram illustrating the workflow of existing pathological diagnosis techniques. Detailed Implementation
[0073] First, the implementation of the pathological diagnosis auxiliary system of the present invention will be described. Figure 1 and Figure 2The apparatus required for assisting pathological diagnosis comprises an optical system, an image processing system, a learning system, a classification system, an integration system, and a summarizing and judgment system. Specifically, it includes a microscope, a camera connected to the microscope, an image processing device, a learning device that learns from the processed images and determines parameters, and a computer system that classifies and judges the images based on the parameters determined by the learning device and integrates the classified and judged images. The computer system comprising the image processing system, classification system, learning system, integration system, summarizing and judgment system is as follows: Figure 1 As illustrated, it is possible to build a system within a server. Here, an example of building the system on different servers is shown, but it is also possible to build a system controlled by different programs on a single server.
[0074] Additionally, if using such Figure 2 As illustrated, a PC including a high-performance CPU and GPU can perform image processing, learning, classification, integration, and judgment functions sequentially within a single PC. Alternatively, although not shown here, it is also possible to configure different PCs to perform the image processing system, learning system, classification system, integration system, and summarization and judgment system separately. Furthermore, it is not limited to this; of course, systems tailored to the specific conditions of each facility can be constructed. The following describes the implementation method.
[0075] [Implementation Method 1]
[0076] This section describes systems that utilize servers. It summarizes and clarifies why computer systems other than PCs are built and hosted on servers. Figure 1 The optical system consists of a bright-field microscope and a camera device. The operator sets the field of view under the microscope and acquires images of pathological tissue through the camera device. Continuous images can be acquired by manually moving the field of view or by using a motorized stage. Alternatively, images can be acquired from WSI (Whole Slide Image) or so-called virtual slides of tissue specimens. Images acquired using the camera device are sent to an image processing system server for image processing.
[0077] In the image processing system server, routine image processing such as RGB channel separation, image histogram parsing, and image binarization are performed. After identifying and extracting epithelial tissue from the HE-stained image (described later), the image is segmented into appropriately sized image tiles (also called image tiles or image grids). Then, for training images, labels such as non-tumor (normal), differentiated, and undifferentiated are assigned before being sent to the training system server. For diagnostic images, no labels are assigned before they are sent to the classification system server.
[0078] As a learning system server, it can utilize not only its own local servers but also cloud services and data centers. Within the learning system, parameters used in the classification system are determined and sent to the classification system server.
[0079] like Figure 1 As shown, in implementations using cloud services or data centers, high-performance GPU resources can be utilized for learning on large-scale image datasets. For example, by uploading images from multiple facilities and annotating which regions are identified as disease areas, pathologists from multiple facilities can participate in the generation of classification criteria, achieving uniformity in diagnostic criteria. Furthermore, to create a high-quality judgment system, the accuracy of classification criteria can be managed by limiting the participation of pathologists in uploading images and annotations based on their experience and expertise.
[0080] The classification system server has the following functions: it classifies diagnostic images received from the image processing system using parameters determined by the learning system, and sends the result data to the integration system.
[0081] Diagnostic images are obtained in multiple, typically 3 to 5, images along the longitudinal direction based on the thickness of the mucosa. Along with this longitudinal order, the specimen slices are also assigned a number to maintain their position in the horizontal direction throughout the specimen. For each slice, the result of the image analysis, which category it is judged to be differentiated, undifferentiated, or non-tumor (normal) (output as 0, 1, 2), is saved and sent to the integrated system.
[0082] In the integrated system, for each vertical column, even if one of the three image blocks in that column is judged to be undifferentiated, it is still classified as undifferentiated. This yields a one-dimensional vector for each slice, which is then converted to grayscale. An image is drawn that corresponds differentiated types to black and undifferentiated types to white, resulting in a straight line that shows differentiated and undifferentiated types using white and black. This line is configured according to the coordinates of the starting point of each slice and the starting point of the cancerous region, resulting in an integrated image of the analysis results.
[0083] [Implementation Method 2]
[0084] Describe the system that performs image processing, classification, integration, judgment, and learning operations within a PC. Figure 2 The optical system is the same as in Implementation Method 1, but the systems that perform the image processing, classification, integration, judgment, and learning operations are all built within a single PC. Although the operations are performed within a single PC, the basic processing is the same as in Implementation Method 1.
[0085] In deep learning using neural networks, the technique of using GPUs as a processing unit when processing repeatedly large amounts of data; GPGPU (General-purpose computing on graphics processing units) can significantly reduce processing time and dramatically improve performance. This is particularly relevant when using servers located on-premises facilities or employing... Figure 2 In any case of the PC system shown, by importing a high-performance GPU, it is possible to learn from a large amount of image data within its own facilities.
[0086] (Example)
[0087] We initially used stomach cancer as an example, but it goes without saying that for any disease requiring pathological diagnosis, learning appropriate images using a suitable classifier is applicable regardless of the disease. Furthermore, as will be discussed later, high accuracy can be achieved by selecting an appropriate model based on the organ. Here, we used neural networks with high image recognition capabilities, particularly convolutional neural networks, but newer image recognition models such as EfficinentNet, BigTransfer (BiT), and ResNeSt can also be used. Additionally, models such as Transformer, Self-Attention, or BERT (Bidirectional Encoder Representations from Transformers), which are used in natural language processing, can be applied to image systems using Vision Transformer (ViT). Furthermore, of course, future models with high image recognition capabilities can be used.
[0088] [Auxiliary methods for pathological diagnosis]
[0089] Taking a gastric cancer tissue specimen that has undergone endoscopic resection as an example, the auxiliary methods for pathological diagnosis will be explained. Figure 3Similar to existing methods, tissue images are continuously acquired via an optical system. Location information, obtained as coordinates on the specimen, is sent to the integrated system server. The acquired images undergo image processing in the image processing system, followed by segmentation to generate image patches of appropriate size. The method of using these image patches for learning significantly reduces the time and effort required for annotation, thus greatly improving the efficiency of the development cycle, including model selection and adjustment.
[0090] Here, an image patch refers to a collection of image data encompassing all forms of images present within a tumor. Figure 4 The example shown illustrates the following: From the original image, the top and bottom blank areas are removed to create a portion containing only the tissue. This portion is then divided into three parts from left to right, creating three smaller images (1-3). Each of these smaller images is then divided into two parts horizontally and vertically, creating four smaller images, resulting in 3 × 4 = 12 image patches. For image 1 in the three-part division, further division by two parts horizontally and vertically yields four smaller patches. If image 1 is 512 × 512 pixels, then four 256 × 256 pixel image patches are obtained. To achieve high accuracy, it is crucial to homogenize the tissue components within each training data point. Because using image patches ensures uniform tissue distribution within each patch, high-quality training data is obtained, which is highly effective for developing high-performance models.
[0091] The location information from the original image is sent to the integration system. In the classification system, each image patch is analyzed and judged based on parameters generated from the learned images. The judgment results are sent to the integration system and mapped onto the original image. In the judgment system, based on the results of the integration system, it is determined whether additional surgical resection is necessary and displayed on the monitor. The physician corrects the AI-based integration map based on their own examination results, confirms the judgment results, and makes a diagnosis. Figure 3 ).
[0092] [Organization Type Classification - Construction of a Judgment System]
[0093] 1. Development Environment
[0094] The development environment uses Python as the programming language, and Keras, with PyTorch and TensorFlow as the backend, is used for development.
[0095] 2. Learn to use data
[0096] To build a diagnostic assistance system using AI, new tissue specimen image data needs to be pre-learned to accurately classify tissue types. Therefore, training data, prepared by experienced pathologists who have identified various tissue types, is required. This training data consists of training images and labels identified by the pathologists as non-tumor (normal), tumor, and tumor tissue types, and is used for supervised learning.
[0097] Images for learning purposes are generated using pathological tissue images of differentiated and undifferentiated gastric cancer. Within gastric cancer tissue, non-cancerous cells, mostly inflammatory cells such as lymphocytes and macrophages, and non-epithelial cells such as fibroblasts, coexist with cancer cells. The differentiated / undifferentiated classification system generated here is based on the morphology and structure of the cancer, thus not requiring non-epithelial cells. Therefore, tissue specimens are first immunohistochemically stained (IHC) using cytokeratin, an intracellular protein found in epithelial cells, instead of the hematoxylin-eosin (HE) staining specimens used for routine diagnosis.
[0098] Cytokeratins are classified into 20 subtypes and are present in all epithelial tissues, exhibiting stable expression even in tumorigenesis. For example, they are expressed in cancers of the esophagus, stomach, colon, liver, bile ducts, breast, and prostate. By binding and staining tissue sections with antibodies targeting cytokeratins expressed in epithelial cells, tissue specimens with only epithelial cells stained are obtained. By acquiring images from the stained tissue specimens, morphological classification of epithelial tissues can be performed.
[0099] Immunohistochemical staining based on cytokeratin can stain only epithelial tissues, making it effective as image data for morphological classification using only the structure of epithelial tissues. Therefore, for cancers originating from epithelial tissues, cytokeratin-based immunohistochemical staining can be used as an auxiliary method for pathological diagnosis, applicable to most cancers. Particularly effective are adenocarcinoma, gastric cancer, colon cancer, pancreatic cancer, bile duct cancer, breast cancer, and lung cancer, among others, which are epithelial-derived cancers suitable for cytokeratin-based immunohistochemical staining. In other words, the method of classification using cytokeratin-based immunohistochemical staining can be applied to a variety of cancers.
[0100] On the other hand, pathological diagnosis in actual clinical practice mainly uses HE-stained specimens. When pathologists observe tissue specimens using an optical microscope, they unconsciously identify and extract epithelial and non-epithelial tissues from the tissue images of HE-stained specimens and perform visual identification, making morphological judgments only on epithelial tissues to diagnose cancer. Regarding this unconscious step performed by pathologists, if AI is used to perform image processing-based conversion of HE-stained specimens, or image conversion and image generation using deep learning, it is possible to convert and generate images equivalent to those based on cytokeratin-based immunohistochemical staining. By performing image conversion or image generation based on HE specimens identical to those used in actual clinical practice, AI can learn or analyze the images. By using HE-stained specimens, it can be applied regardless of the type of tumor. Furthermore, as mentioned later, when HE-stained images cannot be directly used due to faded specimens, lightly stained specimens, or differences in specimen color between facilities, if neural style transfer methods are used to process HE-stained images into simulated cytokeratin staining images, it is possible to make judgments through a pathological diagnosis assistance system.
[0101] By combining conventional image processing methods, such as RGB channel separation, histogram analysis, and binarization, images suitable for morphological classification of epithelial tissues can be obtained. Furthermore, image generation using deep learning methods, such as GANs (Generative Adversarial Networks), can generate training images suitable for morphological classification. In the case of deep learning, by learning from HE-stained specimens and consecutive immunohistochemically stained specimens of cytokeratin in pairs, corresponding transformed images can be obtained pixel by pixel. Such image preprocessing is crucial for learning efficiency and acceleration, and it can correct and standardize image color differences caused by variations in staining states between facilities and specimen thickness, thus also affecting the standardization of training and diagnostic images.
[0102] By using image generation leveraging GANs, it is also possible to identify lesions with relatively low frequency. However, it is difficult to prepare a large number of training images for low-frequency lesions, making it challenging to use existing classification models for identification. But, in methods combining GANs and machine learning-based anomaly detection (such as AnoGAN), it is possible to use a GAN model pre-learned from normal tissue to quantify the anomaly degree of the image for identification. Figure 5In this study, a GAN model (DCGAN; Deep Convolutional GAN) obtained by learning only images of normal fundic gland tissue through a convolutional neural network model was used to perform an analysis using an AnoGAN model that detects anomalies (Non-Patent Literature 3). Regarding the numerical results of the anomaly score, the anomaly score for the learned normal fundic gland tissue (4, 5) was over 1200, but the anomaly score for the images of the unlearned fundic epithelium (1, 2) and poorly differentiated adenocarcinoma (3) was much higher. If an anomaly detection method using images of normal tissue is employed for machine learning, even lesions for which it is difficult to prepare a large number of training images can be detected to distinguish them from normal tissue.
[0103] In addition, regardless of the staining method used, in addition to preparing the pathological specimens evenly, it is important to use quality-controlled equipment, reagents and procedures so that the hue and intensity of the staining can be determined.
[0104] The performance of a classification system is determined by both the performance of the neural network model and the learning images, making the learning images crucial in its construction. Transfer learning leverages excellent models that have undergone repeated refinement in competition using large-scale image data, but users must prepare the learning images tailored to their specific problems. Users need to refine the learning images to suit the tasks to be addressed in various fields. In the following example of an endoscopic resection specimen of gastric cancer, image patches are generated based on images obtained from the normal and cancerous portions of the mucosa, in a way that allows for approximate identification as normal mucosal portion, differentiated carcinoma, and undifferentiated carcinoma.
[0105] Depending on the combination of the objective lens and the imaging lens, the image resolution is in the range of approximately 0.1 μm / pixel to 4 μm / pixel, preferably 0.2 μm / pixel to 4 μm / pixel, and more preferably 0.25 μm / pixel to 2 μm / pixel. This is preferable considering the resolution of commonly used microscopes. The image size is in the range of approximately 34 × 34 pixels to 1024 × 1024 pixels. When it is necessary to identify morphologies limited to a very small area, a high-resolution, small image can be used for learning. Furthermore, in the embodiment based on cytokeratin immunohistochemical specimens, learning was performed based on an image of 256 × 256 pixels, but depending on the model used for learning, a larger image may sometimes be required. The image size can be appropriately selected based on the objective, the state of the tissue, and the learning model.
[0106] Here, image patches used for learning and diagnosis are reduced to a size that can be visually identified. By reducing the size of the image patches, the morphology of the tissues contained within a single image patch can be made closer to uniform. Furthermore, if small image patches containing only a single morphology are used, it is also possible to analyze tumors with boundary morphologies that are difficult to determine using the existing binary classification based on differentiation / undifferentiation from a new perspective. In the image patches generated for diagnostic images, data on the morphology of all pathological tissues contained within each tumor, along with the location data within that tumor tissue, can be obtained together.
[0107] 3. Deep Learning
[0108] Tissue images of cancerous areas were obtained from specimens stained with immunohistochemical staining using anti-cytokeratin antibodies. These images underwent rotation, segmentation, and magnification processing to generate 644 images of 256×256 pixels each, serving as the training dataset. The actual data volume also includes color information, resulting in a total data size of 256×256×3. Dataset preparation included dataset a, which clearly distinguishes between differentiated and undifferentiated cancer; and dataset b, which contains undifferentiated components mixed with differentiated components, and datasets with / without undifferentiated components. Figure 6 ).
[0109] The deep learning used here can be any network suitable for image processing; there are no particular limitations. Examples of high-performance image recognition neural networks include Convolutional Neural Networks (CNNs), or models used in natural language processing such as Transformer, Self-Attention, and BERT (Bidirectional Encoder Representations from Transformers) applied to image systems, such as Vision Transformer (ViT). Examples of CNNs include those developed for the Image Recognition Competition (ILSVRC) using the ImageNet large-scale object recognition dataset, such as ResNet, Inception-v3, VGG, AlexNet, U-Net, SegNet, DenseNet, and SqueezeNet. Newer image recognition models such as EfficinentNet, Big Transfer (BiT), and ResNeSt can also be used. Furthermore, future models with high image recognition performance can certainly be used.
[0110] [Example 1]
[0111] First, using immunohistochemical staining specimens of cytokeratin, through dataset a( Figure 6The datasets shown are easily identifiable as differentiated or undifferentiated, and transfer learning is performed in ResNet. Transfer learning utilizes a learned network and parameter values to learn a new task. It employs two methods: initial parameter (weight) values learned by ResNet on a large dataset, transfer learning using its own image data, and Scratch learning, which uses only the model's structure to learn from its own image data without using the initial values of the learned parameters.
[0112] The following device system is used in this invention.
[0113] A; CPU; Intel Core i7 without GPU
[0114] B; CPU; Intel corei7, GPU NVIDIA GeForce 1080-Ti
[0115] C; CPU; Intel Xeon GPU NVIDIA GeForce RTX 2080-Ti
[0116] D; CPU; Intel corei9, GPU NVIDIA Titan-V
[0117] In learning using dataset a generated from cytokeratin immunohistochemical staining specimens of differentiated / undifferentiated gastric cancer, the recognition accuracy was 0.9609 through Scratch learning and 0.9883 through transfer learning (fine-tuning), achieving sufficient recognition accuracy. Figure 7 ).
[0118] Next, in dataset b, cytokeratin immunohistochemical staining specimens of gastric cancer with a mixture of undifferentiated and undifferentiated components were used. Figure 6 In the learning process, 0.9219 was obtained through Scratch learning and 0.9102 was obtained through transfer learning (fine-tuning). Even with these datasets, sufficient recognition accuracy for diagnostic assistance can be achieved. Figure 8 Specimens with a mixture of undifferentiated and undifferentiated components are also challenging tissue images for pathologists. The AI's accuracy rate decreases for datasets that are difficult for pathologists to interpret, but even in these cases, it achieves an accuracy rate of over 90%, thus demonstrating sufficient precision.
[0119] Overlearning occurs when learning parameters specialized from the training data leads to decreased generalization performance, resulting in reduced accuracy on validation data. To prevent overlearning, increasing the training data is the optimal solution. However, data augmentation is a method to efficiently increase training data with limited data. Data augmentation involves processing images of the training data, such as rotating, deforming, cropping the center, enlarging, or reducing them, to increase the amount of training data. Here, we will also discuss data augmentation of training images based on these processing methods.
[0120] By applying deep neural networks, the dimensionality of the data is successively reduced, and the essential features (also known as latent representations or internal representations) of the input image are extracted. One method for obtaining features (latest representations or internal representations) from data is the autoencoder. This involves an encoder extracting features (latest representations or internal representations) from the original data, and a decoder then reconstructs the original data based on these features. In the case of image data, the encoder portion can utilize convolutional neural networks, but this can also be viewed as a method of extracting latent representations using neural networks. By using a classifier generated in transfer learning as a feature extractor, features are extracted from image patches generated for diagnostic purposes, thus reducing the dimensionality to two dimensions. This allows visualization of the distribution of features derived from tissue morphology as a scatter plot. Further dimensionality reduction to two dimensions is achieved by generating image patches from tissue specimens of differentiated and undifferentiated gastric cancer, allowing visualization of the two-dimensional distribution of morphological features as a scatter plot.
[0121] Figure 9 The left side shows a two-dimensional scatter plot generated by reducing the dimension of morphological features present in the tumor. In image patches generated from diagnostic images, a coverage image patch of the pathological tissue morphology included in each tissue is obtained along with location data within the tissue. Therefore, scatter plots can be used to cluster tumors with boundary morphologies that are difficult to determine using existing differentiated / undifferentiated binary classification methods. Furthermore, by tracing where the image patches from which each point of the scatter plot originates exist within the lesion, it is also possible to know how the feature clusters are distributed within the actual tumor. Figure 9 right).
[0122] By using features and latent manifestations extracted from pathological tissue images, it is possible to concatenate and fuse them with features obtained from other modalities of data beyond images. This allows for multimodal learning, combining features from different modalities to infer outcomes. Instead of existing classifications like differentiated or undifferentiated types, it combines the features inherent in pathological tissue images with features extracted from other modalities, enabling the complementary use of information from different modalities in inferring outcomes. For example, in clinical policy decisions such as determining treatment strategies, a system can be constructed that fuses different types of data, such as categorical data like gender, numerical data like examination data, image data like radiographic images, and information on genetic variations, to form a judgment.
[0123] [Example 2]
[0124] [Generation of the integrated mapping graph]
[0125] Using the aforementioned diagnostic system, analysis and integration were performed on gastric cancer tissue specimens containing a mixture of differentiated and undifferentiated components. Figure 10 In this example, a model was constructed using cytokeratin-based immunohistochemical staining specimens to make judgments.
[0126] Tissue sections are arranged parallel to each other on the slide for examination. In cases where the length of one tissue section is larger than the slide, such as... Figure 10 As shown, to accommodate the specimen on a glass slide, it was cut into two pieces (A and B) of different lengths to prepare the specimen. From a single slice composed of A and B on a cytokeratin-stained specimen, tissue images were continuously acquired while moving the field of view. Here, three consecutive images are shown. Simultaneously with image acquisition, the coordinates of the starting point of each slice of the tissue specimen and the coordinates of the starting point of the cancerous region were obtained. The aforementioned consecutive images were segmented to generate 256×256 pixel images. Furthermore, in the classifier using cytokeratin-based immunohistochemical staining specimens, high accuracy was achieved even when the learning images contained slightly different tissue components, enabling analysis based on weakly magnified images. Therefore, in this example, the longitudinal direction (the thickness direction of the mucosa) could be covered by a single column of image blocks.
[0127] The segmented images are assigned numbers horizontally, maintaining the original order of the specimen slices. The generated classification and judgment system is used to analyze the image groups. The analysis result is determined as differentiated or undifferentiated, outputting the result as 0 or 1 and saving it as a one-dimensional vector. The one-dimensional vector obtained from one slice is then converted to two dimensions, changing the vector value from 1 to 255, thus corresponding to the two colors 0 (black) and 255 (white). The resulting two-dimensional vector is plotted as an image, resulting in lines representing differentiated or undifferentiated types using black or white. In this example, white corresponds to undifferentiated type, and black corresponds to differentiated type.
[0128] Here, the same one-dimensional vector is double-nested using NumPy to achieve two-dimensionality, and then the pyplot module of Matplotlib, a plotting package that functions under NumPy, is used to depict the two-dimensional vector as a straight line. Furthermore, NumPy is a library that provides multi-dimensional manipulation and numerical computation capabilities in the Python development environment. Here, the one-dimensional vector of the analysis results output by the judgment system is double-nested using NumPy to achieve two-dimensionality, and the pyplot module is used to depict it as a straight line.
[0129] Based on the coordinates of the starting point of each slice and the starting point of the cancerous tissue, straight lines are drawn from each slice, thereby obtaining an integrated map of the overall analysis results of the specimen tissue. This further develops the use of deep learning in existing technologies, which only determine the presence of lesions. As shown in this embodiment, it is integrated into a model that comprehensively judges cancerous lesions based on information about their differentiation degree and location. As a result, it is possible to depict a mapping of the undifferentiated degree of cancerous lesions, instantly determining the location and size of undifferentiated components and quickly assessing the risk of metastasis.
[0130] Furthermore, accurately depicting the differentiation degree determined by pathologists through microscopic observation on the secant lines of a specimen's overall map has limitations, especially when tissues with different differentiation degrees are mixed within a very narrow area, making it difficult to faithfully map and reproduce the distribution. Therefore, human-generated integrated maps naturally have limitations in terms of precision, and their generation also requires a significant amount of time. The method developed in this paper, however, enables the generation of precise mapping maps, which can significantly contribute to the selection of treatment methods.
[0131] When pathologists use this system to examine actual specimens, they refer to the AI-generated integrated diagram while simultaneously revising it based on their own examination results to complete the integrated diagram used in the diagnostic report. This not only significantly improves the precision of the integrated diagram but also drastically reduces the time required to generate the diagnostic report.
[0132] [Example 3]
[0133] This paper describes a pathological diagnostic aid method based on HE staining images. In actual clinical practice, pathological diagnosis primarily utilizes HE-stained specimens. To make the classifier constructed from immunohistochemically stained specimens of cytokeratins a more universal classifier, a classifier based on HE-stained specimens is needed. A training dataset from HE-stained specimens is generated, and a two-class classifier for differentiated / undifferentiated gastric cancer is first constructed. Figure 11 ).
[0134] To create a high-accuracy classifier, training with image datasets is essential. Preparing a large dataset is crucial. Figure 11 (A) shows a training image dataset with two levels: differentiated and undifferentiated. The morphological components contained in a single training image are designed to be as singular as possible, preparing the image to avoid containing multiple components. However, regarding the image size itself, depending on the convolutional neural network used in the classifier, specifications such as 224×224 pixels and 299×299 pixels are provided. Therefore, images are prepared such that they contain small regions consisting of only a single morphological component. When segmenting a large image to obtain an image of a specified size, the segmentation size is also studied in a way that the segmented image satisfies the above conditions.
[0135] Next, for a given model, adjustments are made by changing the learning conditions in a way that increases accuracy on the validation data. Figure 11 The image shows an example of a learning curve when ResNet18 is used in the model. In the baseline study, the accuracy varied greatly in the validation data (val) for each epoch. Figure 11 (B)). Therefore, by adjusting the learning rate and block size, the learning curve is improved from the baseline to 1 ( Figure 11 (C)), 2( Figure 11 (D)). In the learning curve of 2, there is little tendency to overlearn, and the highest accuracy value in the validation data is 0.9683.
[0136] [Example 4]
[0137] Using a training dataset comprising three levels of gastric cancer (differentiated and undifferentiated) with the addition of normal tissue, six models were compared to generate a high-accuracy classifier. Figure 12 Typical tissue images of differentiated, undifferentiated, and normal tissues are shown below. Figure 12(A). Using these learning image datasets, six models were used to generate classifiers. Many models are available for image classification based on transfer learning, all of which are evaluated through competition using the ImageNet dataset for image recognition or their improvements. In ResNet18 (… Figure 12 (B)) and InceptionV3 Figure 12 In (C)), the highest accuracy values in the validation data reached above 0.93 and 0.92, in AlexNet ( Figure 12 An accuracy of 0.9639 was obtained in (G)). Furthermore, in Densenet121 ( Figure 12 In (E), generalization performance finally begins to appear after 100 generations of learning; if learning continues beyond 100 generations, sufficient generalization performance may be achieved. On the other hand, in Vgg11_BN( Figure 12 (D)) Generation 60 and later or Squeezenet1 Figure 12 (F)) identifies the tendency to overlearn. In this way, we can observe the learning curve while obtaining parameters from the optimal model and generation that achieves generalization performance without falling into overlearning, and then apply them to actual clinical practice.
[0138] [Example 5]
[0139] Next, the pathological diagnosis assistance method shown in the embodiments is applicable not only to gastric cancer but also to various cancers. This specification describes a pathological diagnosis assistance method that involves preparing a training dataset, selecting a model suitable for the training dataset for transfer learning, saving parameters, and using it in a classification system. This series of methods is not limited to differentiated or undifferentiated gastric cancer but can also be used for cancers of other organs. In various cancers other than gastric cancer, it is known that a mixture of tissue types exists within the same tumor. Examples described in the processing specification include endometrial cancer, thyroid cancer, and breast cancer. For example, in thyroid cancer, when a mixture of well-differentiated components (papillary carcinoma, follicular carcinoma) and poorly differentiated components exists, if the poorly differentiated component accounts for more than 50%, it is described as poorly differentiated. Endometrial cancer is divided into three stages based on the proportion within the tumor. In such cases, by constructing separate training image datasets for well-differentiated and poorly differentiated cancers to generate classifiers, a well-differentiated / poorly differentiated mapping can be generated.
[0140] This represents an example of building a learning model for thyroid cancer. Figure 13 The thyroid gland is classified into three types: well-differentiated (papillary carcinoma), poorly differentiated, and normal thyroid tissue. Figure 13 (A)) Model construction. In this dataset from the thyroid gland, using ResNet18 ( Figure 13(C)) or InceptionV3 Figure 13 (E)), Squeezenet1 ( Figure 13 (D) and others tend to overlearn, but Densenet121 ( Figure 13 In (B), by obtaining the parameters before the learning process ends around generation 40, a model with an accuracy of approximately 90% can be obtained. Additionally, in Vgg11_BN( Figure 13 In (F), the accuracy in the validation data shows a slight tendency to increase, so it is believed that further increasing the number of generations of learning will improve the accuracy. As shown in the examples of gastric cancer and thyroid cancer, the optimal model is different for each organ. Therefore, it is necessary to select a model according to each organ, choosing a model with high accuracy.
[0141] [Example 6]
[0142] For tumors with characteristic pathological tissue images known to be associated with gene changes and expression, the following pathological diagnostic aid method can also be used: prepare a training dataset, select a model suitable for the training dataset, save the parameters, and use it in the classification system. A vast amount of information has accumulated regarding changes in cancer-related genes and proteins, but for those genes and proteins known to be associated with histological morphology, classification based on characteristic pathological tissue images can determine the strategy for further gene and cell biology searches. Furthermore, this also serves as an aid in recommending the level of further examination.
[0143] It is known that gastric cancer is roughly divided into the following four molecular subtypes in The Cancer Genome Atlas (TCGA) based on the NCI (National Cancer Institute) (Non-Patent Literature 4).
[0144] Chromosomal Instability (CIN)
[0145] Genomically stable (GS)
[0146] Microsatellite instability(MSI)
[0147] Epstein-Barr Virus (EBV)
[0148] In MSI or EBV type gastric cancers, a characteristic histological pattern of significant lymphocyte infiltration into the tumor is known. Additionally, a large proportion of the cancers contained in GS are histologically packaged (poorly differentiated adenocarcinomas), and some are known to involve alterations in several genes.
[0149] A method for estimating MSI from HE-stained images using deep learning has been reported (Non-Patent Literature 5). In gastric cancers accompanied by HER2 protein overexpression / gene amplification, treatment with molecularly targeted drugs has established concomitant diagnosis using immunohistochemistry or FISH. Therefore, in addition to MSI, if HER2 protein overexpression / gene amplification can be estimated via HE staining, HER2-positive gastric cancers will not be missed during HE-stained pathological diagnosis, allowing for more appropriate treatment. It is known that in such HER2-positive gastric cancers, the frequency of HER2 protein overexpression is significantly higher in the intestinal type (tubular adenocarcinoma) and petal type (poorly differentiated adenocarcinoma, etc.) based on the Lauren classification. In cases where intestinal (tubular adenocarcinoma) and petal type tissue images are mixed, concomitant diagnosis is recommended in specimens containing more intestinal type images.
[0150] exist Figure 14 In this study, tissue images (A) from cancer tissues with HER2 protein overexpression / gene amplification, (B) from cancer tissues confirmed by MSI, and (C) from cancer tissues with gene variants observed in packaging cancers were used to generate a training image dataset for 3-class classification. Images were trained to classify into 3 classes, and a classifier was generated and validated using diagnostic images. An accuracy of 0.9656 was achieved on AlexNet, while on ResNet18, requiring twice as many generations of training as AlexNet, an accuracy of 0.9089 was achieved.
[0151] When characteristic pathological tissue images associated with gene mutations or expression are known, changes in gene mutations or expression can be inferred from the pathological images. As shown in the example, if the possibility of HER2 protein overexpression / gene amplification can be indicated during pathological diagnosis using HE staining, a definitive diagnosis can be made by immunohistochemical staining, allowing for HER2-targeted therapy. Pathologists in high-volume facilities generally have a good grasp of the general characteristics of pathological tissue images associated with such gene mutations and expression. This accumulation of expert experience can be used as a classifier and installed in the pathological diagnostic assistance system. Furthermore, if diagnostic accuracy is improved, HE staining can also be used as a companion diagnostic tool for selecting molecularly targeted drugs. In the example, classification is performed as the three categories mentioned above, but if new characteristic tissue images associated with gene changes or variations are known, new categories can be added sequentially to generate a classifier.
[0152] [Example 7]
[0153] When using a classifier generated from an image dataset of HE-stained specimens to actually generate differentiated / undifferentiated maps from endoscopically resected specimens, 3 to 5 image blocks are obtained longitudinally based on the thickness of the mucosa. Figure 15 (A) represents an example where three image blocks were obtained longitudinally. This example is an image block of the gastric mucosa lamina propria in an HE specimen, consisting of three longitudinal blocks and ten transverse blocks.
[0154] exist Figure 15 (B) In the middle section, blocks classified as normal within the image patch are represented in gray, differentiated blocks in black, and undifferentiated blocks in white. The lower section shows a mapping where, even if one of the three image patches in a vertical column is classified as undifferentiated, it is still classified as undifferentiated. Here, if there is one undifferentiated patch in each vertical column (i.e., in the thickness direction), a base for classification as undifferentiated is set. However, if several patches in the thickness direction are classified as undifferentiated, the base for creating the horizontal mapping based on that column as undifferentiated can be freely changed. By performing this process, classification can be performed even in cases of raised lesions or when using small image patches.
[0155] [Example 8]
[0156] In actual clinical practice, pathological diagnosis primarily uses HE-stained specimens. However, as a countermeasure against specimens that have faded over time, lightly stained specimens, and differences in specimen color between facilities, neural style transfer using convolutional neural networks (non-patent literature 6) has been utilized. For content images, which are equivalent to the global composition of an image, features, colors, and textures independent of their position within the image are treated as style images and processed using methods such as neural style algorithms. As an example, an example of converting a faded specimen into an immunohistochemically stained specimen of cytokeratin using neural style transfer image processing is shown. Figure 16 Images that have been transformed to simulate cytokeratin (IHC) can be identified and mapped using a classifier constructed from cytokeratin-stained specimens.
[0157] This specification demonstrates that this diagnostic aid method can be applied using HE-stained specimens of gastric and thyroid cancer. However, by selecting an appropriate model, it is possible to determine the tissue type of cancer in all organs, visualize the heterogeneity within the tumor, and thus aid in pathological diagnosis. Furthermore, it illustrates the determination of tissue type in gastric cancer based on immunohistochemical staining images of cytokeratin, but is not limited to cytokeratin. By utilizing immunohistochemical staining images based on proteins that are indicators of cancer malignancy or proteins that show responsiveness to treatment (biomarkers), it is possible to generate tumor malignancy maps and treatment responsiveness maps to visualize the heterogeneity within the tumor. Additionally, as shown in the examples, it is possible to identify cancers suspected of HER2 overexpression based on HE images, indicating its applicability to digital composite diagnostics.
[0158] Morphological findings based on pathological diagnosis—namely, typical normal tissue, benign tumors, and malignant tumors (carcinoma, sarcoma)—are documented and shared in histological and pathological diagnostic literature. These findings are further inherited and guided during pre-pathology training and are explicitly stated as cancer management guidelines and WHO classifications, and shared in academic societies and case study meetings. However, the morphological differences in tissues are continuous, and there are many clinically challenging cases involving the differentiation of normal tissue from tumors, and the distinction between differentiated and undifferentiated types of tumors. Subtle judgments are often delegated to individual pathologists. It is known that the diagnosis of intramucosal carcinoma of the stomach differs between Japan and Europe / America. For histological findings that are inconsistent, this method allows for easy re-examination of the relationship between the system's results and prognosis based on past data, and prognosis can be studied by confirming past pathological data and recurrence rates such as lymph node metastasis. As a result, appropriate treatment can be provided to patients based on the relationship between treatment selection and prognosis, and treatment methods can be established for previously difficult-to-diagnose heterogeneous cancers within the tumor.
Claims
1. A method of aiding in the diagnosis of pathology, characterized by, comprises: an image acquisition step of continuously acquiring microscopic observation image data of a tissue specimen; a step of dividing the image data into image blocks of a prescribed size in a manner that maintains positional information in the entire specimen and positional information on the tissue image; a judgment step of judging the class of each image block based on a feature quantity extracted from learning data by machine learning; and an integration step of displaying the judged class at each position of the tissue image and integrating analysis result maps of the entire specimen.
2. The pathological diagnosis assistance method according to claim 1, wherein: in the integration step, the heterogeneity of a tumor is displayed.
3. The pathological diagnosis assistance method according to claim 1 or 2, wherein: the machine learning is performed using a neural network.
4. The pathological diagnosis assistance method according to claim 3, wherein: the neural network uses an image having a resolution of 0.1 pm / pixel to 4 pm / pixel as learning image.
5. The pathological diagnosis assistance method according to claim 1 or 2, wherein: the tissue specimen is derived from a cancer tissue, the pathological diagnosis assistance method is used to assist the pathological diagnosis of cancer.
6. The pathological diagnosis assistance method according to claim 1 or 2, wherein: the tissue specimen is an HE-stained specimen or an immunohistochemically-stained specimen.
7. The pathological diagnosis assistance method according to claim 6, wherein: in the case where the HE-stained specimen is a non-standardized specimen, a simulated immunohistochemically-stained specimen is generated using a neural style transfer method, a learning model that has learned an immunohistochemical specimen image is used to judge the tissue type.
8. A pathological diagnosis assistance system characterized by comprising: comprises: an image acquisition unit of continuously acquiring microscopic observation image data of a tissue specimen; an image processing unit that divides the image data into image blocks of a prescribed size in a manner that maintains positional information in the entire specimen and positional information on the tissue image; a classification unit that judges the class of each divided image block based on a feature quantity extracted from learning data by machine learning; and an integration unit that displays the judged class at each position of the tissue image and integrates analysis result maps of the entire specimen.
9. The pathological diagnosis assistance system according to claim 8, wherein: the image processing unit includes a unit for standardizing a tissue-stained specimen.
10. The pathological diagnosis assistance system according to claim 8 or 9, wherein: the machine learning is performed using a neural network.
11. The pathological diagnosis assistance system according to claim 10, wherein: the neural network uses a parameter value of a tissue type that is obtained in advance in transfer learning.
12. The pathological diagnosis assistance system according to claim 8 or 9, wherein: the disease for which pathological diagnosis assistance is performed is cancer.
13. A computer program product, comprising a pathological diagnosis assistance program, which, when executed by a computer, enables the following processing: dividing the acquired image data into a predetermined size in a manner that maintains positional information in the entire specimen and positional information on the tissue image, using a learning model learned using the learning data, judging a class for each of the divided images based on a feature quantity extracted from the learning data by machine learning, displaying the judged class at each position of the tissue image, and integrating an analysis result map of the entire specimen.
14. The computer program product according to claim 13, characterized by: including a process of continuously acquiring microscope observation images of a tissue specimen.
15. A pathological diagnosis assistance device, characterized by: judging a class of image data based on microscope observation image data of a tissue specimen using a learned model, the pathological diagnosis assistance device including: an input layer into which microscope observation image data of a tissue specimen to be an object is input after being divided into a predetermined size; and an output layer that displays a class judgment result of the divided image data at each position of a tissue image, the pathological diagnosis assistance device including a classifier that classifies the class based on a feature quantity extracted from learning data by machine learning, judging a class of microscope observation image data that is input in a manner that maintains positional information in the entire specimen and positional information on a tissue image and is divided into a predetermined size, displaying the judged class at each position of a tissue image, and integrating an analysis result map of the entire specimen.
16. The pathological diagnosis assistance device according to claim 15, characterized by: the learning data is data on which a pre-process is performed for an image.
Citation Information
Patent Citations
Pathological tissue diagnosis support device
JP1998197522A
Pathological diagnosis support device, pathological diagnosis support method, control program for supporting pathological diagnosis, and recording medium recorded with control program
JP2012008027A
Pathological diagnosis support device, pathological diagnosis support method, control program for pathological diagnosis support and recording medium with control program recorded thereon
JP2012073179A
Method and system for digital staining of label-free fluorescence images using deep learning
WO2019191697A1