Cancer cell identification method and cancer cell identification system
The cancer cell identification method and system use a CGAN and CNN to estimate fluorescently stained images from unstained images, enhancing classification accuracy and reducing the need for fluorescent staining, making CTC testing more accessible and cost-effective.
Patent Information
- Application Number
- PCT/JP2025/025670
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-23
- Filing Date
- 2025-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Existing methods for detecting circulating tumor cells (CTCs) in blood using unstained images suffer from reduced accuracy due to low brightness and lack of antibody expression information, and conventional AI-based deep learning models do not provide sufficient classification accuracy for cancer cells and blood cells.
A cancer cell identification method and system using a two-stage machine learning approach, comprising a Conditional Generative Adversarial Network (CGAN) to estimate fluorescently stained images from unstained images and a Convolutional Neural Network (CNN) for classification, to enhance the accuracy of cancer cell detection.
Improves the accuracy of cancer cell classification in unstained images by capturing antibody expression information, reducing the need for fluorescent staining and specialized microscopes, thereby lowering costs and increasing accessibility of CTC testing.
Smart Images

Figure JP2025025670_29012026_PF_FP_ABST
Abstract
Description
Cancer cell identification method and cancer cell identification system
[0001] The present invention relates to a cancer cell identification method and a cancer cell identification system that utilizes computer-based machine learning to screen and identify cancer cells from unstained cell images.
[0002] In recent years, the number of deaths from cancer has been increasing in Japan. Therefore, circulating tumor cell (CTC) testing, a blood test that is less invasive for patients, has attracted attention. One method for detecting CTCs in blood is to use fluorescent staining images to detect cancer cells among blood cells. CTC testing is less invasive for patients, and the number of CTCs in blood is expected to enable prognosis prediction and treatment efficacy assessment. However, while blood contains 4,000,000–5,500,000 red blood cells per μL and 5,000–9,000 white blood cells per μL, there are only a few CTCs per μL, making their presence in blood significantly less than that of other blood cells. Therefore, the sensitivity of CTC detection in blood is crucial for CTC testing.
[0003] Methods for detecting CTCs in blood include filtering, which focuses on differences in cell size, and immunoaffinity separation, which focuses on the expression of epithelial cell adhesion molecule (EpCAM) on the cell surface of CTCs. The CellSearch system, one immunoaffinity separation method, refines epithelial-derived cells using magnetic beads coated with anti-EpCAM antibodies. However, the refined cells, even with the CellSearch system, are contaminated with many leukocytes, necessitating a process for separating CTCs from leukocytes. This cell separation involves fluorescent antibody staining for cytokeratin, which is abundantly expressed in epithelial-derived cells, and CD45, which is specifically expressed in leukocytes. Cells that are cytokeratin-positive and CD45-negative are identified, allowing for the number of CTCs to be counted. Immunocytochemistry (IC) is used to detect CTCs using antibodies. Immunocytochemistry is a method for detecting intracellular antigens using antibodies, allowing the visualization of the presence of target components. In immunocytochemistry, a protein-specific antibody called a primary antibody is used to bind to the specific protein or other biomolecule being investigated, and the antigen is then detected by binding a secondary antibody or fluorescent label to the primary antibody.
[0004] As shown in Figure 1(a), CTC testing, which uses immunocytochemical staining to detect CTCs in blood, involves performing immunocytochemical staining 3 on collected test cells 2 to distinguish CTCs from other blood cells. The fluorescently stained cells 4 are then magnified under a fluorescent microscope 6 to observe the antibody reaction, and the fluorescently stained images 8 are then manually observed to detect CTCs. However, only a limited number of institutions are able to observe immunocytochemically stained cells. Furthermore, in addition to cell seeding, immunocytochemical staining requires complex procedures such as membrane permeabilization, blocking, antibody reaction, and counterstaining. The primary antibody reaction takes 1–2 hours, and the secondary antibody reaction takes approximately 1 hour, resulting in a significant workload.
[0005] To address this issue, cell detection methods exist that use unstained images of test cells 2, such as bright-field images, phase-contrast images, and differential interference contrast images. As shown in Figure 1(b), unstained images 16 can be captured using a standard microscope 7. Collected test cells 2 are magnified using the standard microscope 7 to observe antibody reactions. This eliminates the need for the complex staining procedures required for immunocytochemical staining, and the use of artificial intelligence (AI) makes testing easy. Therefore, if CTC detection were possible using unstained images 16, not only would the complex and time-consuming immunocytochemical staining process be unnecessary, but it would also be accessible to most institutions with standard microscopes, making CTC testing readily available at many institutions.
[0006] Additionally, Patent Document 1 discloses a system for automatically generating and evaluating digitally stained images of biological tissues and biopsy specimens. This system acquires multispectral images of the specimen, digitally stains the specimen images, and automatically evaluates the digitally stained images simultaneously with the acquisition of the multispectral images. This system can shorten the time required for evaluation by a histopathologist, thereby shortening the time required before the start of individual diagnostic evaluation by the histopathologist, which serves as the basis for a final decision.
[0007] Patent Document 2 discloses a colorization device for colorizing ordinary monochrome moving images, which includes a feature extractor that compresses monochrome images and extracts feature quantities, a feature corrector that corrects the feature quantities based on auxiliary information input from outside, a color estimator that estimates color information for the monochrome image from the feature quantities, model storage means that stores a color information estimation model that has been trained in advance using a neural network, and switching means that switches whether the feature quantity corrector corrects the feature quantities.The color estimator then estimates color information from the feature quantities before and after correction.
[0008] Patent Literature 3 discloses a method for breast cancer testing, comprising a target image acquisition unit and an estimation unit, in which the target image acquisition unit acquires target image data consisting of images of an HE-stained specimen of a breast tissue sample, and the estimation unit inputs the acquired target image data into an estimation model trained using training data, thereby estimating the test results of a HER2 test on the breast tissue sample corresponding to the target image data. The training data includes tested image data and test result data of a HER2 test based on an IHC method on the breast tissue sample corresponding to the tested image data. The tested image data consists of images of an HE-stained specimen of a breast tissue sample that has already been pathologically diagnosed as invasive breast cancer. As a result, the test results of the HER2 test on the breast tissue sample corresponding to the target image data are estimated from the target image data consisting of images of the HE-stained specimen of the breast tissue sample. Therefore, prior to treatment of invasive breast cancer with Herceptin, the state of HER2 gene amplification and HER2 protein overexpression can be estimated in a short period of time, and when the test results of the HER2 test are estimated using the above-mentioned trained estimation model, the test results can be estimated with significantly higher accuracy than when a person estimates the test results of the HER2 test by observing an HE-stained specimen.
[0009] JP 2014-526700 A JP 2020-144778 A JP 2021-136892 A
[0010] CTC testing using unstained images of blood cells as described above does not provide information on the presence or intensity of antibody expression, as observed in fluorescently stained images. Therefore, the accuracy of cell classification tends to be lower when using unstained images. Furthermore, conventional methods for detecting cancer cells using AI-based deep learning simply apply existing machine learning models to unstained images and classify cancer cells and blood cells, and do not propose machine learning models unique to CTC testing using unstained images. Using unstained images for cancer cell classification using AI presents two challenges: 1) reduced cell detection accuracy due to the low brightness of the background and cell regions in unstained images; 2) reduced classification accuracy when classifying cancer cells and blood cells.
[0011] Furthermore, the digital stain image disclosed in Patent Document 1 uses cells of biological tissue, but is not intended for blood cells or CTCs, and there is no disclosure of classifying or identifying CTCs. Similarly, the colorization device disclosed in Patent Document 2 simply colorizes ordinary monochrome images, but is not intended for blood cells or CTCs, and is not intended for classifying or identifying CTCs. The breast cancer testing method disclosed in Patent Document 3 uses a machine-learned estimation model, but is not intended for blood cells or CTCs, and is not intended for classifying or identifying CTCs.
[0012] The present invention has been made in view of the above-mentioned background art, and aims to provide a cancer cell identification method and a cancer cell identification system that can identify cancer cells by CTC detection using unstained images and classify cancer cells easily and with high accuracy.
[0013] This invention is a cell classification algorithm that detects and captures test target cells in blood using a CTC-chip or CTC enrichment method to capture test target cells, photographs the captured test target cells, estimates a fluorescently stained image from the unstained image, and uses the estimated fluorescent image to identify cancer cells. This algorithm is composed of two machine learning models. The first model is an image estimation model that converts an unstained image into a fluorescently stained image, and the second model is a classification model that actually classifies the image. This invention aims to improve the accuracy of cancer cell classification using unstained images by converting an unstained image into a fluorescently stained image and using that image for cell classification.
[0014] The present invention relates to a cancer cell identification method for identifying cancer cells using blood cell images obtained by photographing cells in collected blood, the method comprising: capturing test subject cells in the blood; photographing the captured test subject cells; using machine learning to learn a pair of unstained images of the test subject cells and fluorescently stained images of the test subject cells photographed after fluorescent staining with an antibody; estimating the fluorescently stained images from the unstained images of the test subject cells to generate an estimated fluorescent image; and identifying and classifying the generated estimated fluorescent image; the machine learning comprises image generation learning using a CGAN to learn a pair of the fluorescently stained images of the test subject cells photographed after fluorescent staining and the unstained images to generate the estimated fluorescent image; and image identification learning using a CNN to learn the unstained images and the fluorescently stained images as a single piece of cell information to identify cancer cells from the unstained images; and using the CGAN to generate the estimated fluorescent image based on the image generation learning. The unstained image and the corresponding estimated fluorescent image are input to the CNN as one piece of cell information, and the input unstained image is identified and classified using the CNN based on the image identification learning to determine whether or not the image is a cancer cell (claim 1).
[0015] In particular, for one unstained image, the estimated fluorescence images using a plurality of antibodies or markers are generated in parallel, and these are combined to form one piece of cell information (claim 2).
[0016] The machine learning involves training the CGAN using the classification results from the CNN, and the training involves feeding back to the CGAN the estimated fluorescence images that were misclassified in the classification identified by the CNN, performing image generation training for the CGAN, and reconstructing training data for the estimated fluorescence images (claim 4). Furthermore, the machine learning involves performing the image generation training by the training using the reconstructed training data, and then again determining whether the CNN made a classification error, and repeating the process of training using the reconstructed training data (claim 5).
[0017] The estimated fluorescence image may be used as training data for the CNN (claims 3 and 6). The estimated fluorescence image of the unstained image and a fluorescently stained image of the test subject cell may be input to the CNN (claim 7).
[0018] For the fluorescent staining, for example, any one of DAPI fluorescent dye, Cytokeratin antibody, CD45 antibody, p75NTR antibody, EpCAM antibody, and CSV antibody can be used (claim 8).
[0019] The present invention also provides a cancer cell identification system for identifying cancer cells using blood cell images of cells in collected blood, which uses machine learning to learn a pair of unstained images of the test subject cells and a fluorescently stained image of the test subject cells fluorescently stained with an antibody and then photographed, estimates the fluorescently stained image from the unstained image of the test subject cells to generate an estimated fluorescent image, and identifies and classifies the cells using the estimated fluorescent image; the cancer cell identification system comprises: a blood cell image photographing device that captures the test subject cells in the blood and photographs blood cell images of the captured test subject cells; a CGAN that uses machine learning to learn a pair of fluorescently stained images of the test subject cells fluorescently stained with an antibody and the unstained image, and performs image generation training to estimate the fluorescently stained image from the unstained image of the test subject cells to generate the estimated fluorescent image; a CNN that receives input of the unstained image and the corresponding estimated fluorescent image as a single piece of cell information, performs machine learning in advance using the fluorescent stained image of the test subject cell to determine whether the unstained image is a cancer cell, and performs image discrimination learning to identify whether the input unstained image is a cancer cell; when the unstained image of cells in the blood of the test subject is input, the CGAN generates the estimated fluorescent image from the unstained image based on the image generation learning, and the CNN identifies and classifies whether the input unstained image is a cancer cell from the unstained image and the corresponding estimated fluorescent image based on the image discrimination learning (claim 9).
[0020] For each unstained image, the estimated fluorescent images using multiple antibodies or markers are generated in parallel, and these are combined to form a single piece of cell information. The antibodies used for the fluorescent staining are selected from DAPI fluorescent dye, Cytokeratin antibody, CD45 antibody, p75NTR antibody, EpCAM antibody, and CSV antibody (Claim 10).
[0021] The machine learning of the cancer cell identification system performs additional training of the CGAN using the classification results of the CNN, and the additional training of the CGAN involves feeding back to the CGAN the estimated fluorescence images that were misclassified in the classification performed by the CNN, performing image generation training of the CGAN, and reconstructing training data for the estimated fluorescence images (Claim 11).Furthermore, the cancer cell identification system performs the image generation training by additional training using the reconstructed training data, and then again determines whether the classification by the CNN is incorrect, and repeats the process of additional training using the reconstructed training data (Claim 12).
[0022] According to the cancer cell identification method and cancer cell identification system of this invention, by using fluorescent stained images estimated from unstained images for cell classification, it is possible to capture information about cells in the blood of the test subject that cannot be obtained from unstained images, thereby improving the accuracy of cancer cell identification and cell classification using unstained images. Furthermore, the high accuracy of cell classification in unstained images eliminates the need for CTC testing using fluorescent stained images, which requires the labor and time required for immune cell staining and a special microscope. This makes it possible to reduce the cost of fluorescent reagents, fluorescent staining time, fluorescent observation equipment, and fluorescent observation time, making it easier to detect cancer cells and improving the accuracy of cancer testing.
[0023] FIG. 1A is a flow diagram showing a CTC inspection method using fluorescently stained cells by immune cell staining, which is background art of the present invention; and FIG. 1B is a flow diagram showing the concept of a CTC inspection method using unstained cell images of the present invention. FIG. 1C is a diagram showing an algorithm of a cancer cell identification method using unstained images according to one embodiment of the present invention. FIG. 1D is a diagram showing an algorithm of an embodiment of the present invention for estimating a fluorescently stained image using an unstained image to generate an estimated fluorescent image. FIG. 1E is a diagram showing an algorithm of an embodiment of the present invention for classifying an estimated fluorescent image estimated from an unstained image. FIG. 1F is a diagram showing a learning method of a cancer cell identification method using unstained images according to one embodiment of the present invention. FIG. 1F is a diagram showing a learning method of a cancer cell identification method using unstained images and a classification evaluation algorithm in a cell classification algorithm according to one embodiment of the present invention. FIG. 1F is a diagram showing an algorithm for follow-up learning by feedback in a learning method of a cancer cell identification method using unstained images according to one embodiment of the present invention. FIG. 1F is a diagram showing examples of unstained images and fluorescently stained images used in the cancer cell identification method of the present invention, and a generated estimated fluorescent image. FIG. 1F is a graph comparing the classification accuracy of a cancer cell identification method according to one embodiment of the present invention with a case where only unstained images are used. FIG. 1F is a diagram showing an algorithm of a cancer cell identification method using unstained images, which is an application example of an embodiment of the present invention. 1 is a diagram showing an algorithm for a cancer cell identification method using an unstained image according to another embodiment of the present invention. 2 is a graph comparing the accuracy of a cancer cell identification method according to another embodiment of the present invention with that of a conventional identification method.
[0024] A cancer cell identification method and cancer cell identification system 10 according to one embodiment of the present invention will be described with reference to the drawings. As shown in FIG. 1(b), the cancer cell identification method and cancer cell identification system 10 of this embodiment identify and classify cancer cells using deep learning with AI (artificial intelligence) configured in a computer system. The cancer cell identification system 10 includes a device (not shown) for capturing CTCs, which are test cells in the blood, using a CTC-chip, which is an element for capturing CTCs in the blood, or other physical or biological CTC enrichment methods. It also includes a blood cell imaging device (not shown) for capturing images of the captured test cells in the blood.
[0025] Furthermore, as described below, the cancer cell identification system 10 performs image generation learning in advance through machine learning, in which a fluorescently stained image 8 of a test subject cell 2 fluorescently stained with an antibody and an unstained image 16 of the fluorescently stained image 8 are paired together. As shown in Figures 2 and 3, the system includes a Conditional Generative Adversarial Network (CGAN) 12 that generates an estimated fluorescent image 18 by estimating the fluorescently stained image 8 from the unstained image 16 of the test subject cell 2 in the blood. The CGAN 12 that generates the image is composed of two neural networks, a Generator and a Discriminator, and these two neural networks are trained in an adversarial manner (GAN: Generative Adversarial Network), and are trained by linking appropriate conditions. By simultaneously inputting information as conditions to the Generator and Discriminator, the CGAN 12 is able to generate more precisely specified images. Specifically, in order to compensate for the low contrast of unstained images and the lack of stained antibody information and to estimate fluorescently stained images from unstained images, CGAN 12 learns pairs of unstained images 16 of each cell and fluorescently stained images 8 of each antibody, as described below, and trains an estimation model to understand how the stained antibodies are expressed.
[0026] In the fluorescent staining image estimation algorithm of the cancer cell identification system 10 of this embodiment, the CGAN 12 outputs multiple estimated fluorescent images 18, as shown in Figure 2. This is because, just as in the case of an actual cytotechnologist who comprehensively determines a cell based on multiple antibody expression information, this algorithm also uses multiple antibody information. Therefore, for one unstained image 16, estimated fluorescent images 18 are generated in parallel for each of the multiple antibodies, and all of these are combined to form a single piece of cell information.
[0027] Furthermore, as shown in FIGS. 2 and 4 , the cancer cell identification system 10 includes a convolutional neural network (CNN) 14, which receives an unstained image 16 and its corresponding estimated fluorescent image 18 as a single piece of cell information and identifies and classifies images based on the unstained image 16 and its corresponding estimated fluorescent image 18. The CNN 14 is a type of neural network used for image recognition. Although not shown, the CNN 14 has a structure in which a convolutional layer for feature extraction and a pooling layer for consolidating local features are repeated. The CNN 14 is used not only for image recognition but also as a base network for segmentation, object detection, pose estimation, and other applications. The CNN 14 is a classifier that classifies input images. As described below, in order to identify and classify the input unstained image 16 as a cancer cell, image identification training is performed in advance by machine learning to learn the fluorescently stained image 8 of the test cell 2 and the unstained image 16.
[0028] In the cell classification algorithm of the cancer cell identification method of this embodiment, an estimated fluorescent image 18, which is a fluorescent stained image estimated from an unstained image 16 by a CGAN 12, is input to a classifier in a CNN 14 to identify and classify cancer cells. Specifically, the image input to the CNN 14 is classified as either a cancer cell or a white blood cell (WBC) in the blood.
[0029] Next, a learning method according to an embodiment of the present invention will be described below. First, a fluorescently stained image will be described. For example, when distinguishing between KYSE140, a type of cancer cell, and WBC, fluorescent staining is performed using three types of dyes: DAPI (4',6-diamidino-2-phenylindole), a cytokeratin antibody, and a CD45 antibody. Because DAPI is a fluorescent dye that binds to DNA, both KYSE140, which contains cell nuclei, and WBC will be positive. Cytokeratin is expressed in epithelial cells, and cancers are typically organized into epithelial cells. Therefore, fluorescent staining with the cytokeratin antibody will result in KYSE140 being positive and WBC, which are blood cells, being negative. On the other hand, CD45 is expressed in hematopoietic cells excluding red blood cells and platelets. Therefore, fluorescent staining with the CD45 antibody will result in WBC being positive and KYSE140 being negative. Thus, by utilizing the fact that cell staining results in differences between positive and negative results depending on the characteristics of the staining antibody, cancer cells can be identified and classified from fluorescently stained images. In this embodiment, the use of multiple types of antibodies enables accurate cell classification, so multiple estimated fluorescence images 18 are generated in parallel for each of the multiple types of antibodies for one unstained image 16, and all of these are input into the cell classification algorithm as cell information for a single test cell.
[0030] As shown in Figure 5, the machine learning for classification begins with training of the classifier CNN 14. Training of the estimator CGAN 12 involves additional training using the trained CNN 14. Therefore, training of CNN 14 is performed first, followed by training of CGAN 12 using CNN 14. After that, estimation of fluorescently stained images is performed using CGAN 12, and the estimated images are classified by CNN 14, whereby the final classification accuracy is confirmed and classification work is performed. This process from training of CNN 14 to evaluation of classification accuracy using the estimated fluorescent image 18 is performed for one sample.
[0031] As shown in Figure 6, the CNN 14 that performs image discrimination learning inputs and learns from microscopic images during the learning process, combining one unstained image 16 and three fluorescently stained images 8 of each antibody as a single piece of cell information. Furthermore, estimated fluorescence images 18 estimated from the unstained images 16 are used for cell classification, thereby improving the classification accuracy of the unstained images 16. Therefore, estimated fluorescence images 18 estimated from the unstained images 16 are used to evaluate the accuracy of cell classification. At this time, the unstained images 16 used for conversion to estimated fluorescence images 18 are also input, so a total of four images are input for evaluation: one unstained image 16 actually captured for each cell and three estimated fluorescence images 18 of each antibody estimated from the unstained images 16.
[0032] The CGAN 12 is trained by inputting a pair of an actually captured unstained image 16 and a fluorescently stained image 8, or, in the case of fake images, a pair of the fake image and an associated image, and the discriminator of the CGAN 12 is trained to determine whether the combination of the input pair of images is Real or Fake. In this embodiment, the unstained image 16 is used as the associated image, and the corresponding actually captured fluorescently stained image 8 is input as the Real image for training. In this way, the generator of the CGAN 12 learns to convert the unstained image 16 into an estimated fluorescence image 18 of the fluorescently stained image 8. The trained generator can then be used to generate an estimated fluorescence image 18 of the fluorescently stained image 8 from the unstained image 16.
[0033] Furthermore, as shown in FIG. 7 , to improve the accuracy of fluorescence estimation, the CGAN 12 feeds back to the CGAN 12 the classification results of the CNN 14 used for image classification during the learning process when generating the estimated fluorescence image 18 in the fluorescence estimation algorithm. Misclassified images are then retrained by the CGAN 12 to generate the correct estimated fluorescence image 18. After a certain amount of training has been performed by the CGAN 12, the unstained image 16 used for training is passed through the CGAN 12's estimator to convert it into an estimated fluorescence image 18. After conversion, the image is passed through the trained classifier of the CNN 14, which determines whether it is a cancer cell or a WBC. Furthermore, if an incorrect determination is made during this determination, the image is considered a weak point of the CGAN 12's estimator, and the training data is reconstructed around that image. The CGAN 12's estimator is then retrained using the reconstructed training data, and this process is repeated a certain number of times. The CNN 14 classifier is then used again. If a misclassification occurs, the training data is reconstructed and further training is performed. This process is repeated, and training ends when the CGAN 12 estimation model reaches its optimum state.
[0034] Next, we will explain the results of an accuracy evaluation of cell classification using a cancer cell identification method and the algorithm of the cancer cell identification system 10 according to one embodiment of the present invention. The classification accuracy obtained from the algorithm of the cancer cell identification system 10 of the present invention was compared with the accuracy of a model that classifies only unstained images as input, and significance was verified. Fluorescence estimated images generated from the fluorescence estimation model were also evaluated. For the accuracy evaluation, 400 training data images and 100 validation data images were extracted from 2,907 images of the cancer cell line KYSE140 and 1,300 images of white blood cells, and 100 samples of the extracted data were used.
[0035] <Estimation Results of the Fluorescence Estimation Algorithm> To evaluate the accuracy of the fluorescent stain image estimation algorithm of this embodiment, the difference between the actual image and the estimated image was calculated using RMSE. Here, the luminance of the actually captured image was used as the true value, and the luminance of the estimated image was used as the predicted value. RMSE is calculated as the square root of the root mean square error (MSE), which is the square of the difference between the true value and the predicted value and the average of these values. The smaller this value, the smaller the difference between the true value and the predicted value, and the smaller the error. Figure 8 shows the estimated fluorescent images generated by the algorithm of this embodiment for each unstained image, along with their RMSE. As shown in Figure 8, the actual fluorescent stained images (real images) and the estimated fluorescent images estimated by CGAN are highly accurate, as they estimate the state of DAPI fluorescent dye expression in the nucleus. Furthermore, the RMSE ranged from approximately 0.5 to 7.0, indicating that the estimation closely matched the actual image, even when compared to the RMSE of approximately 0.65 for the noise level calculated from the captured images.
[0036] <Classification Results of Cell Classification Algorithm> Accuracy, precision, recall, and F-measure were used to evaluate classification accuracy, and a Welch t-test with a significance level of 0.1% was used to test for significance. For Accuracy, Precision, Recall, and F-measure, higher values indicate higher accuracy. Accuracy is the percentage of correct answers among all data, indicating accuracy. Precision is an index of the model's accuracy, with higher values indicating fewer incorrect classifications. Recall is an index of the model's comprehensiveness, with higher values indicating the degree to which positive predictions were actually recovered from all data that should have been predicted. F-measure is the harmonic mean of Precision and Recall, which are in a trade-off relationship. It is an index used to equally evaluate Precision and Recall, allowing for a comprehensive assessment of accuracy and comprehensiveness. Figure 9 shows Accuracy, Precision, Recall, and F-measure on the horizontal axis, and the classification accuracy evaluation index (evaluation index) on the vertical axis (as a percentage). In the figure, "a" shows the result using only unstained images, and "b" shows the result of classification using unstained images and estimated fluorescent images according to this embodiment. As a result of the evaluation, as shown in Figure 9, a 3% improvement in accuracy was observed, and statistical significance was confirmed in all other items, including precision, recall, and F-measure. This confirmed the improvement in classification accuracy achieved by the proposed cell classification algorithm using unstained images.
[0037] According to the cancer cell identification method and cancer cell identification system 10 of the embodiment of the present invention, by using an estimated fluorescent image of a fluorescently stained image estimated from an unstained image for cell classification, antibody information that cannot be obtained from an unstained image can be obtained, improving the accuracy of cell classification in unstained images. Furthermore, the high accuracy of cell classification in unstained images eliminates the need for CTC testing using fluorescently stained images, which requires the time and effort of conventional immune cell staining and a special microscope. This makes it possible to reduce the cost of fluorescent reagents, the time required for fluorescent staining, the fluorescence observation device, and the time required for fluorescence observation, making it easier to detect cancer cells and improving the accuracy of cancer testing.
[0038] Next, an application example of the cancer cell identification method and cancer cell identification system 10 of the present invention will be described below. In the above embodiment, the CGAN 12, which trains on fluorescent stained images and estimates them, and the CNN 14, which performs cell classification, are trained using images actually captured using cell stained images of three types of antibodies. Here, an advantage of the fluorescent stained image is that multiple fluorescent stained images can be created if a trained model for the antibody is available. As shown in Figure 10 , by training multiple fluorescent stained images using multiple appropriate antibodies, many fluorescent stained images can be used for cell classification. However, in the learning algorithm of this embodiment, the CNN used for cell classification is trained using actually captured images. Therefore, when increasing the number of antibodies, securing CNN training data increases the burden of capturing time, microscope limitations, and staining work.
[0039] Therefore, a cancer cell identification method and cancer cell identification system 20 according to another embodiment of the present invention uses actual fluorescently stained images 8 as training data for CNN 14, as well as estimated fluorescent images 18, as shown in Figure 11. With the algorithm of the cancer cell identification method and cancer cell identification system 20 according to this embodiment, if a trained estimation model is available, fluorescently stained images can be obtained without the need for cumbersome processes.
[0040] Furthermore, due to the above advantages, this cancer cell identification method and cancer cell identification system 20 can also be applied to carcinomas with unstable antibody expression. In this embodiment, experiments were conducted using MiaPACA2 and Panc1, cell lines derived from pancreatic cancer. Because these cell lines exhibit unstable expression of the cancer marker CSV (Cell-Surface Vimentin), classification requires the use of other antibodies, such as EpCAM antibodies. Therefore, using the algorithm of this embodiment of the cancer cell identification system 20, an algorithm was used to input estimated EpCAM antibody fluorescence images into images stained with DAPI fluorescent dye, CSV antibody, and CD45 antibody, and an evaluation was conducted using a validation dataset of 100 samples. The evaluation was conducted to determine whether or not the classification accuracy of the CNN 14 was reduced by training the CNN 14 using only the captured images and when training with estimated fluorescence images 18.
[0041] The results are shown in Figure 12. The 4-channel (Conventional) graph on the left side of Figure 12 represents the conventional method, which uses unstained images observed under a microscope and images stained with DAPI fluorescent dye, CSV antibody, and CD45 antibody. The graph on the right side represents unstained images observed under a microscope and images stained with DAPI fluorescent dye, CSV antibody, and CD45 antibody, with the addition of an estimated EpCAM antibody fluorescence image. As shown in Figure 12, accuracy improved by 1.25%, and the variance of accuracy also decreased by 0.6, from 2.32 to 1.72. A t-test at a significance level of 0.1% confirmed the significant difference. These results demonstrate that the use of estimated fluorescence images can improve classification accuracy not only in unstained images alone, but also in fluorescently stained images, confirming its applicability to carcinomas with unstable antibody expression.
[0042] The cancer cell identification method and cancer cell identification system of the present invention are not limited to the above-described embodiments, and the antibodies, fluorescent dyes, and other markers used for fluorescent staining and generating estimated images can be selected appropriately, including p75NTR antibodies, etc. The cancer cells to be applied are also those that can be identified and classified for various cancer cells through estimation learning of fluorescent stained images.
[0043] This invention improves the accuracy of cell classification using unstained images, eliminating the need for fluorescent staining and fluorescent observation. This allows cancer cells to be easily identified using only an optical microscope, and using a computer system makes it possible to more easily and accurately classify and identify cancer cells. This makes it possible to identify cancer cells and detect cancer early, even in general medical institutions equipped with optical microscopes. Furthermore, because unstained images can be obtained using an optical microscope and cell staining is not required, it is expected to reduce the burden on workers and promote the widespread use of cancer cell testing.
[0044] 2 Test cell 4 Fluorescent stained cell 6 Fluorescence microscope 8 Fluorescent stained image 10, 20 Cancer cell identification system 12 CGAN (Conditional Generative Adversarial Network) 14 CNN (Convolutional Neural Network) 16 Unstained image 18 Estimated fluorescent image
Claims
1. A cancer cell identification method for identifying cancer cells using blood cell images of cells in collected blood, the method comprising: capturing test subject cells in the blood and photographing the captured test subject cells; using machine learning to previously learn a pair of unstained images of the test subject cells and a fluorescently stained image of the test subject cells photographed after fluorescently staining them with an antibody; estimating the fluorescently stained image from the unstained image of the test subject cells to generate an estimated fluorescent image; and identifying and classifying the generated estimated fluorescent image; the machine learning comprises: using a CGAN to learn a pair of the fluorescently stained image of the test subject cells photographed after fluorescent staining and the unstained image to generate the estimated fluorescent image; and using a CNN to learn the unstained image and the fluorescently stained image as a single piece of cell information to perform image identification training to identify cancer cells from the unstained image; using the CGAN to generate the estimated fluorescent image from the unstained image based on the image generation training; A cancer cell identification method characterized by inputting the unstained image and the corresponding estimated fluorescent image into the CNN as one piece of cell information, and using the CNN to identify and classify whether the input unstained image is a cancer cell based on the image identification learning.
2. A cancer cell identification method according to claim 1, wherein the estimated fluorescence images using a plurality of antibodies or markers are generated in parallel for each unstained image, and these are combined to form a single piece of cell information.
3. A cancer cell identification method according to claim 1 or 2, wherein the estimated fluorescence image is used as training data for the CNN.
4. The cancer cell identification method according to claim 1 or 2, wherein the machine learning involves performing additional training of the CGAN using the classification results of the CNN, and the additional training involves feeding back to the CGAN the estimated fluorescence images that were misclassified in the classification identified by the CNN, performing image generation training of the CGAN, and reconstructing training data for the estimated fluorescence images.
5. A cancer cell identification method according to claim 3, wherein the image generation learning is performed by the additional learning using the reconstructed learning data, and then the classification error by the CNN is determined again, and the additional learning process using the reconstructed learning data is repeated.
6. The cancer cell identification method according to claim 5, wherein the estimated fluorescence image is used as training data for the CNN.
7. A cancer cell identification method according to claim 1 or 2, wherein the estimated fluorescent image of the unstained image and a fluorescently stained image of the test subject cell are input to the CNN, and the unstained image is classified.
8. A method for identifying cancer cells according to claim 1 or 2, wherein the fluorescent staining is performed using a dye selected from the group consisting of DAPI fluorescent dye, Cytokeratin antibody, CD45 antibody, p75NTR antibody, EpCAM antibody, and CSV antibody.
9. A cancer cell identification system for identifying cancer cells using blood cell images taken of cells in collected blood, the system using machine learning to train a pair of unstained images of the test subject cells and fluorescently stained images of the test subject cells taken after fluorescently staining them with an antibody, estimates the fluorescently stained images from the unstained images of the test subject cells to generate an estimated fluorescent image, and identifies and classifies the cells using the estimated fluorescent image; the system includes: a blood cell image capturing device that captures the test subject cells in the blood and captures blood cell images of the captured test subject cells; a CGAN that uses machine learning to train a pair of fluorescently stained images of the test subject cells taken with an antibody and the unstained images, and performs image generation training to estimate the fluorescently stained images from the unstained images of the test subject cells to generate the estimated fluorescent image; a CNN that receives the unstained image and the corresponding estimated fluorescent image as input as one piece of cell information, performs machine learning in advance using the fluorescent stained image of the test subject cell to determine whether the unstained image is a cancer cell, and performs image discrimination learning to identify whether the input unstained image is a cancer cell, wherein when the unstained image of cells in the blood of the test subject is input, the CGAN generates the estimated fluorescent image from the unstained image based on the image generation learning, and the CNN identifies and classifies whether the input unstained image is a cancer cell from the unstained image and the corresponding estimated fluorescent image based on the image discrimination learning.
10. A cancer cell identification system as described in claim 9, wherein for each unstained image, each of the estimated fluorescent images using multiple antibodies or markers is generated in parallel, and these are combined to form a single piece of cell information, and the antibodies used for the fluorescent staining are selected from DAPI fluorescent dye, Cytokeratin antibody, CD45 antibody, p75NTR antibody, EpCAM antibody, and CSV antibody.
11. A cancer cell identification system as described in claim 9 or 10, wherein the machine learning of the cancer cell identification system uses the classification results of the CNN to perform additional training of the CGAN, and the additional training of the CGAN involves feeding back to the CGAN the estimated fluorescence images that were misclassified in the classification identified by the CNN, performing image generation training of the CGAN, and reconstructing training data for the estimated fluorescence images.
12. A cancer cell identification system as described in claim 11, wherein the image generation learning is performed by the additional learning using the reconstructed learning data, and then the classification error by the CNN is determined again, and the additional learning process using the reconstructed learning data is repeated.
Citation Information
Patent Citations
Convolutional neural networks for cancer diagnosis
US9739783B1
Virtual staining logic
WO2021198252A1
Label-free virtual immunohistochemical staining of tissue using deep learning
WO2023107844A1