Apparatus for classifying circulating tumor cells from cell images and classification method using same

The device and method automate the classification of circulating tumor cells from cell images by using a computing device and a deep neural network, addressing inefficiencies in existing methods and enhancing the accuracy and speed of cancer diagnosis.

WO2025110352A1PCT designated stage expired Publication Date: 2025-05-30CYTOGEN CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/002450
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing methods for classifying circulating tumor cells (CTCs) from cell images are inefficient and require manual confirmation, which can be time-consuming and prone to errors.

Method used

A device and method that utilize a computing device to process cell images, extract morphological variables, and input them into a learned CTC classification model, such as a deep neural network, to accurately and quickly classify CTCs from images stained with specific biomarkers.

Benefits of technology

The solution enables rapid and accurate classification of CTCs, improving the efficiency and reliability of cancer diagnosis by automating the classification process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024002450_30052025_PF_FP_ABST
    Figure KR2024002450_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to an apparatus for classifying circulating tumor cells (CTCs) from cell images and a classification method using same, in which CTCs can be accurately and quickly classified from cell images stained with a specific biomarker. The purpose of the present invention is to provide a method, a computer program, and a computing device for classifying CTCs from cell images.
Need to check novelty before this filing date? Find Prior Art

Description

Device for classifying circulating tumor cells from cell images and classification method using the same

[0001] The present invention was made under the support of the Ministry of Trade, Industry and Energy under the project identification number 1415187522 and project number 20008829. The research management specialized institution of the project is the Korea Institute of Industrial Technology Evaluation and Planning, the research project name is “Customized Diagnostic Treatment Product”, the research project name is “Development of Commercialization Technology for Precision Cancer Diagnosis Platform Based on Circulating Tumor Cells and Tumor-Derived Exosomes of Liquid Biopsy”, the main institution is Cytogen Co., Ltd., and the research period is from 2023-01-01 to 2023-12-31.

[0002] The present invention relates to a device for classifying circulating tumor cells (CTCs) from a cell image and a classification method using the same, and more particularly, to a device capable of accurately and quickly classifying circulating tumor cells from a cell image stained with a specific biomarker and a classification method using the same.

[0003] Tissue biopsies, the primary method used in traditional cancer diagnosis, directly extract a patient's tissue to determine the presence and type of cancer. While generally accurate, this method can be physically burdensome for patients and requires specialized equipment and personnel at hospitals. Recently, interest in liquid biopsy, which tests circulating tumor cells (CTCs) in the blood, has been growing.

[0004] Among liquid biopsy techniques, the "In situ Padlock Probe Assay" is particularly noteworthy. This technique targets and stains specific biomarkers within cells. The staining results can be used to distinguish CTCs from normal cells, thereby identifying the presence of cancer.

[0005] Previously, CTC classification work utilized software called CellProfiler to preprocess stained CTC images and extract feature data. Afterwards, the CTCs were identified using CellProfiler Analysis software, which verified them visually.

[0006] However, recent research trends have shifted beyond the "in situ padlock probe assay" to include the use of immunofluorescence, which focuses on and stains specific proteins in cells. Research has shown that this technology, by analyzing the characteristics of biomarker images using machine learning, is effective in more precisely classifying cancer. This is expected to significantly improve the accuracy and efficiency of cancer diagnosis, and is emerging as a hot topic among researchers and medical professionals.

[0007] The present inventors have developed a device for classifying circulating tumor cells (CTCs) from cell images and a classification method using the same, and confirmed that circulating tumor cells can be quickly classified from images stained with a specific biomarker.

[0008] Accordingly, an object of the present invention is to provide a method for classifying circulating tumor cells (CTCs) in cell images.

[0009] Another object of the present invention is to provide a computer program for classifying circulating tumor cells in cell images.

[0010] Another object of the present invention is to provide a computing device for classifying circulating tumor cells in cell images.

[0011] The present invention relates to a device for classifying circulating tumor cells from a cell image and a classification method using the same, which can accurately and quickly classify circulating tumor cells from a cell image stained with a specific biomarker.

[0012] Hereinafter, the present invention will be described in more detail.

[0013]

[0014] One aspect of the present invention is a method for classifying circulating tumor cells (CTCs) in a cell image using a computing device, the method comprising: a processing step of generating input data based on data including a cell image; an extraction step of extracting morphological variables from the input data; and a generation step of inputting the morphological variables into a learned circulating tumor cell classification model to generate classification information.

[0015] The term “cell image” as used herein may include multidimensional cell image data composed of discrete image elements (e.g., pixels in a two-dimensional image). Specifically, a cell image may include a visible object or a digital representation of the object, and for example, a cell image may be a cell image obtained from a blood sample obtained from a subject and then stained with a specific biomarker and then photographed.

[0016] The term "circulating tumor cells (CTCs)" used herein refers to cancer cells that shed from cancerous tissue and migrate into the bloodstream. These cells are known to play a crucial role in the spread and metastasis of cancer, and analyzing CTCs in blood samples can noninvasively assess cancer progression and treatment response. CTCs are considered a useful indicator for cancer diagnosis and prognosis, and their quantity and characteristics are directly related to patient survival.

[0017] The term "subject" as used herein may be an object (or subject) that has or is suspected of having cancer or a tumorous disease. In one embodiment of the present invention, the subject may be a mammal, for example, the subject may be one or more selected from the group consisting of humans, monkeys, dogs, cats, mice, rats, cows, horses, pigs, goats, and sheep.

[0018] The term "morphological variable" in this specification refers to a variable that represents a characteristic related to the shape of an object or structure within image data. Morphological analysis is primarily used to extract morphological features, such as the size, shape, surrounding boundaries, texture, and density of objects within an image or video.

[0019] In one embodiment of the present invention, the circulating tumor cell classification model may be a deep neural network.

[0020] The term “deep neural network” as used herein may refer to a neural network that includes multiple hidden layers in addition to input and output layers. The term “deep neural network” may be used interchangeably with the terms “neural network,” “network function,” and “neural network” throughout this specification. Using a deep neural network, it is possible to identify latent structures of data. That is, it is possible to identify latent structures of photos, text, videos, voices, and music (e.g., what objects are in a photo, what the content and emotion of a text is, what the content and emotion of a voice is, etc.). In the present invention, a deep neural network may include, but is not limited to, a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a Q network, a U network, a Siamese network, and the like.

[0021] In one embodiment of the present invention, the cell image may be stained with one or more biomarkers selected from the group consisting of DAPI, EpCAM, CD45, and Vimentin.

[0022] In one embodiment of the present invention, the morphological variable may include at least one selected from the group consisting of the area of ​​the object (Tatal Area), the area ratio (% Area), the perimeter of the object (Perimeter), the circularity (Circularity), the solidity (Solidity), the major axis distance (Feret), the minor axis distance (MinFeret), and the inregrated density (Inregulated Density).

[0023] In one embodiment of the present invention, the processing step may include a removal step of removing non-cellular image background values ​​from the cell image by applying a filter.

[0024] In one embodiment of the present invention, the filter may be a mean filter.

[0025] In one embodiment of the present invention, the processing step may include a distinguishing step of distinguishing a single cell from a cell cluster in the cell image.

[0026] In one embodiment of the present invention, the processing step may include an equalization step that equalizes the staining intensity of the cell image by applying MinMaxScaling.

[0027] In this embodiment of the present invention, the processing step may include a recognition step of recognizing a cell object by removing extracellular noise from the cell image.

[0028] In one embodiment of the present invention, the recognition step may be performed through an object recognition algorithm.

[0029] In the present invention, the circulating tumor cell classification model is supervised learning based on one or more learning images and guide labels corresponding to the learning images and including the circulating tumor cell classification results, and the supervised learning may be performed based on a comparison result between the learning information generated for the learning images using the circulating tumor cell classification model and the guide labels.

[0030] In the present invention, supervised learning may be performed based on a result value calculated by substituting learning information and the guide label into a binary cross-entropy loss function or a log loss function.

[0031] In one embodiment of the present invention, the processing step may include applying a robust filter to remove the background of the cell image and applying MinMaxScaling.

[0032] The term "robust filter" in this specification refers to a filter designed to be unduly affected by specific types of noise or outliers when removing noise or outliers from an image or signal. A robust filter is insensitive to outliers or extreme values, effectively removing noise while preserving the fundamental characteristics of the data.

[0033] Another aspect of the present invention is a computer program stored in a storage medium, which, when executed on one or more processors, causes the computer program to perform the following operations for classifying circulating tumor cells (CTCs), the operations including: a generation operation for generating input data based on data including cell images; an extraction operation for extracting morphological variables from the input data; and a generation operation for inputting morphological variables into a learned circulating tumor cell classification model to generate classification information.

[0034] Another aspect of the present invention is a computing device for classifying circulating tumor cells (CTCs), comprising: a processor including one or more cores; and a memory; wherein the processor generates input data based on data including cell images, extracts morphological variables from the input data, and inputs the morphological variables into a learned circulating tumor cell classification model to generate classification information.

[0035] The present invention relates to a device for classifying circulating tumor cells from a cell image and a classification method using the same, which can accurately and quickly classify circulating tumor cells from a cell image stained with a specific biomarker.

[0036] FIG. 1 is a block diagram illustrating a computing device that performs an operation to classify circulating tumor cells (CTCs) from a cell image according to one embodiment of the present invention.

[0037] Figure 2 is a schematic diagram illustrating a deep neural network according to one embodiment of the present invention.

[0038] FIG. 3 is a block diagram illustrating a process for classifying circulating tumor cells from cell images according to one embodiment of the present invention.

[0039] FIG. 4 is a diagram illustrating an overview of primary and secondary data according to one embodiment of the present invention.

[0040] FIG. 5 is a diagram illustrating an overview of 3rd party data according to one embodiment of the present invention.

[0041] FIG. 6 is a diagram showing a synthesis process of primary and secondary data according to one embodiment of the present invention.

[0042] FIG. 7 is a diagram showing the importance of variables in predicting 3rd data after learning 1st data in morphological analysis using DAPI, EpCAM, and CD45 biomarkers according to one embodiment of the present invention.

[0043] FIG. 8 is a diagram showing the importance of variables in predicting primary data after learning 3rd data in morphological analysis using DAPI, EpCAM, and CD45 biomarkers according to one embodiment of the present invention.

[0044] FIG. 9 is a diagram showing the process and result of applying a dyeing robustness filter according to one embodiment of the present invention.

[0045] A method for classifying circulating tumor cells (CTCs) in a cell image using a computing device, comprising: a processing step of generating input data based on data including a cell image; an extraction step of extracting morphological variables from the input data; and a generation step of inputting the morphological variables into a learned circulating tumor cell classification model to generate classification information.

[0046] Below, with reference to the attached drawings, embodiments of the present invention are described in detail so that those skilled in the art can easily implement the invention. However, the present invention can be implemented in various different forms and is not limited to the embodiments described herein. In addition, in the drawings, parts irrelevant to the description are omitted for clarity of description, and similar reference numerals are used throughout the specification to indicate similar parts.

[0047] Throughout the specification, whenever a part is said to "include" a component, this means that it may include other components, but not to the exclusion of other components, unless otherwise stated. The terms "and / or" and "and / or" include any and all combinations of the associated listed items.

[0048] Additionally, the terms “…part” and “…module” described in the specification mean a unit that processes at least one function or operation, which may be implemented by hardware, software, or a combination of hardware and software.

[0049] It should be appreciated that the various illustrative logical blocks, configurations, modules, circuits, means, logics, and algorithm steps described in connection with the embodiments herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, configurations, means, logics, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application. However, such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0050]

[0051] FIG. 1 is a block diagram illustrating a computing device that performs an operation to classify circulating tumor cells (CTCs) from a cell image according to one embodiment of the present invention.

[0052] Referring to FIG. 1, a computing device (1000) that performs an operation of classifying circulating tumor cells from a cell image according to one embodiment may include a processor (100) and a memory (200).

[0053] The processor (100) can perform an operation of generating input data based on data including cell images, extracting morphological variables from the input data, and inputting the morphological variables into a learned circulating tumor cell classification model to generate classification information.

[0054] The processor (100) may, when generating input data based on data including cell images, perform an operation of removing non-cellular image background values ​​from the cell images by applying a filter. At this time, the processor may perform an operation of removing pixels having non-cellular values ​​from the cell images based on a specific filter. The filter is not limited to any filter known in the art for removing a specific portion from an image, but, for example, an operation of removing non-cellular images may be performed using a median filter. When a median filter is used to remove non-cellular background, the filter sets a window of a fixed size (e.g., 3x3, 5x5) centered on each pixel of the image, calculates an average of all pixel values ​​within the window, and sets the calculated average value as a new value for the pixel, thereby reducing high-frequency noise in the image and smoothly modifying the image. Thereafter, a difference is calculated between the image smoothed through the median filter and the original image, and a threshold value is set based on the difference image to distinguish between the background and the target object (cell), and then the background value can be removed.

[0055] The processor (100) may perform an operation of equalizing the staining intensity of the cell image by applying maximum-minimum scaling (MinMaxScaling) when generating input data based on data including a cell image. At this time, in addition to maximum-minimum scaling, the processor may equalize the staining intensity by any other known method capable of equalizing the staining color intensity in the cell image. When equalizing the intensity of the staining image by maximum-minimum scaling, all pixel values ​​of the image are first examined to obtain the minimum value (min) and the maximum value (max), and a new value for each pixel is calculated using the MinMaxScaling formula to equalize the pixel values ​​within a specific range (e.g., [0, 255] or [0, 1]). When the staining intensity in the image is equalized by the processor (100), it may be easy to learn the classification of a subsequent circulating tumor cell classification model, and thus the classification accuracy of circulating tumor cells through the learned circulating tumor cell classification model may be improved.

[0056] Specifically, the processor (100) can remove noise by calculating the mean of a cell image using a median filter and converting pixel values ​​below the mean in the image to 0 to remove them. In addition, when processing multiple images, the processor (100) can calculate the mean of pixels for each cell image for maximum-minimum scaling, calculate the minimum (Min) and maximum (Max) of the mean of each image, and scale the image pixels of each image between the minimum and maximum to equalize the staining intensity of each image.

[0057] Meanwhile, the processor (100) can remove the background from the cell images and equalize the staining intensity of each image by using a robust filter to remove the background and equalize the staining intensity of the aforementioned images. The robust filter has a characteristic of being robust to outliers or noise, and can be used for image processing tasks such as background removal or color equalization. It can remove backgrounds other than cells while removing noise and outliers of the image (preserving main features), generate a histogram of the image, and equalize the color or brightness of the image.

[0058] The processor (100) can perform an operation of counting objects in a cell image. Specifically, the processor (100) can apply an object recognition algorithm to process pixels in the cell image that are below or above a preset threshold value into black and white, and can recognize and remove objects having a typical cell shape, i.e., a circle or an abnormal size, as noise for each cell image.

[0059] The cell image may be an image of cells obtained by staining a blood sample obtained from a subject, and specifically, may be an image of cells stained with a specific biomarker. In this case, the biomarker may be used without limitation as long as it is a biomarker for distinguishing circulating tumor cells from other cells, and for example, the biomarker may be one or more biomarkers selected from the group consisting of DAPI, EpCAM, CD45, and Vimentin.

[0060] A cell image may be a merged image of individual cells stained with one or more biomarkers. For example, when using images of individual cells stained with the biomarkers DAPI, EpCAM, and CD45, the input cell image may be a merged image of cells stained with the biomarkers DAPI, EpCAM, and CD45.

[0061] The processor (100) may perform an operation of extracting morphological variables from input data. At this time, the morphological variables mean variables that represent characteristics related to the shape of an object or structure within the image data. In one embodiment of the present invention, the morphological variables may include one or more selected from the group consisting of the total area of ​​the object, the area ratio (% Area), the perimeter of the object, the circularity, the solidity, the major axis distance (Feret), the minor axis distance (MinFeret), and the intensity of staining (Inregrated Density), but are not limited thereto, and other known morphological variables may be used without limitation.

[0062] The processor (100) can perform an operation of distinguishing between single cells and cell clusters in a cell image when generating input data based on data including cell images. In the cell classification operation, if there are cell clusters along with single cells, smooth learning of a circulating tumor cell classification model may be difficult, and the classification performance of the constructed classification model may deteriorate. Therefore, the processor (100) can perform an operation of distinguishing between single cells and cell clusters, and at this time, the processor can distinguish single cells through object identification using specific software.

[0063] The classification information may include the results of classifying objects recognized in the cell image as circulating tumor cells or other cells. For example, when classifying circulating tumor cells and peripheral blood mononuclear cells (PBMCs), the classification information may include the results of classifying each object recognized in the cell image into two types: circulating tumor cells and peripheral blood mononuclear cells. Alternatively, the classification information may include a score indicating the degree of affinity of the recognized object with circulating tumor cells. For example, when the degree of affinity with circulating tumor cells is classified on a scale of 1 to 10, the classification information may include the results of assigning a score of 10 if it is definitely circulating tumor cells and a score of 1 if it is definitely not circulating tumor cells.

[0064] The learned CTC classification model may be a deep neural network (DNN). The DNN may include an input layer and an output layer, or may include multiple hidden layers separate from the input layer and the output layer. The DNN may include a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a Q network, a U network, a Siamese network, etc. For example, the DNN may be a convolutional neural network.

[0065] The CTC classification model may be trained through transfer learning. Transfer learning involves applying the knowledge of a model learned on a specific task to a different, similar task. This can involve taking a model trained on a large dataset (e.g., ImageNet) for image classification and retraining it to fit a smaller dataset in a specific domain. For example, in the present invention, the CTC classification model may be a model trained on a large dataset, retrained on a dataset for distinguishing between CTCs and other cells (e.g., PBMCs).

[0066] A convolutional neural network is a type of multilayer perceptron and may include a neural network that includes convolutional layers. A convolutional neural network can utilize weights in the computational process through the neural network. A convolutional neural network may be composed of one or more convolutional layers and neural network layers combined with them. A convolutional layer can extract features from input data using a filter. At this time, the convolutional layer may include a filter and an activation function that converts the filter into a non-linear value. A convolutional neural network can process image data by representing it as a matrix with dimensions, and thus, a convolutional neural network can be used to recognize objects in an image. For example, image data encoded in red, green, and blue can be represented as a two-dimensional matrix for each of the R, G, and B colors. That is, the color value of each pixel can be an element of the matrix, and the size of the matrix can be the same as the size of the image. Convolutional neural networks can include a pooling layer, which allows them to utilize input data with a two-dimensional structure.

[0067] A convolutional neural network may include one or more convolutional layers and sub-sampling layers. A sub-sampling layer may be connected to the output of a convolutional layer to simplify the output of the convolutional layer. For example, when the output of a convolutional layer is input to a pooling layer having a 2*2 average pooling filter, the image can be compressed by outputting the average value contained in each 2*2 patch from each pixel of the image. The aforementioned pooling may output the minimum value from the patch or the maximum value from the patch, and any pooling method may be used. A convolutional neural network can extract features from a given image by repeatedly performing a convolution process and a sub-sampling process such as pooling. At this time, the output from the convolutional layer and / or the sub-sampling layer may be input to a fully connected layer. A fully connected layer is a layer in which all neurons in one layer are connected to all neurons in the neighboring layer.

[0068] The processor (100) may be configured with one or more cores and may include a processor for data analysis and deep learning, such as a central processing unit (CPU), a graphics processing unit (GPU), and a tensor processing unit (TPU) of a computing device. The processor may read a computer program stored in a memory and perform data processing for machine learning according to an embodiment. According to an embodiment, the processor may perform operations for learning a neural network. The processor may perform calculations for learning a neural network, such as processing input data for learning in deep learning (DL), extracting features from input data, calculating errors, and updating weights of a neural network using backpropagation. At least one of the CPU, GPU, and TPU of the processor (100) may process learning of a network function.

[0069] The memory (200) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk.

[0070] Figure 2 is a schematic diagram illustrating a deep neural network according to one embodiment of the present invention.

[0071] A neural network can represent a model of a machine learning structure designed to extract feature data from input data and provide inference operations using the feature data. In this case, the feature data can represent data regarding features abstracted from the input data. For convenience of explanation, Figure 2 depicts the hidden layer as comprising three layers; however, the hidden layer can comprise a variety of layers. A neural network can comprise one or more layers, and each layer can comprise one or more nodes.

[0072] Nodes (or units) are elements that constitute each layer, and each layer can be composed of a node or a set of nodes. Within a neural network, nodes in layers other than the output layer can be connected to nodes in the next layer via links for transmitting output signals. At this time, nodes in each layer can be connected to each other via links, and nodes in connected layers can be in the relationship of input nodes and output nodes depending on whether signals are transmitted or received. Among the nodes connected via links, the data of the output node can have a value determined according to the data input to the input node. Each node included in the hidden layer can be input with the output of an activation function regarding the weighted inputs of the nodes included in the previous layer. The weighted input is the input of the nodes included in the previous layer with the weight reflected. At this time, the weight can be variable and can vary depending on the function and algorithm of the neural network. Weights can be referred to as parameters of a neural network, and activation functions can include sigmoid, hyperbolic tangent (tanh), and rectified linear unit (ReLU).

[0073] An initial input node may refer to one or more nodes within a neural network into which data is directly input without going through links with other nodes. Alternatively, within a neural network, it may refer to nodes that do not have other input nodes connected by links in the relationship between nodes based on links. Similarly, a final output node may refer to one or more nodes within a neural network that do not have output nodes in their relationship with other nodes. Furthermore, a hidden node may refer to nodes that constitute a neural network other than the initial input node and the final output node.

[0074] A deep neural network (DNN) can refer to a neural network that includes multiple hidden layers in addition to input and output layers. Using a deep neural network, one can identify latent structures in data. That is, one can identify the latent structures of photos, text, videos, voices, and music (e.g., what objects are in the photo, what the content and emotion of the text are, what the content and emotion of the voice are, etc.). Deep neural networks can include convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, generative adversarial networks (GANs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), Q networks, U networks, Siamese networks, and generative adversarial networks (GANs).

[0075] A convolutional neural network (CNN) is a type of deep learning model inspired by the structure of the visual cortex of animals, designed to process data with grid patterns, such as images. A convolutional neural network typically includes convolutional layers, pooling layers, and fully connected layers. Convolutional and pooling layers can exist repeatedly within a neural network, and input data can be transformed into output through these layers. Convolutional layers utilize kernels (or masks) to extract features. The element-wise product between each element of the kernel and the input value is computed at each location and summed to obtain the output, which is called a feature map. This process can be repeated, applying multiple kernels to form any number of feature maps. In a convolutional neural network, convolutional and pooling layers perform feature extraction, while fully connected layers map the extracted features to the final output, such as a classification operation.

[0076] Convolutional neural networks, such as neural networks, can be trained to minimize output errors. Separate from the forward propagation process, which extracts values ​​from the input layer to the output layer, backpropagation occurs within the neural network, calculating the error between the input training data and the corresponding neural network output values ​​and updating the weights of the nodes in each layer to reduce this error. The training process in convolutional neural networks can be summarized as finding the kernel that extracts the output values ​​with the lowest error based on the given training data. The kernel is the only parameter that is automatically learned during the training of the convolutional layer. On the other hand, in convolutional neural networks, kernel size, number of kernels, padding, etc. are hyperparameters that must be set before the training process begins. Therefore, convolutional neural network models can be distinguished based on the kernel size, number of kernels, and number of convolutional and pooling layers.

[0077] A neural network can be trained by at least one of supervised learning, which uses training data with labeled answers for each training data, unsupervised learning, semi-supervised learning, or reinforcement learning, in which the training data are not labeled with answers. In this case, an error can be calculated by comparing the output from the neural network with the label or training data, and the calculated error is backpropagated in the backward direction (i.e., from the output layer to the input layer) in the neural network, and the connection weights of each node in each layer of the neural network can be updated according to the backpropagation. The amount of change in the connection weights of each node that are updated can be determined according to a learning rate.

[0078] Overfitting occurs when neural networks learn excessively from training data, resulting in increased errors even as the number of training data increases. Overfitting can increase errors in machine learning algorithms, and various optimization methods can be used to prevent it. Methods to prevent overfitting include increasing the training data, regularization, dropout (inactivating some network nodes during the learning process), and the use of batch normalization layers.

[0079] FIG. 3 is a block diagram illustrating a process for classifying circulating tumor cells from cell images according to one embodiment of the present invention.

[0080] Referring to FIG. 3, a computing device that performs an operation of classifying circulating tumor cells from a cell image according to one embodiment can generate input data based on data including a cell image (S101), extract morphological variables from the input data (S102), and input the morphological variables into a learned circulating tumor cell classification model to generate classification information (S103).

[0081] When generating input data based on data including cell images, the computing device can perform an operation of applying a filter to remove non-cell image background values ​​from the cell images.

[0082] The computing device can perform an operation of equalizing the staining intensity of the cell image by applying maximum-minimum scaling (MinMaxScaling) when generating input data based on data including a cell image.

[0083] The computing device can perform an operation of distinguishing a single cell and a cell cluster in a cell image when generating input data based on data including a cell image.

[0084]

[0085] Hereinafter, the present invention will be described in more detail with reference to the following examples. However, these examples are only intended to illustrate the present invention, and the scope of the present invention is not limited by these examples.

[0086]

[0087] Example 1: First and Second Experimental Data

[0088] After artificially culturing circulating tumor cells (CTCs), primary and secondary data were constructed. The primary data was constructed for the purpose of classifying CTCs and peripheral blood mononuclear cells (PBMCs), and the secondary data was constructed for the purpose of classifying non-peripheral blood mononuclear cells (NBMs) and peripheral blood mononuclear cells (PBMCs).

[0089] The primary and secondary image data consisted of a 256x256 8bit-RGB image stained with DAPI, EpCAM, and CD45 biomarkers for one cell, and a Merge 256x256 8bit-RGB image that merged the three images into one. In this case, the primary data consisted of 1514 CTC data sheets and 520 PBMC data sheets, and the secondary data consisted of 4262 Non-PBMC data sheets and 250 PBMC data sheets (see Fig. 4).

[0090] During the data verification process, 29 data points that were designated as Non-PBMC and PBMC in the secondary data were identified and removed. Since the legend indicating the size of the merged image data cells existed, it was judged that this could act as noise during the analysis process, so the DAPI, EpCAM, and CD45 images were synthesized and used as shown in Figure 6.

[0091]

[0092] Example 2: Morphological analysis and deep learning application using tertiary clinical data

[0093] Data Overview

[0094] The tertiary data used clinical data extracted from actual patients, and was divided into early stage patients SMC-045, SMC-046, and SMC-047, and late stage patients SMC-015, SMC-019, and SMC-022 (see Figure 5). The tertiary data used the same biomarker as the 1st and 2nd data, plus vimentin. The data was divided into single cell data and cell cluster data, and the following experiments were conducted based on single cell data for clear analysis. The tertiary data consisted of 48 CTC data sheets and 1312 PBMC data sheets.

[0095]

[0096] Data preprocessing

[0097] The tertiary data, when provided, identified PBMCs as single cells and cell clusters. However, even within the single-cell data, some cells contained more than one cell. To identify only single cells, Fiji software was used to identify image objects and classify them as single cells. The entire single-cell data consisted of 35 CTCs and 768 PBMCs.

[0098]

[0099] Data feature extraction for morphological analysis

[0100] To conduct morphological analysis on the image data, Fiji software was used to extract a total of eight variables for each biomarker, as shown in Table 1 below.

[0101]

[0102] Variable name Description Toral Area Area of ​​the object %Area The percentage of the entire image occupied by the object Perimeter The perimeter of the object Circularity The circularity of the object Solidity The degree of curvature of the object Feret The longest distance between two points on the boundary of the object (extension distance) MinFeret The shortest distance between two points on the boundary of the object (inregulation distance) Inregrated Density The intensity (intensity) that can represent the brightest area in the image

[0103] Building models for deep learning analysis

[0104] The ResNet-18 model was utilized based on deep learning analysis. Because the amount of data was insufficient for the deep learning model to learn, a transfer learning technique was applied using weights pretrained on ImageNet data. CTC and PBMC classification was performed using image data classified as single cells.

[0105]

[0106] Prediction results

[0107] In a morphological analysis involving the DAPI, EpCAM, and CD45 biomarkers, the accuracy of the tertiary data validation process after primary data training was 96%, but the recall of CTCs was 6%. In the tertiary data validation process after primary data training, the feature variables of the EpCAM biomarker played a key role in the classification process (see Table 2 and Figure 7).

[0108]

[0109] Precision, Recall, F1-Score, Accuracy, CTC 100%, 6%, 11%, 96%, PBMC 96%, 100%, 98%

[0110] After the third data training, the first data validation process achieved a respectable accuracy of 94%. In the first data validation process after the third data training, the CD45 biomarker's feature variables played a key role in the classification process (see Table 3 and Figure 8).

[0111] Precision, Recall, F1-Score, Accuracy, CTC 94%, 97%, 95%, 94%, PBMC 93%, 86%, 89%

[0112] Analysis results using deep learning showed that for images using all of the DAPI, EpCAM, and CD45 biomarkers, the accuracy was 99% after the first data training and the third data verification (see Table 4), and the accuracy was 86% after the first data training and the third data verification (see Table 5).

[0113] Precision, Recall, F1-Score, Accuracy, CTC 93%, 80%, 86%, 99%, PBMC 99%, 100%, 99%

[0114] Precision, Recall, F1-Score, Accuracy, CTC 83%, 100%, 91%, 86%, PBMC 100%, 57%, 72%

[0115] Example 3: Morphological analysis and deep learning application using 4th clinical data

[0116] Data Overview

[0117] Similar to the tertiary data, this study consisted of clinical data obtained from actual patients. Unlike the tertiary clinical data, the Vimentin biomarker was not used. For cross-validation, analysis was conducted based on single-cell data, with 68 CTC data sheets and 157 PBMC data sheets.

[0118]

[0119] Applying data dyeing robustness filter

[0120] To equalize the staining intensity between data, a data staining robustness filter was applied. The process of applying the staining robustness filter included applying a mean filter to the image to remove background values ​​that had values ​​other than cells, and then applying minmax scaling to equalize the data staining intensity. To address the issue of noise being equalized in a specific image, object counting was performed, and if two or more objects were captured during the process, they were judged to be noise and removed.

[0121]

[0122] Data feature extraction for morphological analysis

[0123] To conduct morphological analysis on image data, feature variables were extracted using the same process as data feature extraction in the 3rd clinical data.

[0124]

[0125] Cross-validation of morphological analysis and deep learning analysis

[0126] To ensure clear cross-validation, in addition to training and validating data by round, a mixed-data validation process was added. The second-round data, which included non-PBMC and PBMC classification data, was excluded due to differences from the first, third, and fourth rounds (CTC and PBMC classification data). The purpose of this study was to understand the performance of the robustness filter by comparing data before and after applying the dye robustness filter.

[0127]

[0128] Prediction results

[0129] In the process of verifying the morphological analysis model, the accuracy was generally over 96% before applying the robustness filter, but in the process of verifying the 3rd data for the 4th data learning, the recall of CTC dropped to 3% (see Table 6).

[0130] After applying the robustness filter, the CTC recall rate of the 4th data learning and 3rd data verification increased to 100% during the model verification process (see Table 7).

[0131]

[0132] Applying robustness filter X class Precision Recall F1-Score Accuracy 3rd data training 4th data verification CTC 97% 100% 98% 99% PBMC 100% 99% 99% 4th data training 3rd data verification CTC 100% 3% 6% 96% PBMC 96% 100% 98% 1+ 3rd data training 4th data verification CTC 97% 100% 99% 99% PBMC 100% 99% 99% 1+ 4th data training 3rd data verification CTC 100% 97% 99% 99% PBMC 100% 100% 100% 1+ 3+ 4th data training 1+ 3+ 4th data verification CTC 96% 100% 98% 98% PBMC 100% 96% 98%

[0133] Applying robust filter O class Precision Recall F1-Score Accuracy 3rd data training 4th data verification CTC 92% 100% 96% 97% PBMC 100% 96% 98% 4th data training 3rd data verification CTC 97% 100% 99% 99% PBMC 100% 100% 100% 1+3rd data training 4th data verification CTC 92% 100% 96% 97% PBMC 100% 96% 98% 1+4th data training 3rd data verification CTC 97% 97% 97% 99% PBMC 100% 100% 100% 1+3+4th data training 1+3+4th data Validation CTC96%100%98%98%PBMC100%96%98%

[0134] During the deep learning analysis model validation process, accuracy decreased by approximately 20% compared to other cases, reaching 74% when training with the 1st and 4th data sets and then verifying with the 3rd data set before applying the robustness filter (see Table 8). After applying the robustness filter, accuracy increased by 23% compared to before applying the filter, reaching 97% when training with the 1st and 4th data sets and then verifying with the 3rd data set (see Table 9).

[0135]

[0136] Applying robustness filter X class Precision Recall F1-Score Accuracy 3rd data training 4th data verification CTC 99% 99% 99% 99% PBMC 99% 99% 99% 4th data training 3rd data verification CTC 57% 100% 72% 96% PBMC 100% 96% 98% 1+ 3rd data training 4th data verification CTC 98% 100% 99% 99% PBMC 100% 99% 100% 1+ 4th data training 3rd data verification CTC 14% 100% 25% 74% PBMC 100% 73% 84% 1+ 3+ 4th data training 1+ 3+ 4th data verification CTC 95% 100% 97% 97% PBMC 100% 94% 97%

[0137] Applying robust filter O class Precision Recall F1-Score Accuracy 3rd data training 4th data verification CTC 90% 82% 86% 92% PBMC 93% 96% 94% 4th data training 3rd data verification CTC 85% 97% 91% 99% PBMC 100% 99% 100% 1+3rd data training 4th data verification CTC 92% 100% 96% 97% PBMC 100% 96% 98% 1+4th data training 3rd data verification CTC 57% 100% 73% 97% PBMC 100% 97% 98% 1+3+4th data training 1+3+4th data Verification CTC100%100%100%100%PBMC100%100%100%

[0138] The present invention relates to a device for classifying circulating tumor cells (CTCs) from a cell image and a classification method using the same, and more particularly, to a device capable of accurately and quickly classifying circulating tumor cells from a cell image stained with a specific biomarker and a classification method using the same.

Claims

1. A method for classifying circulating tumor cells (CTCs) in a cell image using a computing device, A processing step of generating input data based on data including cell images; An extraction step for extracting morphological variables from the above input data; and A generation step of generating classification information by inputting the above morphological variables into a learned circulating tumor cell classification model; A method for classifying circulating tumor cells in a cell image, comprising:

2. In paragraph 1, A method for classifying circulating tumor cells in a cell image, wherein the cell image is stained with one or more biomarkers selected from the group consisting of DAPI, EpCAM, CD45 and Vimentin.

3. In paragraph 1, A method for classifying CTCs in a cell image, wherein the morphological variables include at least one selected from the group consisting of total area, percent area, perimeter, circularity, solidity, major axis distance (Feret), minor axis distance (MinFeret), and inregrated density.

4. In paragraph 1, A method for classifying circulating tumor cells in a cell image, wherein the processing step includes a removal step of removing image background values ​​other than cells from the cell image by applying a filter.

5. In paragraph 4, A method for classifying circulating tumor cells in a cell image, wherein the above filter is a mean filter.

6. In paragraph 1, A method for classifying circulating tumor cells in a cell image, wherein the processing step includes an equalization step for equalizing the staining intensity of the cell image by applying MinMaxScaling.

7. In paragraph 1, A method for classifying circulating tumor cells in a cell image, wherein the above circulating tumor cell classification model is a convolutional neural network (CNN)-based model.

8. In the first paragraph, the circulating tumor cell classification model, Supervised learning is performed based on one or more learning images and guide labels corresponding to the learning images and including the results of circulating tumor cell classification. A method for classifying circulating tumor cells in cell images, wherein the above supervised learning is performed based on the comparison results of learning information generated for learning images using a circulating tumor cell classification model and guide labels.

9. In paragraph 8, the supervised learning is A method for classifying circulating tumor cells in a cell image, the method being performed based on a result value calculated by substituting the above learning information and the above guide label into a binary cross-entropy loss function or a log loss function.

10. In paragraph 1, A method for classifying circulating tumor cells in a cell image, wherein the processing step includes an application step of applying a robust filter to remove the background of the cell image and applying MinMaxScaling.

11. A computer program stored in a storage medium, When a computer program runs on more than one processor, The following actions are performed to classify circulating tumor cells (CTCs): The above actions are: A generating operation that generates input data based on data containing cell images; An extraction operation for extracting morphological variables from the above input data; and A generation operation that generates classification information by inputting the above morphological variables into a learned CTC classification model; A computer program stored on a storage medium, comprising:

12. In paragraph 11, A computer program stored in a storage medium, wherein the cell image is stained with one or more biomarkers selected from the group consisting of DAPI, EpCAM, CD45 and Vimentin.

13. In paragraph 11, A computer program stored in a storage medium, wherein the above morphological variables include at least one selected from the group consisting of total area, percent area, perimeter, circularity, solidity, major axis distance (Feret), minor axis distance (MinFeret), and inregrated density.

14. A computing device for classifying circulating tumor cells (CTCs), A processor comprising one or more cores; and memory; including; The processor is, A computing device that generates input data based on data including cell images, extracts morphological variables from the input data, and inputs the morphological variables into a learned circulating tumor cell classification model to generate classification information.

15. In paragraph 14, A computing device, wherein the above morphological variables include at least one selected from the group consisting of total area, percent area, perimeter, circularity, solidity, major axis distance, minor axis distance, and inregrated density.

Citation Information

Patent Citations

  • Integrated management system for employment support projects integrating and managing employment support projects in conjunction with company's personnel management system and government's employment support system

    KR1020240159193A

  • Put-on type mounted work platform of main tower for bridge, and mounting method of the work platform

    KR102537075B1

  • KR20230037339A