Methods, apparatuses and computer programs for assisting disease analysis, and methods, apparatuses and programs for training computer algorithms
By classifying and extracting features from cell images, and utilizing deep learning and machine learning algorithms, the problem of inaccurate disease identification in existing technologies has been solved, achieving high-precision diagnosis of hematopoietic system diseases.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-24
- Publication Date
- 2026-03-31
AI Technical Summary
Current technologies have failed to accurately identify diseases based on individual cell images, especially hematopoietic system diseases.
By classifying multiple analyte cell images collected from specimens, deep learning and machine learning algorithms, especially gradient boosting trees, are used to identify cell morphology and abnormalities, thereby identifying diseases.
It enables high-precision identification of diseases, especially accurate diagnosis of hematopoietic system diseases such as aplastic anemia and myelodysplastic syndrome.
Smart Images

Figure CN111861975B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to methods, apparatus, and computer programs for assisting in disease analysis, as well as methods, apparatus, and programs for training computer algorithms for assisting in disease analysis. Background Technology
[0002] Patent document 1 describes a method that involves inputting the following data from a tissue image into a neural network: the number, area, shape, roundness, color, and chromaticity of the nuclear region; the number, area, shape, and roundness of the lacunar region; the number, area, shape, roundness, color, and chromaticity of the stroma region; the number, area, shape, and roundness of the lumen region; image texture; feature quantities calculated using wavelet transform values; the degree to which epithelial cells show a double-layer structure with myoepithelial cells; the degree of fibrosis; the presence or absence of papillary patterns; the presence or absence of sieve patterns; the presence or absence of necrotic material; the presence or absence of solid patterns; and feature quantities calculated using the color or chromaticity of the image. This data is then used to differentiate pathological tissues of "scleroderma patterns" and "intratubular fibroma patterns."
[0003] Existing technical documents
[0004] Patent documents
[0005] Patent Document 1: Japanese Patent Application Publication No. 10-197522 Summary of the Invention
[0006] The problem that the invention aims to solve
[0007] Although Patent Document 1 discloses a scheme for identifying diseases based on tissue images, it does not disclose a method for identifying diseases based on individual cell images.
[0008] One of the objectives of this invention is to identify diseases with high accuracy from individual cell images.
[0009] Methods for solving problems
[0010] This invention relates to a method for assisting in disease analysis. In this method, the morphology of each analytical cell is classified based on images obtained from multiple analytical target cells contained in a sample collected from a subject. Based on the classification results, relevant information corresponding to the cell morphology classification of the sample is obtained. Based on this cell morphology classification information, a computer algorithm is used to analyze the diseases present in the subject. Based on these components, diseases can be identified based on individual cell images.
[0011] Preferably, in the process of classifying the morphology of the cells of each analytical object, the cell type of each analytical object is identified. More preferably, the relevant information for the cell morphology classification is the information on the number of cells of each cell type (64). Based on these components, diseases can be identified based on each cell type.
[0012] Preferably, in the process of classifying the morphology of the cells of each analytical object, the types of abnormalities observed in the cells of each analytical object are identified. More preferably, the relevant information for the cell morphology classification is information on the number of cells of each type of abnormality observed (63). Based on these configurations, diseases can be identified based on the types of abnormalities observed in each cell.
[0013] In the process of classifying the morphology of cells in each of the above-mentioned analytical objects, the types of abnormalities observed are identified for each cell type in each analytical object. Based on these characteristics, diseases can be identified more effectively based on the precision of individual cell images.
[0014] In the process of classifying the morphology of the cells in each analysis object, analytical data containing relevant information about each cell is input into a deep learning algorithm with a neural network structure. This deep learning algorithm is then used to classify the morphology of each cell. Based on these characteristics, diseases can be identified with greater accuracy.
[0015] The aforementioned computer algorithm is a machine learning algorithm. The analysis of the diseases present in the subject is performed by inputting the relevant information of the cell morphology classification as features into the machine learning algorithm (67). Based on these components, diseases can be identified with greater accuracy.
[0016] The machine learning algorithm (67) is preferably selected from trees, regression, neural networks, Bayesian methods, clustering, or ensemble learning. More preferably, the aforementioned machine learning algorithm is gradient boosting trees. By using these machine learning algorithms, diseases can be identified more accurately.
[0017] In the process of obtaining the aforementioned cell morphology classification information, the probability that each analyzed cell belongs to each category in multiple cell morphology classifications is calculated, and the sum of the probabilities for each category in the aforementioned cell morphology classifications is obtained as the relevant information for the aforementioned cell morphology classification. Based on these components, disease identification can be performed with higher accuracy.
[0018] Preferably, the sample is a blood sample. Blood cells reflect the condition of various diseases, thus enabling more accurate disease identification.
[0019] Preferably, the aforementioned diseases are hematopoietic system diseases. According to the present invention, hematopoietic system diseases can be identified with high accuracy.
[0020] The aforementioned hematopoietic system diseases are aplastic anemia or myelodysplastic syndrome. According to the present invention, hematopoietic system diseases can be identified with high accuracy.
[0021] Preferably, the aforementioned abnormalities are selected from at least one of the following: nuclear morphology abnormalities, granular abnormalities, abnormal cell size, cell malformation, cell destruction, vacuoles, immature cells, presence of inclusion bodies, Dürer bodies, satellite phenomena, abnormal nuclear network, petal-like nuclei, high N / C ratio, bleb-like morphology, fragmentation (smudge), and hairy cell-like morphology. Evaluating these abnormalities in these cells allows for more precise disease identification.
[0022] Preferably, the aforementioned nuclear morphological abnormalities include at least one selected from the following: highly lobed, poorly lobed, pseudo-Perger's nuclear abnormalities, ring-shaped nuclei, spherical nuclei, elliptical nuclei, apoptosis, polynuclearity, nuclear collapse, enucleation, naked nuclei, irregular nuclear margins, nuclear fragmentation, intranuclear bridging, multinucleation, split nuclei, nuclear division, and nucleolar abnormalities. The aforementioned granular abnormalities include at least one selected from the following: degranulation, abnormal granule distribution, toxic granules, Orl bodies, bundled cells, and pseudo-Chediak-Higashi granular granules. The aforementioned cell size abnormalities include giant platelets. By evaluating these abnormalities, disease identification can be performed with greater precision.
[0023] Preferably, the cell types mentioned above include at least one selected from neutrophils, eosinophils, platelets, lymphocytes, monocytes, and basophils. Evaluating these cell types allows for more precise disease identification.
[0024] More preferably, the aforementioned cell types further include at least one selected from metamyelocytes, myelocytes, promyelocytes, granulocytes, plasma cells, atypical lymphocytes, immature eosinophils, immature basophils, erythroblasts, and megakaryocytes. By evaluating these cell types within these cell ranges, disease identification can be performed with greater precision.
[0025] The present invention relates to an apparatus (200) for assisting in disease analysis. The apparatus (200) includes a processing unit (20) that classifies the morphology of each analytical target cell based on images obtained from multiple analytical target cells contained in a sample collected from a subject, obtains relevant information corresponding to the cell morphology classification of the sample based on the classification results, and analyzes the disease present in the subject using a computer algorithm based on the relevant information of the cell morphology classification.
[0026] This invention relates to a program for assisting disease analysis, which enables a computer to perform the following steps: classifying the morphology of each analytical cell based on images obtained from multiple analytical cell samples collected from a subject, obtaining relevant information corresponding to the cell morphology classification of the sample based on the classification results; and analyzing the disease present in the subject based on the relevant information of the cell morphology classification and using a computer algorithm.
[0027] High-precision disease identification is possible based on the devices or programs used to assist in disease analysis.
[0028] This invention relates to a training method for a computer algorithm used to assist in disease analysis. The training method includes: classifying the morphology of each analytical cell based on images obtained from multiple analytical cell samples collected from a subject; obtaining relevant information corresponding to the cell morphology classification of the sample based on the classification results; inputting the obtained cell morphology classification information as first training data and the disease information of the subject as second training data into the computer algorithm.
[0029] The present invention relates to a training apparatus (100) for a computer algorithm for assisting in disease analysis. The training apparatus (100) includes a processing unit (10) that classifies the morphology of each analytical cell based on images obtained from multiple analytical cells contained in a sample collected from a subject, obtains relevant information corresponding to the cell morphology classification of the sample based on the classification results, and inputs the obtained cell morphology classification information as first training data and the disease information (55) of the subject as second training data into the computer algorithm.
[0030] This invention relates to a training program for a computer algorithm used to assist in disease analysis. The computer performs the following steps: classifying the morphology of each analytical cell based on images obtained from multiple analytical cell samples collected from a subject, and obtaining relevant information corresponding to the cell morphology classification of the sample based on the classification results; inputting the relevant information of the cell morphology classification as first training data and the disease information (55) of the subject as second training data into the computer algorithm.
[0031] Based on the training method, training device (100) or training procedure, high-precision disease identification can be performed.
[0032] The effects of the invention
[0033] According to the present invention, diseases can be identified with high precision based on individual cell images. Attached Figure Description
[0034] Figure 1 A diagram illustrating an outline of the auxiliary method using a recognizer.
[0035] Figure 2 This is a schematic diagram illustrating the steps for generating training data for deep learning and the training steps for the first deep learning algorithm, the second deep learning algorithm, and the machine learning algorithm.
[0036] Figure 3 Examples of label values are shown.
[0037] Figure 4 This shows an example of using training data for machine learning.
[0038] Figure 5 A schematic diagram illustrating an example of the steps involved in generating analytical data and the steps involved in disease analysis using computer algorithms.
[0039] Figure 6 A diagram illustrating an example of the structure of disease analysis system 1.
[0040] Figure 7 A block diagram illustrating an example of the hardware configuration of the supplier-side device 100.
[0041] Figure 8 A block diagram illustrating an example of the hardware configuration of the user-side device 200.
[0042] Figure 9 This is a block diagram illustrating an example of the function of the training device 100A.
[0043] Figure 10 A flowchart illustrating an example of the deep learning processing flow.
[0044] Figure 11 This is a schematic diagram used to illustrate a neural network.
[0045] Figure 12 A block diagram illustrating the functionality of the machine learning device 100A.
[0046] Figure 13 A flowchart illustrating an example of a machine learning process using the first piece of information.
[0047] Figure 14 A flowchart illustrating an example of a machine learning process using the first and second information.
[0048] Figure 15 This is a block diagram illustrating an example of the function of the disease analysis device 200A.
[0049] Figure 16 A flowchart illustrating an example of a disease analysis process using the first piece of information.
[0050] Figure 17 A flowchart illustrating an example of a disease analysis process using information 1 and information 2.
[0051] Figure 18 A diagram illustrating an example of the structure of the disease analysis system 2.
[0052] Figure 19 This is a block diagram illustrating an example of the functionality of the comprehensive disease analysis device 200B.
[0053] Figure 20 A diagram illustrating an example of the composition of the disease analysis system 3.
[0054] Figure 21 This is a block diagram illustrating an example of the functionality of the integrated image analysis device 100B.
[0055] Figure 22 A diagram illustrating the structure of the recognizer used in the embodiments.
[0056] Figure 23 A table is provided showing the number of cells used as training data for a deep learning algorithm and the number of cells used in validation, which is used to evaluate the performance of the trained deep learning algorithm.
[0057] Figure 24 A table showing the results of evaluating the performance of the trained second deep learning algorithm.
[0058] Figure 25 A table showing the results of evaluating the performance of the first deep learning algorithm after training.
[0059] Figure 26 A heatmap showing abnormalities that can aid in disease analysis.
[0060] Figure 27 A graph showing the ROC curves of the disease analysis results. Detailed Implementation
[0061] The following describes in detail the methods for carrying out the invention with reference to the accompanying drawings. It should be noted that in the following description and drawings, the same symbols denote the same or similar constituent elements; therefore, descriptions of the same or similar constituent elements are omitted.
[0062] Methods for assisting in the analysis of diseases present in the subject (hereinafter sometimes simply referred to as "auxiliary methods") are described. The auxiliary methods include: classifying the morphology of each analytical object cell, and analyzing the diseases present in the subject based on the classification results. The morphology of each analytical object cell is classified based on images obtained from multiple analytical object cells contained in a sample collected from the subject, and relevant information corresponding to the cell morphology classification of the sample is obtained based on the classification results. The auxiliary methods also include: analyzing the diseases present in the subject based on abnormality type information (hereinafter sometimes referred to as "first information"), which is related to cell morphology classification. First information is relevant information corresponding to the abnormality type detected in the sample, obtained based on the types of abnormalities detected in each of the multiple analytical object cells contained in the sample. Abnormalities are identified based on images taken of the analytical object cells. Furthermore, the auxiliary methods also include: analyzing the diseases present in the subject based on cell type information (hereinafter sometimes referred to as "second information"), which is related to cell morphology classification. The second piece of information is information related to the cell types mentioned above, obtained based on the multiple analyte cell types contained in the sample. Cell types are identified based on images taken of the analyte cells.
[0063] There are no restrictions on the subject of the examination, as long as it is an animal that will be the subject of the disease analysis. Examples of animals include humans, dogs, cats, rabbits, and monkeys. Humans are preferred.
[0064] There are no restrictions on the diseases that affect the animals mentioned above. For example, diseases can include tumors in tissues other than the hematopoietic system, hematopoietic system diseases, metabolic diseases, kidney diseases, infectious diseases, allergic diseases, autoimmune diseases, and injuries.
[0065] Tumors in tissues outside the hematopoietic system can include benign epithelial tumors, benign non-epithelial tumors, malignant epithelial tumors, and malignant non-epithelial tumors. Among tumors in tissues outside the hematopoietic system, malignant epithelial tumors and malignant non-epithelial tumors are preferred examples.
[0066] Diseases of the hematopoietic system include tumors, anemia, polycythemia, platelet abnormalities, and myelofibrosis. Tumors of the hematopoietic system preferably include myelodysplastic syndromes, leukemia (acute myeloid leukemia, acute myeloid leukemia (with neutrophil differentiation), acute promyelocytic leukemia, acute myelomonocytic leukemia, acute monocytic leukemia, erythroleukemia, acute megakaryocytic leukemia, acute myeloid leukemia, acute lymphoblastic leukemia, lymphoblastic leukemia, chronic myeloid leukemia, and chronic lymphocytic leukemia), malignant lymphomas (Hodgkin's lymphoma and non-Hodgkin's lymphoma), multiple myeloma, and granulomas. Malignant tumors of the hematopoietic system are preferably myelodysplastic syndromes, leukemia, and multiple myeloma, with myelodysplastic syndromes being more preferred.
[0067] Examples of anemia include aplastic anemia, iron deficiency anemia, megaloblastic anemia (including vitamin B12 deficiency, folic acid deficiency, etc.), hemorrhagic anemia, renal anemia, hemolytic anemia, thalassemia, sideroblastic anemia, and atransferrinemia. Among these, aplastic anemia, pernicious anemia, iron deficiency anemia, and sideroblastic anemia are preferred, with aplastic anemia being more preferred.
[0068] Polycythemia can include polycythemia vera and secondary polycythemia. Polycythemia vera is preferred as a type of polycythemia.
[0069] Platelet abnormalities can include thrombocytopenia, thrombocytosis, and megakaryocyte abnormalities. Thrombocytopenia can include disseminated intravascular coagulation, idiopathic thrombocytopenic purpura, MYH9 abnormality, Bernard-Soulier syndrome, etc. Thrombocytosis can include essential thrombocytemia. Megakaryocyte abnormalities can include mini-megakaryocytes, multinucleated megakaryocytes, platelet hypoplasia, etc.
[0070] Myelofibrosis can include primary myelofibrosis and secondary myelofibrosis.
[0071] Metabolic diseases can include disorders of carbohydrate metabolism, lipid metabolism, electrolyte imbalance, and metal metabolism. Disorders of carbohydrate metabolism include mucopolysaccharidosis, diabetes, and glycogenopathies. Mucopolysaccharidosis and diabetes are preferred examples of carbohydrate metabolism disorders. Disorders of lipid metabolism include Gaucher's disease, Niemann-Pick disease, hyperlipidemia, and arteriosclerotic diseases. Arteriosclerotic diseases include arteriosclerosis, atherosclerosis, thrombosis, and embolism. Electrolyte imbalances include hyperkalemia, hypokalemia, hypernatremia, and hyponatremia. Disorders of metal metabolism include iron metabolism disorders, copper metabolism disorders, calcium metabolism disorders, and inorganic phosphorus metabolism disorders.
[0072] Kidney damage can include nephrotic syndrome, decreased kidney function, acute kidney failure, chronic kidney disease, and kidney failure.
[0073] Infectious diseases can include bacterial infections, viral infections, rickettsial infections, chlamydial infections, fungal infections, protozoan infections, and parasitic infections.
[0074] There are no particular restrictions on the pathogens causing bacterial infections. Examples of pathogens include Escherichia coli, Staphylococcus, Streptococcus, Haemophilus spp., Neisseria spp., Moraxella spp., Listeria spp., Corynebacterium diphtheriae, Clostridium spp., Helicobacter pylori, and Mycobacterium tuberculosis.
[0075] There are no restrictions on the viruses that cause viral infections. Examples of pathogenic viruses include influenza virus, measles virus, rubella virus, varicella virus, dengue virus, cytomegalovirus, Epstein-Barr virus, enterovirus, human immunodeficiency virus, HTLV-1 (human T-lymphotropic virus type-I), and rabies virus.
[0076] There are no specific restrictions on the causative fungi of fungal infections. Pathogenic fungi can include yeast-like fungi, filamentous fungi, etc. Yeast-like fungi can include Cryptococcus and Candida species, etc. Filamentous fungi can include Aspergillus species, etc.
[0077] There are no specific restrictions on the causative protozoa of protozoan infections. Pathogenic protozoa may include Plasmodium, Leishmaniasis protozoa, etc.
[0078] Parasitic infections can be caused by parasites such as roundworms, nematodes, and hookworms.
[0079] As infectious diseases, bacterial infections, viral infections, protozoan infections, and parasitic infections can be preferentially listed, with bacterial infections being more preferred. Furthermore, infectious diseases may include pneumonia, sepsis, meningitis, and urinary tract infections.
[0080] Allergic diseases may include those belonging to types I, II, III, IV, or V. Type I allergic diseases may include hay fever, anaphylactic shock, allergic rhinitis, conjunctivitis, bronchial asthma, urticaria, atopic dermatitis, etc. Type II allergic diseases may include immune-incompatible blood transfusions, autoimmune hemolytic anemia, autoimmune thrombocytopenia, autoimmune granulocytopenia, Hashimoto's disease, pulmonary hemorrhage-nephritis syndrome, etc. Type III allergic diseases may include immune complex nephritis, Altus syndrome, serum sickness, etc. Type IV allergic diseases may include tuberculosis, contact dermatitis, etc. Type V allergic diseases may include Graves' disease, etc. Types I, II, III, and IV are preferred as allergic diseases; Types I, II, and III are more preferred; and Type I is even more preferred. Types II, III, and V allergic diseases partially overlap with the autoimmune diseases described later.
[0081] Autoimmune diseases can include systemic lupus erythematosus, rheumatoid arthritis, multiple sclerosis, Sjögren's syndrome, scleroderma, dermatomyositis, primary biliary cirrhosis, primary sclerosing cholangitis, ulcerative colitis, Crohn's disease, psoriasis, vitiligo vulgaris, bullous pemphigoid, alopecia areata, idiopathic dilated cardiomyopathy, type 1 diabetes mellitus, Graves' disease, Hashimoto's disease, myasthenia gravis, IgA nephropathy, membranous nephropathy, megaloblastic anemia, etc. Systemic lupus erythematosus, rheumatoid arthritis, multiple sclerosis, Sjögren's syndrome, scleroderma, and dermatomyositis are preferred autoimmune diseases. Autoimmune diseases for which antinuclear antibodies are detected are preferred.
[0082] Trauma can include fractures, burns, etc.
[0083] There are no restrictions on the type of sample that can be collected from the subject. Preferred samples include blood, bone marrow, urine, and body fluids. Examples of blood samples include peripheral blood, venous blood, and arterial blood. Peripheral blood is preferred. Examples of blood samples include peripheral blood collected using anticoagulants such as ethylenediaminetetraacetic acid (sodium or potassium salt) or heparin sodium. Body fluids are those other than blood and urine. Examples of such body fluids include ascites, pleural effusion, and bone marrow fluid.
[0084] The choice of sample depends on the disease being analyzed. In particular, in the aforementioned diseases, blood cells often exhibit characteristics different from normal cells in terms of the quantity distribution and / or abnormally observed cell types (described later). Therefore, blood samples can be used to analyze various diseases. Additionally, bone marrow can be used for the analysis, especially for diseases of the hematopoietic system. Cells contained in ascites, pleural effusion, and medullary fluid are effective for the diagnosis of tumors, hematopoietic system diseases, and infections, especially in tissues outside the hematopoietic system. Urine can be analyzed, especially for tumors and infections in tissues outside the hematopoietic system.
[0085] There are no restrictions on the cells to be analyzed, as long as they are cells contained in the sample. The cells to be analyzed are preferably cells used for disease analysis. The cells to be analyzed may contain multiple cells. "Multiple" here can refer to multiple quantities of a single cell type or multiple cell species. Normally, the sample may contain multiple cell species morphologically classified by histological and cytological microscopy. Cell morphological classification (also called "cell morphology classification") includes classification of cell species and classification of the types of cell abnormalities observed. Preferably, the cells to be analyzed are a group of cells belonging to a specified cell line. A specified cell line refers to a group of cells belonging to the same cell line differentiated from stem cells of a particular tissue. As a specified cell line, the hematopoietic system is preferred, and cells in the blood (also called "blood cells") are more preferred.
[0086] In conventional methods, hematopoietic cells are morphologically classified by observing specimens stained with bright-field microscopy under a microscope. The preferred staining agents are Reiter's stain, Giemsa stain, Swiss Giemsa stain, and Maggee stain. Maggee stain is preferred. There are no restrictions on the type of specimen, as long as the morphology of each cell belonging to a specified cell group can be observed individually. Examples include smears and imprints. Smears using peripheral blood or bone marrow as samples are preferred, with peripheral blood smears being more preferred.
[0087] In morphological classification, blood cells include the following cell types: neutrophils including segmented and band neutrophils, neutrophils, metamyelocytes, myelocytes, promyelocytes, granulocytes, lymphocytes, plasma cells, atypical lymphocytes, monocytes, eosinophils, basophils, erythroblasts (nucleated red blood cells, including proerythroblasts, basophils, polychromatic erythroblasts, orthochromatic erythroblasts, promegablasts, basophilic megablasts, polychromatic megablasts, and orthochromatic megablasts), platelets, platelet clots, and megakaryocytes (nucleated megakaryocytes, including micromegakaryocytes); etc.
[0088] In addition to normal cells, the aforementioned cell population may also include abnormal cells exhibiting morphological abnormalities. Abnormalities are characterized by morphologically classified cell features. Examples of abnormal cells are those present in the presence of the specified diseases, such as tumor cells. In the case of the hematopoietic system, the specified diseases are those selected from myelodysplastic syndromes, leukemias (including acute myeloid leukemia, acute granulocytic leukemia, acute promyelocytic leukemia, acute myelomonocytic leukemia (with neutrophil differentiation), acute monocytic leukemia, erythroleukemia, acute megakaryocytic leukemia, acute myeloid leukemia, acute lymphoblastic leukemia, lymphoblastic leukemia, chronic myeloid leukemia, and chronic lymphocytic leukemia), malignant lymphomas (Hodgkin's lymphoma and non-Hodgkin's lymphoma), and multiple myeloma. In addition, in the case of the hematopoietic system, abnormalities include cells with at least one morphological feature selected from abnormal nuclear morphology, presence of vacuoles, abnormal granule morphology, abnormal granule distribution, presence of abnormal granules, abnormal cell size, presence of inclusion bodies, and naked nuclei.
[0089] Abnormalities in nuclear morphology can be listed as follows: smaller nucleus; larger nucleus; more lobes in the nucleus; nuclei that should normally be lobed but are not lobed (including pseudo-Perger's nuclear abnormalities); presence of vacuoles; enlarged nucleoli; nuclei with incision marks; and, where a cell should normally have one nucleus but one cell abnormally has two nuclei, etc.
[0090] Abnormalities in the overall morphology of cells can include: vacuoles in the cytoplasm (also known as vacuolar degeneration); abnormal morphology of granules such as giant platelets, azurophilic granules, neutrophilic granules, eosinophilic granules, and basophilic granules; abnormal distribution of the above granules (excess, reduction, or disappearance); presence of abnormal granules (such as toxic granules); abnormal cell size (larger or smaller than normal); presence of inclusion bodies (Dürer bodies, Orl bodies, etc.); and naked nuclei, etc.
[0091] Preferably, the above-mentioned abnormalities are selected from at least one of the following: abnormal nuclear morphology, abnormal granules, abnormal cell size, abnormal cell shape, cell damage, vacuoles, immature cells, presence of inclusion bodies, Dürer bodies, satellite phenomena, abnormal nuclear network, petal-like nuclei, large N / C ratio, bleb-like, ulcerated cells, and hair-like cell morphology.
[0092] Preferably, the aforementioned nuclear morphological abnormalities include at least one selected from the following: highly lobed, poorly lobed, pseudo-Perger's nuclear abnormalities, ring-shaped nuclei, spherical nuclei, elliptical nuclei, apoptosis, fused nuclei, nuclear rupture, enucleation, naked nuclei, irregular nuclear margins, nuclear fragmentation, intranuclear bridging, multinucleation, split nuclei, nuclear division, and nucleolar abnormalities. The aforementioned granular abnormalities include at least one selected from the following: degranulation, abnormal granule distribution, toxic granules, Orl bodies, bundled cells, and pseudo-Chediak-Higashi granular granules. Granular abnormalities in eosinophils and basophils include phenomena such as granule distribution bias within the cell, which are considered abnormal granules. The aforementioned cell size abnormalities include giant platelets.
[0093] Preferably, the cell types include at least one selected from neutrophils, eosinophils, platelets, lymphocytes, monocytes, and basophils.
[0094] More preferably, the aforementioned cell types further include at least one selected from metamyelocytes, myelocytes, promyelocytes, granulocytes, plasma cells, atypical lymphocytes, immature eosinophils, immature basophils, erythroblasts, and megakaryocytes.
[0095] More preferably, the hematopoietic system disease is aplastic anemia or myelodysplastic syndrome. When the cell type is neutrophils, the abnormal findings are selected from at least one of granular abnormalities and polylobulation; when the cell type is eosinophils, the abnormal findings are abnormal granules. Furthermore, the abnormal cellular findings include giant platelets. By evaluating these findings, aplastic anemia and myelodysplastic syndrome can be identified.
[0096] <Summary of Auxiliary Methods>
[0097] In the auxiliary methods, there are no restrictions on identifying abnormal findings and / or cell types, as long as the abnormal findings and / or cell types can be identified from the image. Identification can be performed by the examiner or using the identifier described below.
[0098] use Figure 1This section outlines an auxiliary method using a recognizer. The recognizer used in the auxiliary method comprises a computer algorithm. Preferably, the computer algorithm comprises a first computer algorithm and a second computer algorithm. More preferably, the first computer algorithm comprises multiple deep learning algorithms with a neural network structure. The second computer algorithm comprises a machine learning algorithm. Preferably, the deep learning algorithm comprises: a first neural network 50 for extracting feature quantities that quantitatively represent the morphological characteristics of cells, and a second neural network 51 for identifying the types of abnormalities seen in cells and / or a second neural network 52 for identifying cell types. The first neural network 50 extracts feature quantities of cells, and the second neural networks 51 and 52 are downstream of the first neural network, identifying abnormalities or cell types based on the feature quantities extracted by the first neural network 50. More preferably, the second neural networks 51 and 52 may comprise a neural network trained for identifying cell types, and multiple neural networks corresponding to each abnormality seen in the cell, trained on each abnormality seen in the cell. For example, in Figure 1 In this algorithm, the first deep learning algorithm is used to identify a first anomalous appearance (e.g., particle anomalous appearance), and includes a first neural network 50 and a second neural network 51 trained for identifying the first anomalous appearance. The first' deep learning algorithm is used to detect a second anomalous appearance (e.g., multi-lobed appearance), and includes a first neural network 50 and a second neural network 51 trained for identifying the second anomalous appearance. The second deep learning algorithm is used to identify cell types, and includes a first neural network 50 and a second neural network 52 trained for identifying cell types.
[0099] For each sample, the machine learning algorithm analyzes the diseases present in the subject who collected the sample based on the feature quantities output by the deep learning algorithm, and outputs the disease name or a label representing the disease name as the analysis result.
[0100] Then, use Figures 2-4 The examples shown illustrate training data 75 for deep learning, methods for generating training data for machine learning, and methods for analyzing diseases. For ease of explanation, the following uses a first neural network, a second neural network, and a gradient boosting tree as a machine learning algorithm.
[0101] <Generation of Training Data for Deep Learning>
[0102] The training images 70 used to train the deep learning algorithm are images of cells from the subject of analysis, which are contained in a specimen collected from the patient and labeled with a disease name. Multiple training images 70 are taken from one specimen. The cells from the subject of analysis contained in each image are associated with the cell types and abnormalities identified by the examiner based on morphological classification. Preferably, the specimens used for taking the training images 70 are prepared from specimens containing the same type of cells as the cells from the subject of analysis, and using the same specimen preparation and staining methods as those used for specimens containing the cells from the subject of analysis. Furthermore, it is preferable to take the training images 70 under the same conditions as those used for taking the specimens from the subject of analysis.
[0103] Training images 70 for each cell can be pre-acquired using imaging devices such as well-known optical microscopes or virtual slide scanners. Figure 2 In the example shown, the original image captured at 360 pixels × 365 pixels using the DI-60 blood image automatic analysis device (manufactured by Sysmex Corporation) is reduced to 255 pixels × 255 pixels to generate training image 70, but this reduction is not necessary. There is no limit to the number of pixels in training image 70, as long as it can be analyzed; preferably, one side should exceed 100 pixels. Furthermore, in Figure 2 In the example shown, a neutrophil is at the center, surrounded by red blood cells. The image can also be cropped to retain only the target cell. As long as at least one image contains one cell to be trained (which can include red blood cells and normal-sized platelets), and the pixels corresponding to the cell to be trained occupy more than 1 / 9 of the total pixels of the image, it can be used as a training image 70.
[0104] As an example, it is preferable to use an imaging device that captures images using RGB color and CMY color. In the color image, it is preferable to represent the shades or brightness of each primary color, such as red, green, and blue, or cyan, magenta, and yellow, using 24-bit values (8 bits × 3 colors). The training image 70 may contain at least one hue and the shade or brightness of that hue, and more preferably, it may contain at least two hues and the shade or brightness of each hue. The information containing the hue and the shade or brightness of that hue is also called hue.
[0105] Then, the hue information of each pixel is converted from RGB color to a format that includes both luminance and hue information. Examples of formats that include luminance and hue information include YUV (YCbCr, YPbPr, YIQ, etc.). Here, we will use the conversion to YCbCr format as an example. Since the training image is RGB color, it is converted to luminance 72Y, first hue (e.g., blue) 72Cb, and second hue (e.g., red) 72Cr. This conversion from RGB to YCbCr can be done using well-known methods. For example, it can be done according to the international standard ITU-R BT.601. Figure 2 As shown, the converted luminance 72Y, hue 72Cb, and hue 72Cr can be represented by rows and columns of level values (hereinafter also referred to as hue rows and columns 72y, 72cb, and 72cr). Luminance 72Y, hue 72Cb, and hue 72Cr are represented by 256 levels from 0 to 255. Alternatively, the three primary colors of red (R), green (G), and blue (B) or the three primary colors of cyan (C), magenta (M), and yellow (Y) can be used to convert the training image, replacing luminance 72Y, hue 72Cb, and hue 72Cr.
[0106] Then, based on the hue rows 72y, 72cb, and 72cr, a hue vector data 74 is generated for each pixel, consisting of three levels of values: a combined luminance of 72y, a first hue of 72cb, and a second hue of 72cr.
[0107] Then, for example Figure 2 The training image 70 represents segmented neutrophils; therefore, the tonal vector data 74 generated from the training image 70 in Figure 2 is assigned a "1" as a label value 77 representing segmented neutrophils, becoming the training data 75 for deep learning. Figure 2 In this example, for convenience, the training data 75 for deep learning is represented by 3 pixels × 3 pixels, but in reality, there exists tone vector data corresponding to the number of pixels when the training image 70 was captured.
[0108] Figure 3 An example of label value 77 is shown. Different label values 77 are assigned based on cell type and the presence or absence of abnormalities observed in each cell.
[0109] <Summary of Recognizer Generation>
[0110] by Figure 2 Let's take an example to illustrate the general outline of the recognizer generation method. The generation of the recognizer may include the training of deep learning algorithms and machine learning algorithms.
[0111] Training of deep learning algorithms
[0112] The first deep learning algorithm comprises a first neural network 50 and a second neural network 51, used to generate first information 53, which is information related to the types of anomalies observed. The second deep learning algorithm comprises a first neural network 50 and a second neural network 52, used to generate second information 54, which is information related to cell types.
[0113] The number of nodes in the input layer 50a of the first neural network 50 corresponds to the product of the number of pixels in the input deep learning training data 75 and the number of luminance and hue in the image (e.g., luminance 72y, hue 72cb, and hue 72cr in the example above). Hue vector data 74, as its set 76, is input to the input layer 50a of the first neural network 50. The label values 77 of each pixel in the deep learning training data 75 are input to the output layer 50b of the first neural network to train the first neural network 50.
[0114] The first neural network 50 extracts feature quantities based on the deep learning training data 75 for the aforementioned morphological cell types and cell features reflecting anomalies. The output layer 50b of the first neural network outputs results reflecting these feature quantities. The results output by the flexible maximum function of the output layer 50b of the first neural network 50 are input to the input layer 51a of the second neural network 51 and the input layer 52a of the second neural network 52. In addition, since cells belonging to a specified cell line have similar cell morphologies, the second neural networks 51 and 52 are trained specifically to further identify morphologically specific cell types and cell features reflecting specific anomalies. Therefore, the label values 77 of the deep learning training data 75 are also input to the output layers 51b and 52b of the second neural networks. Figure 2 The symbols 50c, 51c, and 52c represent intermediate layers. Additionally, one second neural network 51 can be trained for each anomaly observed. In other words, a second neural network 51 can be trained corresponding to the number of anomaly types to be analyzed. The second neural network 52 used to identify cell types is of one type.
[0115] Preferably, the first neural network 50 is a convolutional neural network, and the second neural networks 51 and 52 are fully connected neural networks.
[0116] Thus, a first deep learning algorithm with a trained first neural network 60 and a second neural network 61, and a second deep learning algorithm with a trained first neural network 61 and a second neural network 62 are generated (see reference). Figure 5 ).
[0117] For example, the second neural network 61, used to identify anomalies, outputs the probability of whether or not an anomaly is present, as the identification result. This probability can be an anomaly name or a label value corresponding to the anomaly. The second neural network 62, used to identify cell types, outputs the probability that each analyzed cell belongs to any of the various cell types input as training data, as the identification result. This probability can be a cell type name or a label value corresponding to the cell type.
[0118] Training machine learning algorithms
[0119] As training data for training machine learning algorithm 57, using Figure 4 The machine learning training data 90 shown contains feature quantities and disease information 55. For each sample, the probability of anomalies and / or cell types output by the deep learning algorithm, or the value obtained by converting the above probabilities into cell counts, can be used as the feature quantity of the learning object. In the machine learning training data 90, the feature quantities are associated with disease information such as disease names and disease label values held by the subjects of each sample.
[0120] The features input to the machine learning algorithm 57 are at least one of information related to the type of anomaly and information related to the cell type. Preferably, information related to the type of anomaly and information related to the cell type are used as features. The anomaly type used as a feature can be one or more. Similarly, the cell type used as a feature can be one or more.
[0121] The trained first deep learning algorithm and / or second deep learning algorithm are used to analyze the training images 70 taken from each sample for training the deep learning algorithms, identifying abnormalities and / or cell types in each training image 70. The second neural network 61 outputs the probability of each abnormality and a label value representing the abnormality for each cell. The probability of each abnormality and the label value representing the abnormality constitute the identification result of the abnormality type. The second neural network 62 outputs the probability corresponding to each cell type and the label value representing the cell type. The probability corresponding to each cell type and the label value representing the cell type constitute the cell type identification result. Based on this information, feature quantities are generated to input the machine learning algorithm 57.
[0122] The example of machine learning using 90 training data is shown below. Figure 4 .exist Figure 4For ease of explanation, three cells (cells No. 1 to 3) are shown, representing five abnormalities: degranulation, Orr bodies, globular nuclei, multilobulated nuclei, and giant platelets. Eight cell types are also shown: segmented neutrophils, band neutrophils, lymphocytes, monocytes, eosinophils, basophils, granulocytes, and platelets. Figure 4 In the table, the A to F markers at the top indicate the column numbers. The 1 to 19 markers on the left indicate the row numbers.
[0123] For each sample, the first neural network 50 and the second neural network 51 calculate the probability of each abnormality observed for each analyte cell, and the second neural network 51 outputs the calculated probability. Figure 4 In this context, the probability of retaining an anomaly is represented by a number between 0 and 1. For example, retaining an anomaly can be represented by a value close to "1", while not retaining an anomaly can be represented by a value close to "0". Figure 4 In the diagram, the values recorded in cells A through E and rows 1 through 5 are the output values of the second neural network 51. Then, for each sample, the sum of probabilities of the types of anomalies observed is calculated. Figure 4 In the table, the values recorded in cells F and rows 1 to 5 are the sum of all anomalies observed. For a single sample, the data set obtained by associating the label representing the name of an anomaly with the sum of the probabilities of each anomaly observed in a single sample is called "anomaly type related information". Figure 2 In the information provided, the type of anomaly observed is listed as information number 1, 53. Figure 4 In this context, the data group obtained by establishing associations between cells B1 and F1, B2 and F2, B3 and F3, B4 and F4, and B5 and F5 is designated as "Abnormality Type Related Information," becoming Information 53 (Section 1). Associating Information 53 with disease names or disease information 55 representing disease name label values creates training data 90 for machine learning. Figure 4 In the middle, line 19 indicates disease information 55.
[0124] In addition, such as Figure 4 As shown, sometimes the probability of a single anomaly being displayed may not reach "1". In such cases, for example, a specified cutoff value can be set, and the probability of displaying anomaly types with values below the cutoff value will be considered as "0". Alternatively, a specified cutoff value can be set, and the probability of displaying anomaly types with values above the cutoff value will be considered as "1".
[0125] Here, the probability of the type of each anomaly can also be expressed in terms of the number of cells of each type of anomaly.
[0126] Furthermore, for each cell type, the first neural network 50 and the second neural network 52 calculate the probability corresponding to each cell type for each analysis target cell, and the second neural network 52 outputs the calculated probability. The first neural network 50 and the second neural network 52 calculate the probability corresponding to each cell type for all cell types being analyzed. Figure 4 In the example, for a single cell sample, the probability corresponding to each cell type is calculated for all categories: segmented neutrophils, band neutrophils, lymphocytes, monocytes, eosinophils, basophils, granulocytes, and platelets. The values recorded in cells A through E and rows 6 through 13 are the output values of the second neural network 52. Then, for each sample, the sum of the probabilities for each cell type is calculated. Figure 4 In the table, the values recorded in cells 6 to 13 of column F are the sum of all cell types. The data set obtained by associating the labels representing cell types with the sum of the probabilities of a sample corresponding to each cell type is called "cell type related information". Figure 2 In the text, information related to cell type is item 2, number 54. Figure 4 In this process, the data group obtained by establishing associations between cells B6 and F6, B7 and F7, B8 and F8, B9 and F9, B10 and F10, B11 and F11, B12 and F12, and B13 and F13 is designated as "cell type related information," becoming the second piece of information 54. This second piece of information 54 is then associated with disease information 55, which consists of disease names or label values representing disease names, becoming training data 90 for machine learning. Figure 4 In the middle, line 19 indicates disease information 55.
[0127] Here, the probability of each cell type can be represented by the number of cells of each cell type. Additionally, as... Figure 4 As shown, for items involving one analysis cell and multiple cell types, the probability of a value higher than 0 is displayed. In this case, for example, a predetermined cutoff value can be set, and the probability of cell types displaying values below the cutoff value can be considered as "0". Alternatively, a predetermined cutoff value can be set, and the probability of cell types displaying values above the cutoff value can be considered as "1".
[0128] Furthermore, the preferred feature is relevant information about the types of abnormalities observed for each cell type. If using... Figure 4To illustrate, when generating deep learning training data 75 as shown in rows 14 to 18, such as neutrophil degranulation (cell B14), granulocyte eustomas (cell B15), neutrophil globules (cell B16), neutrophil segmentation (cell B17), and giant platelets (cell B18), training is performed by associating specific cell types with specific types of anomalies. The generation of features is the same as when using anomaly types not associated with cell types. The relevant information about anomalies obtained for each cell type is called third information. This third information is associated with disease information 55, either the disease name or a label value representing the disease name, to become machine learning training data 90.
[0129] The training data 90 is input into the machine learning algorithm 57 to train the machine learning algorithm 57, generating the trained machine learning algorithm 67 (see reference). Figure 5 ).
[0130] As a training method for the machine learning algorithm 57, it is preferable to use at least one of the following: machine learning training data 90 obtained by establishing a correlation between the first information and disease information 55, machine learning training data 90 obtained by establishing a correlation between the second information and disease information 55, and machine learning training data 90 obtained by establishing a correlation between the third information and disease information 55. More preferably, the machine learning algorithm 57 is trained using machine learning training data 90 obtained by establishing a correlation between the first information and disease information 55 and machine learning training data 90 obtained by establishing a correlation between the second information and disease information 55, or machine learning training data 90 obtained by establishing a correlation between the third information and disease information 55 and machine learning training data 90 obtained by establishing a correlation between the second information and disease information 55. The optimal method is to use both the machine learning training data 90 obtained by associating the second information 54 with the disease information 55 (or the label value representing the disease name) and the machine learning training data 90 obtained by associating the third information 54 with the disease information 55 (or the label value representing the disease name) as training data and input them into the machine learning algorithm 57. In this case, the cell type in the second information 54 can be the same as or different from the cell type associated with the third information.
[0131] There are no restrictions on machine learning algorithms as long as they can analyze diseases based on the aforementioned features. For example, they can be chosen from regression, trees, neural networks, Bayesian methods, time series models, clustering, and ensemble learning.
[0132] Regression models can include linear regression, logistic regression, and support vector machines. Tree models can include gradient boosting trees, decision trees, regression trees, and random forests. Neural networks can include perceptrons, convolutional neural networks, recurrent neural networks, and residual networks. Time series models can include moving averages, autoregressive models, autoregressive moving averages, and autoregressive integral moving averages. Clustering can include k-nearest neighbors. Ensemble learning can include boosting learning and bundled learning. Gradient boosting trees are preferred.
[0133] <Auxiliary Methods for Disease Analysis>
[0134] Figure 5 An example of an auxiliary method for disease analysis is shown. In this auxiliary method, analytical data 81 is generated from an analytical image 78 obtained by photographing the cells of the target organism. The analytical image 78 is an image obtained by photographing the cells of the target organism contained in a sample collected from a subject. For example, the analytical image 78 can be obtained using a known imaging device such as an optical microscope or a virtual slide scanner. Figure 5 In the example shown, the original image, captured using the same 360 pixel × 365 pixel format as the training image 70 using the DI-60 automatic blood image analysis device (manufactured by Sysmex Corporation), is reduced to 255 pixels × 255 pixels to generate the analysis image 78. However, this reduction is not mandatory. The number of pixels in the analysis image 78 is not limited as long as analysis is possible; preferably, one side has more than 100 pixels. Furthermore, in... Figure 5 In the example shown, the image centers on a segmented neutrophil surrounded by red blood cells. Alternatively, the image can be cropped to retain only the target cell. An image 78 can be used for analysis as long as it contains at least one cell to be analyzed (which can include red blood cells and normal-sized platelets) and the pixels corresponding to the cell to be analyzed occupy approximately 1 / 9 or more of the total image pixels.
[0135] As an example, it is preferable to use an imaging device that captures images using RGB color and CMY color. In the color image, it is preferable to represent the shades or brightness of each primary color, such as red, green, and blue, or cyan, magenta, and yellow, using 24-bit values (8 bits × 3 colors). The image 78 for analysis only needs to contain at least one hue and the shade or brightness of that hue, and more preferably, it needs to contain at least two hues and the shade or brightness of each hue. The information containing the hue and the shade or brightness of that hue is also called hue.
[0136] For example, converting from RGB color to a format that includes both luminance and hue information. Examples of formats that include luminance and hue information include YUV (YCbCr, YPbPr, YIQ, etc.). Here, we will use the conversion to YCbCr format as an example. Since the image used for analysis is RGB color, it is converted to luminance 79Y, primary hue (e.g., blue) 79Cb, and secondary hue (e.g., red) 79Cr. Well-known methods can be used to convert from RGB to YCbCr. For example, it can be done according to the international standard ITU-R BT.601. Figure 5 As shown, the converted luminance 79Y, hue 79Cb, and hue 79Cr can be represented by rows and columns of level values (hereinafter also referred to as hue rows and columns 79y, 79cb, and 79cr). Luminance 79Y, hue 79Cb, and hue 79Cr are represented by 256 levels from 0 to 255. Alternatively, the three primary colors of red (R), green (G), and blue (B) or the three primary colors of cyan (C), magenta (M), and yellow (Y) can be used to convert and analyze the image, replacing luminance 79Y, hue 79Cb, and hue 79Cr.
[0137] Then, based on the hue rows 79y, 79cb, and 79cr, a hue vector data 80 is generated for each pixel, consisting of three levels of values: a combined luminance of 79y, a first hue of 79cb, and a second hue of 79cr. A set of hue vector data 80 generated from one analysis image 78 is generated as analysis data 81.
[0138] For the generation of analysis data 81 and the generation of training data 75 for deep learning, it is preferable that at least the shooting conditions and the generation conditions of the vector data input into the neural network from each image are the same.
[0139] The first deep learning algorithm comprises a first neural network 60 and a second neural network 62, which are used to generate first information 63, which is information related to the types of anomalies observed. The second deep learning algorithm comprises a first neural network 60 and a second neural network 62, which are used to generate second information 64, which is information related to cell types.
[0140] The analysis data 81 is input into the input layer 60a of the trained first neural network 60. The first neural network 60 extracts cell features from the analysis data 81 and outputs the results from the output layer 60b of the first neural network 60. The results output by the flexible maximum function of the output layer 60b of the first neural network 60 are input into the input layer 61a of the second neural network 61 and the input layer 62a of the second neural network 62.
[0141] Then, the result output by the output layer 60b is input into the input layer 61a of the trained second neural network 61. For example, the second neural network 61, used to identify anomalies, outputs the probability of whether or not an anomaly is present based on the input feature quantity, as the anomaly identification result.
[0142] Furthermore, the results output by output layer 60b are input into input layer 62a of the trained second neural network 62. Based on the input features, the second neural network 62 outputs from output layer 62b the probability that the cells in the image being analyzed belong to each of the cell types input as training data. Figure 5 In the text, the symbols 60c, 61c, and 62c represent intermediate layers.
[0143] Then, based on the identification results of the anomalies, information related to the type of anomalies observed in each sample is obtained. Figure 5 The first information 63 is, for example, the sum of probabilities of the types of each anomaly observed, output by the output layer 61b of the second neural network 61 and obtained through analysis. The method for generating the first information 63 is the same as the method for generating training data in machine learning.
[0144] In addition, based on the cell species identification results, relevant information on the cell species corresponding to each sample is obtained. Figure 5 (Second information 64 in the second information). Second information 64 is, for example, the sum of probabilities for each cell type, output by the output layer 62b of the second neural network 62 and obtained through analysis. The method for generating second information 64 is the same as the method for generating training data 90 for machine learning.
[0145] The generated first information 63 and second information 64 are input into the trained machine learning algorithm 67, which then generates an analysis result 83. The analysis result 83 can be a disease name or a label value representing a disease name.
[0146] As input data for the machine learning algorithm 67, at least one of the following can be used: first information 63, second information 64, and third information. More preferably, first information 63 and second information 64, or third information and second information 64, can be used. Most preferably, both second information 64 and third information can be used as the method for analyzing data 81. In this case, the cell types in second information 64 can be the same as or different from the cell types associated with third information. The third information is information generated when generating analysis data 81 to associate specific cell types with specific types of anomalies, and its generation method is the same as that described in the method for generating training data 90 for machine learning.
[0147] [Disease Analysis Support System 1]
[0148] <The Composition of Disease Analysis Support System 1>
[0149] The auxiliary system 1 for disease analysis is described below. (Refer to...) Figure 6 The disease analysis assistance system 1 includes a training device 100A and a disease analysis device 200A. The supplier-side device 100 operates as the training device 100A, and the user-side device 200 operates as the disease analysis device 200A. The training device 100A uses deep learning training data 75 and machine learning training data 90 to generate a recognizer and provides it to the user. The recognizer is provided from the training device 100A to the disease analysis device 200A via a recording medium 98 or a network 99. The disease analysis device 200A uses the recognizer provided by the training device 100A to perform image analysis of the target cells.
[0150] The training device 100A, for example, is composed of a general-purpose computer and performs deep learning processing based on the flowchart described later. The disease analysis device 200A, for example, is composed of a general-purpose computer and performs disease analysis processing based on the flowchart described later. The recording medium 98 is, for example, a non-transitory tangible recording medium that can be read by a computer, such as a DVD-ROM or a USB memory.
[0151] The training device 100A is connected to the imaging device 300. The imaging device 300 includes an imaging element 301 and a fluorescence microscope 302, and captures bright-field images of the training specimen 308 placed on the stage 309. The training specimen 308 has undergone the aforementioned staining. The training device 100A acquires the training image 70 captured by the imaging device 300.
[0152] The disease analysis device 200A is connected to the imaging device 400. The imaging device 400 includes an imaging element 401 and a fluorescence microscope 402, and captures a bright-field image of the analyte, i.e., specimen 408, placed on the stage 409. The analyte, i.e., specimen 408, has been pre-stained as described above. The disease analysis device 200A acquires an image 78 of the analyte captured by the imaging device 400.
[0153] The imaging devices 300 and 400 can use known optical microscopes or virtual slide scanners with specimen imaging capabilities.
[0154] <Hardware Configuration of the Training Device>
[0155] Reference Figure 7 The supplier-side device 100 (training device 100A) includes a processing unit 10 (10A), an input unit 16, and an output unit 17.
[0156] The processing unit 10 includes a CPU (Central Processing Unit) 11 for data processing (described later), a memory 12 for a working area for data processing, a storage unit 13 for recording programs and processed data (described later), a bus 14 for transmitting data between the units, an interface unit 15 for data input / output with external devices, and a GPU (Graphics Processing Unit) 19. An input unit 16 and an output unit 17 are connected to the processing unit 10. For example, the input unit 16 may be an input device such as a keyboard or mouse, and the output unit 17 may be a display device such as a liquid crystal display. The GPU 19 functions as an accelerator to assist the computational processing (e.g., parallel computation processing) performed by the CPU 11. That is, in the following description, the processing performed by the CPU 11 also includes processing performed by the CPU 11 using the GPU 19 as an accelerator.
[0157] In addition, in order to perform the following Figure 10 , Figure 13 and Figure 14 In the processing of each step described herein, the processing unit 10 pre-records the program and recognizer of the present invention in the storage unit 13 in, for example, an executable format. The executable format is, for example, a format generated by a compiler from a programming language. The processing unit 10 uses the program recorded in the storage unit 13 to perform training processing on the first neural network 50, the second neural network 51, the second neural network 52, and the machine learning algorithm 57.
[0158] In the following description, unless otherwise stated, the processing performed by the processing unit 10 refers to the processing performed by the CPU 11 based on the program stored in the storage unit 13 or the memory 12, as well as the first neural network 50, the second neural network 51, the second neural network 52, and the machine learning algorithm 57. The CPU 11 uses the memory 12 as its working area to erasurely and temporarily store necessary data (intermediate data during processing, etc.), and non-erasably stores data that needs to be stored for a long time, such as calculation results, in the storage unit 13.
[0159] <Hardware Composition of the Disease Analysis Device>
[0160] Reference Figure 8 The user-side device 200 (disease analysis device 200A, disease analysis device 200B, disease analysis device 200C) includes a processing unit 20 (20A, 20B, 20C), an input unit 26, and an output unit 27.
[0161] The processing unit 20 includes a CPU (Central Processing Unit) 21 for data processing (described later), a memory 22 for a working area for data processing, a storage unit 23 for storing programs and processing data (described later), a bus 24 for transmitting data between the units, an interface unit 25 for data input / output with external devices, and a GPU (Graphics Processing Unit) 29. An input unit 26 and an output unit 27 are connected to the processing unit 20. For example, the input unit 26 may be an input device such as a keyboard or mouse, and the output unit 27 may be a display device such as a liquid crystal display. The GPU 29 functions as an accelerator to assist the computational processing (e.g., parallel computation processing) performed by the CPU 21. That is, in the following description, the processing performed by the CPU 21 also includes processing performed by the CPU 21 using the GPU 29 as an accelerator.
[0162] Furthermore, in order to perform the steps described in the following disease analysis process, the processing unit 20 pre-records the program and recognizer of the present invention in the storage unit 23, for example, in an executable format. The executable format is, for example, a format generated by a compiler from a programming language. The processing unit 20 performs processing using the program and recognizer recorded in the storage unit 23.
[0163] In the following description, unless otherwise stated, the processing performed by the processing unit 20 refers to the actual processing performed by the CPU 21 of the processing unit 20 based on the program and deep learning algorithm 60 stored in the storage unit 23 or memory 22. The CPU 21 uses the memory 22 as its working area to erasurely and temporarily store necessary data (intermediate data during processing, etc.), and non-erasably records data that needs to be stored for a long time, such as calculation results, in the storage unit 23.
[0164] <Functional Modules and Processing Steps>
[0165] (Deep learning processing)
[0166] Reference Figure 9 The processing unit 10A of the training device 100A includes a deep learning training data generation unit 101, a deep learning training data input unit 102, and a deep learning algorithm update unit 103. These functional modules are implemented by installing a program that enables a computer to perform deep learning processing in the storage unit 13 or memory 12 of the processing unit 10A, and by having the CPU 11 execute the program. A deep learning training data database (DB) 104 and a deep learning algorithm database (DB) 105 are recorded in the storage unit 13 or memory 12 of the processing unit 10A.
[0167] The training images 70 are pre-captured using the imaging device 300, and are correlated with, for example, the morphological cell type and abnormalities observed in the target cells, and are pre-stored in the storage unit 13 or memory 12 of the processing unit 10A. The deep learning algorithm database 105 pre-stores untrained first neural network 50, second neural network 52, and second neural network 52. The deep learning algorithm database 105 also pre-stores first neural network 50, second neural network 52, and second neural network 52 that have undergone one training iteration and have been updated.
[0168] The processing unit 10A of the training device 100A performs... Figure 10 The process shown. Using Figure 9 The functional modules shown are explained below. Steps S11, S12, S16, and S17 are processed by the deep learning training data generation unit 101. Step S13 is processed by the deep learning training data input unit 102. Step S14 is processed by the deep learning algorithm update unit 103.
[0169] use Figure 10 An example of deep learning processing performed by the processing unit 10A will be described. First, the processing unit 10A acquires a training image 70. The training image 70 is acquired by the imaging device 300 through operator operation, or by the recording medium 98, or via the network using the I / F unit 15. When acquiring the training image 70, information is also acquired regarding which cell type and / or abnormality is represented by the training image 70 according to morphological classification. This information regarding which cell type and / or abnormality is represented by the training image 70 can be associated with the training image 70, or it can be input by the operator from the input unit 16.
[0170] In step S11, the processing unit 10A converts the acquired training image 70 into luminance Y, first hue Cb and second hue Cr, and generates hue vector data 74 according to the steps described in the above training data generation method.
[0171] In step S12, the processing unit 10A assigns a label value corresponding to the tone vector data 74 based on information associated with the training image 70 indicating which of the cell types and / or abnormalities are classified morphologically, and the label values stored in the memory 12 or storage unit 13 associated with the cell types or abnormalities classified morphologically. Thus, the processing unit 10A generates deep learning training data 75.
[0172] exist Figure 10In step S13 shown, the processing unit 10A trains the first neural network 50 and the second neural network 51 using deep learning training data 75. Each time training is performed using multiple deep learning training data 75, the training results of the first neural network 50 and the second neural network 51 are accumulated.
[0173] Then, in step S14, the processing unit 10A determines whether training results for a predetermined number of attempts have been accumulated. When training results for the predetermined number of attempts have been accumulated (YES), the processing unit 10A proceeds to step S15; when training results for the predetermined number of attempts have not been accumulated (NO), the processing unit 10A proceeds to step S16.
[0174] Then, when the training results of a predetermined number of trials have been accumulated, in step S15, the processing unit 10A updates the connection weights w of the first neural network 50 and the second neural network 51, or the first neural network 50 and the second neural network 52, using the training results accumulated in step S13. Since the disease analysis method uses stochastic gradient descent, the connection weights w of the first neural network 50 and the second neural network 51, or the first neural network 50 and the second neural network 52, are updated at the stage where the learning results of a predetermined number of trials have been accumulated. The process of updating the connection weights w specifically involves performing the gradient descent operation shown in equations (11) and (12) described later.
[0175] In step S16, the processing unit 10A determines whether the first neural network 50 and the second neural network 51, or the first neural network 50 and the second neural network 52, have been trained using a predetermined amount of training data 75. When the predetermined amount of training data 75 has been used (YES), the deep learning process ends.
[0176] When the first neural network 50 and the second neural network 51, or the first neural network 50 and the second neural network 52, are not trained with the specified amount of training data 75 (NO), the processing unit 10A proceeds from step S16 to step S17, and performs the processing from step S11 to step S16 for the next training image 70.
[0177] Train the first neural network 50 and the second neural network 51, or the first neural network 50 and the second neural network 52, according to the above description, to obtain the second deep learning algorithm and the second deep learning algorithm.
[0178] (Structure of a neural network)
[0179] Figure 11(a) The structure of the first neural network 50 and the second neural networks 51 and 52 is illustrated. The first neural network 50 and the second neural network 51 have: input layers 50a, 51a, and 52a; output layers 50b, 51b, and 52b; and intermediate layers 50c, 51c, and 52c located between the output layers 50b, 51b, and 52b. The intermediate layers 50c, 51c, and 52c are composed of multiple layers. The number of layers constituting the intermediate layers 50c, 51c, and 52c can be, for example, five or more.
[0180] In the first neural network 50 and the second neural network 51, or the first neural network 50 and the second neural network 52, multiple nodes 89 configured in a layered manner are connected between layers. Thus, information propagates from the input-side layers 50a, 51a, 52a only along the unidirectional path indicated by arrow D in the figure to the output-side layers 50b, 51b, 52b.
[0181] (Calculations for each node)
[0182] Figure 11 (b) is a schematic diagram illustrating the operations at each node. Each node 89 receives multiple inputs and calculates one output (z). Figure 11 In the example shown in (b), node 89 receives 4 inputs. The total input (u) received by node 89 is represented by the following (Equation 1).
[0183]
Number 1
[0184] u=w1x1+w2x2+w3x3+w4x4+b (Formula 1)
[0185] Each input is multiplied by a different weight value. In Equation 1, b is the value called the bias. The output (z) of the node is the output of the defined function f relative to the total input (u) represented by Equation 1, and is expressed as follows (Equation 2). The function f is called the activation function.
[0186]
Number 2
[0187] z = f(u) (Equation 2)
[0188] Figure 11 (c) is a schematic diagram illustrating the operations between nodes. In the first neural network 50 and the second neural networks 51 and 52, the nodes that output the result (z) shown in Equation 2 for the total input (u) shown in Equation 1 are arranged in a layered manner. The output of the node in the previous layer becomes the input of the node in the next layer. Figure 11In the example shown in (c), the output of node 89a in the left layer of the diagram is the input of node 89b in the right layer of the diagram. Each node 89b in the right layer receives the output from node 89a in the left layer. The connections between each node 89a in the left layer and each node 89b in the right layer are multiplied by different weight values. If the outputs of the multiple nodes 89a in the left layer are x1 to x4 respectively, then the inputs of the three nodes 89b in the right layer are as shown in equations (3-1) to (3-3).
[0189]
Number 3
[0190] u1 = w 11 x1+w 12 x2+w 13 x3+w 14 x4+b1
[0191] u2 = w 21 x1+w 22 x2+w 23 x3+w 24 x4+b2
[0192] u3=w 31 x1+w 32 x2+w 33 x3+w 34 x4+b3
[0193] The above equations (3-1) to (3-3) can be summarized as equation (3-4). Here, i = 1, ... I, j = 1, ... J.
[0194]
Number 4
[0195]
[0196] Applying equation (3-4) to the activation function yields the output. The output is shown in equation (4) below.
[0197]
Number 5
[0198] z j =f(u j (j=1,2,3) (Equation 4)
[0199] (Activation function)
[0200] In disease analysis methods, a rectified linear unit function is used as the activation function. The rectified linear unit function is shown in Equation 5 below.
[0201]
Number 6
[0202] f(u) = max(u, 0) (Equation 5)
[0203] Equation 5 represents the part of a linear function z = u where u < 0 is the part where u = 0. Figure 9 In the example shown, according to (Equation 5), the output of the node j=1 is as follows.
[0204]
Number 7
[0205] z1=max((w 11 x1+w 12 x2+w 13 x3+w 14 x4+b1), 0)
[0206] (Learning Neural Networks)
[0207] Let y(x:w) be the function represented using a neural network. Changing the parameters w of the neural network will change the function y(x:w). Adjusting the function y(x:w) so that the neural network selects more suitable parameters w for the input x is called the learning of the neural network. Provide an array of inputs and outputs of the function represented using the neural network. If the preferred output for a given input x is d, then provide {(x1, d1), (x2, d2), ..., (x...} n d n The set of groups (x, d) as input and output is called the training data. Specifically, Figure 2 The set of groups consisting of the color density value of each pixel in the monochrome images of Y, Cb, and Cr, and the labels of the true image, is shown below. Figure 2 The training data shown.
[0208] Learning a neural network refers to adjusting the weight values w so that the set (x) for any input and output is optimized. n d n The output y(x) of the neural network when given input xn n :w) Try to get as close as possible to the output d n The error function is a measure of how closely the function represented by the neural network approximates the training data.
[0209]
Number 8
[0210] y(x n :w)≈d n
[0211] The error function is also called the loss function. The error function E(w) used in disease analysis methods is shown in Equation 6 below. Equation 6 is called cross entropy.
[0212]
Number 9
[0213]
[0214] The method for calculating the cross-entropy in Equation 6 is explained. In the output layer 50b of the neural network 50 used in the disease analysis method, i.e., the final layer of the neural network, an activation function is used to classify the input x into a finite number of levels based on its content. The activation function is called the softmax function, as shown in Equation 7 below. It should be noted that the output layer 50b has the same number of nodes as the number of levels k. The total input u of each node k (k = 1, ... K) of the output layer L is provided by the output of the previous layer L-1 through uk(L). Thus, the output of the kth node of the output layer is as shown in Equation 7 below.
[0215]
Number 10
[0216]
[0217] Equation 7 is the flexible maximum value function. The outputs y1, ..., y1 determined by Equation 7 are... K The sum is usually 1.
[0218] If we represent each level as C1, ..., C K Then the output y of node k in the output layer L K (i.e. u k (L) The probability that the given input x belongs to level CK is shown in Equation 8 below. The probability of input x being classified into the level shown in Equation 8 reaches its maximum.
[0219]
Number 11
[0220]
[0221] In neural network learning, the function represented by the neural network is regarded as a model of posterior probability at each level. Under such a probability model, the likelihood of the weight value w relative to the training data is evaluated, and the weight value w that maximizes the likelihood is selected.
[0222] Only when the output level is correct will the target output d be obtained through the flexible maximum function of (Equation 7).n The output is recorded as 1, and otherwise as 0. When the target output is represented by d... n =[d n1 ,···,d nK When represented in this vector form, for example, when the input x... n When the correctness level is C3, then only the target output d is considered. n3 The target output is denoted as 1, and anything else is denoted as 0. With this encoding, the posterior distribution is as shown in Equation 9.
[0223]
Number 12
[0224]
[0225] With training data {(x n d n The likelihood L(w) of the relative weight values w for n = 1, ..., N is shown in Equation 10 below. Taking the logarithm of the likelihood L(w) and reversing the sign, we obtain the error function in Equation 6.
[0226]
Number 13
[0227]
[0228] Learning refers to minimizing the error function E(w) calculated based on the training data for the parameters w of a neural network. In disease analysis methods, the error function E(w) is shown in Equation 6.
[0229] Minimizing the error function E(w) with respect to parameter w has the same meaning as finding a local minimum of the function E(w). Parameter w represents the weight value of the connection between nodes. The minimum point of the weight value w is obtained through iterative calculation, which starts with arbitrary initial values and repeatedly updates the parameter w. As an example of this calculation, there is the gradient descent method.
[0230] In the gradient descent method, the vector shown in Equation 11 is used.
[0231]
Number 14
[0232]
[0233] In gradient descent, the process is repeated multiple times to make the current value of the parameter w move in the negative gradient direction (i.e., ...). The handling of movement. If the current weight value is w (t) The weight after the shift is w (t+1)The gradient descent method is calculated as shown in Equation 12 below. The value t refers to the number of times the parameter w is shifted.
[0234]
Number 15
[0235]
[0236] symbol
[0237]
Number 16
[0238] ∈
[0239] The constant that determines the magnitude of the update of parameter w is called the learning coefficient. By repeatedly performing the operation shown in Equation 12, the error function E(w) increases with the value t. (t) As the parameter w decreases, it reaches its minimum.
[0240] It should be noted that the operation of (Equation 12) can be performed on all training data (n = 1, ..., N), or it can be performed on only a portion of the training data. The gradient descent method performed on only a portion of the training data is called stochastic gradient descent. Stochastic gradient descent is used in disease analysis methods.
[0241] (Machine Learning Process 1)
[0242] The first machine learning process trains a machine learning algorithm based on the first or second information.57
[0243] Reference Figure 12 The processing unit 10A of the training device 100A trains a machine learning algorithm 57 based on the first information 53 and the disease information 55. The processing unit 10A includes a machine learning training data generation unit 101a, a machine learning training data input unit 102a, and a machine learning algorithm update unit 103a. These functional modules are implemented by installing a program that enables a computer to perform machine learning processing in the storage unit 13 or memory 12 of the processing unit 10A, and executing the program by the CPU 11. A machine learning training data database (DB) 104a and a machine learning algorithm database (DB) 105a are recorded in the storage unit 13 or memory 12 of the processing unit 10A.
[0244] The first or second information is generated by the processing unit 10A, for example, establishing a correspondence with the morphological cell type and abnormalities observed in the analyzed cell, and is pre-stored in the storage unit 13 or memory 12 of the processing unit 10A. The machine learning algorithm database 105 pre-contains untrained first neural network 50, second neural network 51, and second neural network 52. In the machine learning algorithm database 105a, the first neural network 50, second neural network 51, and second neural network 52 that have undergone one training and received an update are pre-contained.
[0245] The processing unit 10A of the training device 100A performs... Figure 13 The processing shown. If using Figure 12 The functional modules shown will be explained below. Steps S111, S112, S114, and S115 are processed by the machine learning training data generation unit 101a. Step S113 is processed by the machine learning training data input unit 102a.
[0246] use Figure 13 An example of the first machine learning process performed by the processing unit 10A will be explained.
[0247] In step S111, the processing unit 10A of the training device 100A generates first information or second information according to the method described in the training section of the machine learning algorithm. Specifically, the processing unit 10A uses the first deep learning algorithm or the second deep learning algorithm trained through steps S11 to S16 to identify the types of abnormalities seen in the cells of each training image 70, and obtains identification results. The second neural network 61 outputs identification results of the types of abnormalities seen for each cell. In step S111, the processing unit 10A generates first information for each sample of the training image 70 based on the identification results of the types of abnormalities seen. In addition, the processing unit 10A uses the second neural network 62 to identify the cell types of the cells in each training image 70, and obtains identification results. The processing unit 10A generates second information for each sample of the training image 70 based on the cell type identification results.
[0248] Then, in step S112, the processing unit 10A generates machine learning training data 90 from the first information and the disease information 55 associated with the training image 75. Alternatively, the processing unit 10A generates machine learning training data 90 from the second information and the disease information 55 associated with the training image 75.
[0249] Then, in step S113, the processing unit 10A inputs the machine learning training data 90 into the machine learning algorithm to train the machine learning algorithm.
[0250] Then, in step S114, the processing unit 10A determines whether processing has been performed on all training samples. If processing has been performed on all training samples, the processing ends. If processing has not been performed on all training samples, the process proceeds to step S115, where the identification result of the type of anomaly or cell type of another sample is obtained, and the process returns to step S111 to repeatedly train the machine learning algorithm.
[0251] (Machine Learning Processing 2)
[0252] The second machine learning process uses the first and second information to train the machine learning algorithm 57.
[0253] Reference Figure 12 The processing unit 10A of the training device 100A trains a machine learning algorithm 57 based on the first information 53 and the disease information 55. The processing unit 10A includes a machine learning training data generation unit 101b, a machine learning training data input unit 102b, and a machine learning algorithm update unit 103b. These functional modules are implemented by installing a program that enables a computer to perform machine learning processing in the storage unit 13 or memory 12 of the processing unit 10A, and executing the program by the CPU 11. The storage unit 13 or memory 12 of the processing unit 10A stores a machine learning training data database (DB) 104b and a machine learning algorithm database (DB) 105b.
[0254] The first or second information is generated by the processing unit 10A, for example, establishing a correspondence with the morphological cell type and abnormalities observed in the analyzed cell, and is pre-stored in the storage unit 13 or memory 12 of the processing unit 10A. The machine learning algorithm database 105 pre-contains untrained first neural network 50, second neural network 51, and second neural network 52. The machine learning algorithm database 105b pre-contains first neural network 50, second neural network 51, and second neural network 52 that have undergone one training iteration and have been updated.
[0255] The processing unit 10A of the training device 100A performs... Figure 14 The processing shown. If using Figure 12 The functional modules shown will be explained below. Steps S1111, S1112, S1114, and S1115 are processed by the machine learning training data generation unit 101b. Step S1113 is processed by the machine learning training data input unit 102.
[0256] use Figure 14An example of the second machine learning processing performed by the processing unit 10A will be described below. In step S1111, the processing unit 10A of the training device 100A generates first information and second information according to the method described in the training section of the machine learning algorithm described above. Specifically, the processing unit 10A uses the first deep learning algorithm and the second deep learning algorithm trained through steps S11 to S16 to identify the types of abnormalities seen in the cells of each training image 70 and obtains identification results. The second neural network 61 outputs the identification results of the types of abnormalities seen for each cell. In step S1111, the processing unit 10A generates first information for each sample of the training image 70 based on the identification results of the types of abnormalities seen. In addition, the processing unit 10A uses the second neural network 62 to identify the cell types of the cells in each training image 70 and obtains identification results. The processing unit 10A generates second information for each sample of the training image 70 based on the cell type identification results.
[0257] Then, in step S1112, the processing unit 10A generates machine learning training data 90 from the first information, the second information, and the disease information 55 associated with the training image 75.
[0258] Then, in step S1113, the processing unit 10A inputs the machine learning training data 90 into the machine learning algorithm to train the machine learning algorithm.
[0259] Then, in step S1114, the processing unit 10A determines whether processing has been performed on all training samples. If processing has been performed on all training samples, the processing ends. If processing has not been performed on all training samples, the process proceeds to step S1115, where the identification result of the type of anomaly or cell type of another sample is obtained, and the process returns to step S1111 to repeatedly train the machine learning algorithm.
[0260] The machine learning algorithms used in steps S113 and S1113 are summarized below.
[0261] As a machine learning algorithm, ensemble learning (a classifier composed of multiple classifiers) can be used, such as gradient boosting. Examples of ensemble learning include extreme gradient boosting (EGB) and stochastic gradient boosting. Gradient boosting is a type of boosting algorithm used to construct multiple weak learners. Weak learners can be, for example, regression trees.
[0262] For example, in the case of regression trees, the input vector is set to x, and the label is set to y. The weak learners fm(x), m = 1, 2, ..., M, learn and integrate successively, so that the overall learner...
[0263]
Number 17
[0264] F(x)=f0(x)+f1(x)+f2(x)+...+f M (x)
[0265] The loss function L(y, F(x)) is minimized. That is, at the beginning of learning, a function F0(x) = f0(x) is provided, and a weak learner fm(x) is determined in the m-th learning step so that the learner contains m weak learners.
[0266]
Number 18
[0267] F m (x)=F m-1 (x)+f m (x)
[0268] The loss function L(y, F(x)) is minimized. In ensemble learning, when optimizing a weak learner, instead of using all the data in the training set, a fixed set of data is randomly selected.
[0269]
Number 19
[0270] (1) Obtain the constant function F0(x) that minimizes the loss.
[0271] (2) For m=1 to M
[0272] (a) Take N data points from the training set to obtain set D.
[0273] (b) For each element (x1, y1), (x2, y2), ... (x...) of set D... N y N ),
[0274] Calculate gradient
[0275] (c) Generate a regression tree T(x) of the predicted gradient.
[0276] That is, generation makes To achieve the minimum regression tree.
[0277] This regression tree is a weak learner f m (x).
[0278] (d) Optimize the weights of the leaves of the regression tree T to minimize the loss.
[0279] To reach the minimum.
[0280] (e)F i (x)=F m-1 (x)+vT(x)·v represents shrinkage
[0281] The parameter is a constant that satisfies 0 < v ≤ 1.
[0282] (3) Output F M (x), as F(x).
[0283] Specifically, the learner F(x) is obtained through the following algorithm. The shrinkage parameter v is set to 1, and F0(x) can be transformed by a constant function.
[0284] (Disease Analysis and Processing)
[0285] Figure 15 A functional block diagram of a disease analysis apparatus 200A is shown, illustrating disease analysis processing performed on an image 78 of the analysis object until an analysis result 83 is generated. The processing unit 20A of the disease analysis apparatus 200A includes an analysis data generation unit 201, an analysis data input unit 202, and an analysis unit 203. These functional modules are implemented by installing a program that enables the computer of the present invention to perform disease analysis processing in the storage unit 23 or memory 22 of the processing unit 20A, and executing the program by the CPU 21. A deep learning training data database (DB) 104 and a deep learning algorithm database (DB) 105 are provided by the training device 100A via a recording medium 98 or a network 99 and recorded in the storage unit 23 or memory 22 of the processing unit 20A. A machine learning training data database (DB) 104a, b and a deep learning algorithm database (DB) 105a, b are provided by the training device 100A via a recording medium 98 or a network 99 and recorded in the storage unit 23 or memory 22 of the processing unit 20A.
[0286] The image 78 of the analysis object is captured by the imaging device 400 and stored in the storage unit 23 or memory 22 of the processing unit 20A. A first neural network 60 and a second neural network 61, 62, containing trained connection weights w, establish correspondences with, for example, cell types based on morphological classification to which the analysis object cells belong, and types of abnormalities observed, and are contained in the deep learning algorithm database 105, functioning as a program module. This program module is part of a program that enables a computer to perform disease analysis processing. That is, the first neural network 60 and the second neural networks 61, 62 are used in a computer equipped with a CPU and memory to output the determination result of the type of abnormality observed or the identification result of the cell type. The CPU 21 of the processing unit 20A enables the computer to perform information operations or processing specific to the intended use. Furthermore, a trained machine learning algorithm 67 is contained in the machine learning algorithm databases 105a, b, functioning as a program module. This program module is part of a program that enables a computer to perform disease analysis processing. That is, the machine learning algorithm 67 is used in a computer equipped with a CPU and memory to output disease analysis results.
[0287] Specifically, the CPU 21 of the processing unit 20A uses a first deep learning algorithm stored in the storage unit 23 or the memory 22 to generate an identification result of the type of anomaly observed in the analysis data generation unit 201. The processing unit 20A generates first information 63 based on the identification result of the type of anomaly observed in the analysis data generation unit 201. The generated first information 63 is input into the analysis data input unit 202 and stored in the machine learning training data DB104a. The processing unit 20A performs disease analysis in the analysis unit 203 and outputs the analysis result 83 to the output unit 27. Alternatively, the CPU 21 of the processing unit 20A uses a second deep learning algorithm stored in the storage unit 23 or the memory 22 to generate an identification result of the type of anomaly observed in the analysis data generation unit 201. The processing unit 20A generates second information 64 based on the identification result of the type of anomaly observed in the analysis data generation unit 201. The generated second information 64 is input into the analysis data input unit 202 and stored in the machine learning training data DB104b. The processing unit 20A analyzes the disease in the analysis unit 203 and outputs the analysis results 83 to the output unit 27.
[0288] If using Figure 16 The functional modules shown are explained below. Steps S21 and S22 are processed by the analysis data generation unit 201. Steps S23, S24, S25, and S27 are processed by the analysis data input unit 202. Step S26 is processed by the analysis unit 203.
[0289] (Disease Analysis and Processing 1)
[0290] use Figure 16 This describes an example of a first disease analysis process performed by the processing unit 20A, from analyzing the target image 78 to outputting the analysis result 83. The first disease analysis process outputs the analysis result 83 based on either first information or second information.
[0291] First, the processing unit 20A acquires the analysis image 78. The analysis image 78 can be acquired by the imaging device 400 through user operation, or by the recording medium 98, or via the network using the I / F unit 25.
[0292] In step S21, with Figure 10 Similarly, step S11 converts the obtained analysis image 78 into luminance Y, first hue Cb, and second hue Cr, and generates hue vector data 80 according to the steps described in the above analysis data generation method.
[0293] Then, in step S22, the processing unit 20A generates analysis data 81 from the hue vector data 80 according to the steps described in the above analysis data generation method.
[0294] Then, in step S23, the processing unit 20A obtains the first deep learning algorithm or the second deep learning algorithm stored in the algorithm database 105.
[0295] Then, in step S24, the processing unit 20A inputs the analysis data 81 into the first neural network 60 constituting the first deep learning algorithm. Following the steps described in the disease analysis method above, the processing unit 20A inputs the feature values output by the first neural network 60 into the second neural network 61, and the second neural network 61 outputs an identification result of the type of abnormality. The processing unit 20A stores this identification result in the memory 22 or the storage unit 23. Alternatively, in step S24, the processing unit 20A inputs the analysis data 81 into the first neural network 60 constituting the second deep learning algorithm. Following the steps described in the disease analysis method above, the processing unit 20A inputs the feature values output by the first neural network 60 into the second neural network 62, and the second neural network 62 outputs an identification result of the cell type. The processing unit 20A stores this identification result in the memory 22 or the storage unit 23.
[0296] In step S25, the processing unit 20A first determines whether all the acquired analytical images 78 have been identified. If the identification of all analytical images 78 is completed (YES), the process proceeds to step S26, where first information 63 or second information is generated based on the cell type identification result. If the identification of all analytical images 78 is not completed (NO), the process proceeds to step S27, where the analytical images 78 that have not yet been identified are processed according to steps S21 to S25.
[0297] Then, in step S28, the processing unit 20A obtains the machine learning algorithm 67. Next, in step S29, the processing unit 20A inputs either the first information or the second information into the machine learning algorithm 67.
[0298] Finally, in step S30, the processing unit 20A outputs the analysis result 83 to the output unit 27 in the form of a disease name or a tag value associated with the disease name.
[0299] (Disease Analysis and Processing 2)
[0300] use Figure 17 This describes an example of a second disease analysis process performed by the processing unit 20A, from analyzing the target image 78 to outputting the analysis result 83. The second disease analysis process outputs the analysis result 83 based on the first information and the second information.
[0301] First, the processing unit 20A acquires the analysis image 78. The analysis image 78 can be acquired by the imaging device 400 through user operation, or by the recording medium 98, or via the network using the I / F unit 25.
[0302] In step S121, with Figure 10 Similarly, step S11 converts the obtained analysis image 78 into luminance Y, first hue Cb, and second hue Cr, and generates hue vector data 80 according to the steps described in the above analysis data generation method.
[0303] Then, in step S122, the processing unit 20A generates analysis data 81 from the hue vector data 80 according to the steps described in the above-described method for generating analysis data.
[0304] Then, in step S123, the processing unit 20A obtains the first deep learning algorithm and the second deep learning algorithm stored in the algorithm database 105.
[0305] Then, in step S124, the processing unit 20A inputs the analysis data 81 into the first neural network 60 constituting the first deep learning algorithm. Following the steps described in the disease analysis method above, the processing unit 20A inputs the feature values output by the first neural network 60 into the second neural network 61, and the second neural network 61 outputs an identification result of the type of abnormality. The processing unit 20A stores this identification result in the memory 22 or the storage unit 23. Additionally, in step S124, the processing unit 20A inputs the analysis data 81 into the first neural network 60 constituting the second deep learning algorithm. Following the steps described in the disease analysis method above, the processing unit 20A inputs the feature values output by the first neural network 60 into the second neural network 62, and the second neural network 62 outputs an identification result of the cell type. The processing unit 20A stores this identification result in the memory 22 or the storage unit 23.
[0306] In step S125, the processing unit 20A first determines whether all the acquired analytical images 78 have been identified. If the identification of all analytical images 78 is completed (YES), the process proceeds to step S126, where first information 63 is generated from the cell type identification result, and second information is generated from the cell type identification result. If the identification of all analytical images 78 is not completed (NO), the process proceeds to step S127, where the analytical images 78 that have not yet been identified are processed according to steps S121 to S125.
[0307] Then, in step S128, the processing unit 20A obtains the machine learning algorithm 67. Next, in step S129, the processing unit 20A inputs the first information and the second information into the machine learning algorithm 67.
[0308] Finally, in step S130, the processing unit 20A outputs the analysis result 83 to the output unit 27 in the form of a disease name or a tag value associated with the disease name.
[0309] <Computer Programs>
[0310] A computer program for assisting disease analysis is described, which causes the computer to perform the processes of steps S21 to S30 or steps S121 to S130. The computer program may include a program for training a machine learning algorithm that causes the computer to perform the processes of steps S11 to S17 and steps S111 to S115, or a program for training a machine learning algorithm that causes the computer to perform the processes of steps S11 to S17 and steps S1111 to S1115.
[0311] Furthermore, a program article, such as a storage medium storing the aforementioned computer program, will be described. The computer program is stored in a semiconductor storage element such as a hard disk or flash memory, or a storage medium such as an optical disc. There are no restrictions on the storage format of the program stored on the aforementioned storage medium, as long as the processing unit can read the program. Preferably, the storage on the aforementioned storage medium is non-erasable.
[0312] [Disease Analysis System 2]
[0313] <Composition of Disease Analysis System 2>
[0314] Another approach to disease analysis systems will be described. Figure 18 An example of the configuration of the second disease analysis system 2 is shown. The disease analysis system 2 includes a user-side device 200, which operates as a comprehensive disease analysis device 200B. The disease analysis device 200B is, for example, composed of a general-purpose computer, and performs both deep learning processing and disease analysis processing as described in the disease analysis system 1 above. That is, the disease analysis system 2 is a stand-alone system that performs deep learning and disease analysis on the user side. In the second disease analysis system, the comprehensive disease analysis device 200B located on the user side performs the functions of both the training device 100A and the disease analysis device 200A.
[0315] Figure 18 In this embodiment, the disease analysis device 200B is connected to the imaging device 400. The imaging device 400 captures training images 70 during deep learning processing and captures images of the analysis target 78 during disease analysis processing.
[0316] <Hardware Configuration>
[0317] Hardware configuration of the disease analysis device 200B and Figure 9 The hardware configuration of the user-side device 200 shown is the same.
[0318] <Functional Modules and Processing Steps>
[0319] Figure 19 A functional block diagram of the disease analysis device 200B is shown. The processing unit 20B of the disease analysis device 200B includes a deep learning training data generation unit 101, a deep learning training data input unit 102, a deep learning algorithm update unit 103, machine learning training data generation units 101a, b, machine learning training data input units 102a, b, machine learning algorithm update units 103a, b, analysis data generation unit 201, analysis data input unit 202, and analysis unit 203.
[0320] The processing unit 20B of the disease analysis device 200B performs deep learning processing. Figure 9 The processing shown is performed during machine learning. Figure 13 or Figure 14 The processing is carried out during disease analysis and processing. Figure 16 or Figure 17 The processing shown. If using Figure 19 The functional modules shown are explained below, so that during deep learning processing, Figure 9 The processing of steps S11, S12, S16 and S17 is performed by the deep learning training data generation unit 101. Figure 9 The processing in step S13 is performed by the deep learning training data input unit 102. Figure 9 The processing in step S14 is performed by the deep learning algorithm update unit 103. During machine learning processing, Figure 13 The processing of steps S111, S112, S114, and S115 is performed by the machine learning training data generation unit 101a. The processing of step S113 is performed by the machine learning training data input unit 102a. Alternatively, during machine learning, Figure 14 The processing of steps S1111, S1112, S1114 and S1115 is performed by the machine learning training data generation unit 101b. Figure 14 The processing in step S1113 is performed by the machine learning training data input unit 102. During disease analysis processing, Figure 16 The processing in steps S21 and S22 is performed by the analysis data generation unit 201. Figure 16 The processing in steps S21 and S22 is performed by the analysis data generation unit 201. Figure 16 The processing of steps S23, S24, S25 and S27 is performed by the analysis data input unit 202. Figure 16 The processing in step S26 is performed by the analysis unit 203. Alternatively, the processing in steps S121 and S122 of FIG17 is performed by the analysis data generation unit 201. Figure 17 The processing of steps S123, S124, S125 and S127 is performed by the analysis data input unit 202. Figure 16 The processing in step S126 is performed by the analysis unit 203.
[0321] The deep learning processing steps and disease analysis processing steps performed by the disease analysis device 200B are the same as those performed by the training device 100A and the disease analysis device 200A, respectively. However, the disease analysis device 200B acquires the training image 70 from the imaging device 400.
[0322] In the disease analysis device 200B, the user can verify the recognition accuracy of the recognizer. If the recognizer's result differs from the result obtained by the user observing the image, the analysis data 81 can be used as training data 78, and the recognition result obtained by the user observing the image can be used as the label value 77 to retrain the first deep learning algorithm and the second deep learning algorithm. This can further improve the training efficiency of the first neural network 50 and the first neural network 51.
[0323] [Disease Analysis System 3]
[0324] <Composition of Disease Analysis System 3>
[0325] Another approach to disease analysis systems will be described. Figure 20 An example configuration of the third disease analysis system 3 is shown. The disease analysis system 3 includes a supplier-side device 100 and a user-side device 200. The supplier-side device 100 includes a processing unit 10 (10B), an input unit 16, and an output unit 17. The supplier-side device 100 operates as a comprehensive disease analysis device 100B, and the user-side device 200 operates as a terminal device 200C. The disease analysis device 100B is, for example, a general-purpose computer, serving as a cloud service-side device for performing the deep learning processing and disease analysis processing described in the disease analysis system 1 above. The terminal device 200C is, for example, a general-purpose computer, which sends images of the analysis object to the disease analysis device 100B via a network 99, and receives images of the analysis results from the disease analysis device 100B via the network 99.
[0326] In the disease analysis system 3, the integrated disease analysis device 100B, located on the supplier side, performs the functions of both the training device 100A and the disease analysis device 200A. On the other hand, the third disease analysis system includes a terminal device 200C, which provides an input interface for the analysis image 78 and an output interface for the analysis result image to the user-side terminal device 200C. That is, in the case of the third disease analysis system, the supplier side performing deep learning processing and disease analysis processing is a cloud service-type system that provides the analysis image 78 to the user-side input interface and the analysis result 83 to the user-side output interface. The input and output interfaces can be integrated.
[0327] The disease analysis device 100B is connected to the imaging device 300 to acquire training images 70 captured by the imaging device 300.
[0328] The terminal device 200C is connected to the imaging device 400 to obtain the image 78 of the analysis object captured by the imaging device 400.
[0329] <Hardware Configuration>
[0330] Hardware configuration of the disease analysis device 100B and Figure 7 The hardware configuration of the supplier-side device 100 shown is the same. The hardware configuration of the terminal device 200C is the same as that of the supplier-side device 100 shown. Figure 8 The hardware configuration of the user-side device 200 shown is the same.
[0331] <Functional Modules and Processing Steps>
[0332] Figure 21 A functional block diagram of the disease analysis device 100B is shown. The processing unit 10B of the disease analysis device 100B includes a deep learning training data generation unit 101, a deep learning training data input unit 102, a deep learning algorithm update unit 103, machine learning training data generation units 101a, b, machine learning training data input units 102a, b, machine learning algorithm update units 103a, b, analysis data generation unit 201, analysis data input unit 202, and analysis unit 203.
[0333] The processing unit 20B of the disease analysis device 200B performs deep learning processing. Figure 9 The processing shown is performed during machine learning. Figure 13 or Figure 14 The processing is carried out during disease analysis and processing. Figure 16 or Figure 17 The processing shown. If using Figure 19 The functional modules shown are explained below, so that during deep learning processing, Figure 9 The processing of steps S11, S12, S16 and S17 is performed by the deep learning training data generation unit 101. Figure 9 The processing in step S13 is performed by the deep learning training data input unit 102. Figure 9 The processing in step S14 is performed by the deep learning algorithm update unit 103. During machine learning processing, Figure 13 The processing of steps S111, S112, S114, and S115 is performed by the machine learning training data generation unit 101a. The processing of step S113 is performed by the machine learning training data input unit 102a. Alternatively, during machine learning, Figure 14 The processing of steps S1111, S1112, S1114 and S1115 is performed by the machine learning training data generation unit 101b. Figure 14 The processing in step S1113 is performed by the machine learning training data input unit 102. During disease analysis processing, Figure 16 The processing in steps S21 and S22 is performed by the analysis data generation unit 201. Figure 16 The processing in steps S21 and S22 is performed by the analysis data generation unit 201. Figure 16 The processing of steps S23, S24, S25 and S27 is performed by the analysis data input unit 202. Figure 16 The processing in step S26 is performed by the analysis unit 203. Alternatively, Figure 17 The processing of steps S121 and S122 is performed by the analysis data generation unit 201. Figure 17 The processing of steps S123, S124, S125 and S127 is performed by the analysis data input unit 202. Figure 16 The processing in step S126 is performed by the analysis unit 203.
[0334] The deep learning processing steps and disease analysis processing steps performed by the disease analysis device 100B are the same as those performed by the training device 100A and the disease analysis device 200A, respectively.
[0335] The processing unit 10B receives the image 78 of the object to be analyzed from the user-side terminal device 200C, and then processes it according to... Figure 9 The steps S11 to S17 shown generate training data 75 for deep learning.
[0336] exist Figure 12 In step S26, the processing unit 10B sends the analysis results, including the analysis result 83, to the user-side terminal device 200C. In the user-side terminal device 200C, the processing unit 20C outputs the received analysis results to the output unit 27.
[0337] In summary, the user of terminal device 200C can obtain analysis results 83 by sending the image 78 of the object to be analyzed to disease analysis device 100B.
[0338] According to the disease analysis device 100B, users can use the recognizer without obtaining the training data database 104 and algorithm database 105 from the training device 100A. Therefore, the service of identifying cell types and cell characteristics based on morphological classification can be provided as a cloud service.
[0339] [Other methods]
[0340] The present invention is not limited to the above-described methods.
[0341] The above method illustrates a way to generate training data 75 for deep learning by converting hue to brightness Y, primary hue Cb, and secondary hue Cr. However, the hue conversion is not limited to this. For example, red (R), green (G), and blue (B) can be used directly without hue conversion. Alternatively, it can be a binary primary color obtained by subtracting any hue from the above primary colors. Or, it can be any one of the three primary colors (red (R), green (G), and blue (B)) (e.g., green (G)), i.e., a single primary color). It can also be converted to cyan (C), magenta (M), and yellow (Y), i.e., the three primary colors. For example, for the image 78 being analyzed, it is not limited to a color image with three primary colors (red (R), green (G), and blue (B); it can be a color image with two primary colors, as long as it contains one or more primary colors.
[0342] In the above-described methods for generating training data and analysis data, in step S11, processing units 10A, 20B, and 10B generate hue rows 72y, 72cb, and 72cr from the training image 70. However, it is also possible to convert the training image 70 into an image of luminance Y, first hue Cb, and second hue Cr. That is, processing units 10A, 20B, and 10B can initially obtain luminance Y, first hue Cb, and second hue Cr directly from, for example, a virtual slice scanner. Similarly, in step S21, processing units 20A, 20B, and 10B generate hue rows 72y, 72cb, and 72cr from the analysis target image 78. However, processing units 20A, 20B, and 10B can also initially obtain luminance Y, first hue Cb, and second hue Cr directly from, for example, a virtual slice scanner.
[0343] In addition to RGB and CMY, YUV and CIE L can also be used. * a * b * Various color spaces are used for image acquisition and hue conversion.
[0344] In tone vector data 74 and tone vector data 80, for each pixel, tone information is stored in the order of luminance Y, first hue Cb, and second hue Cr, but the order in which the tone information is stored and the processing order are not limited to this. However, it is preferable that the arrangement order of the tone information in tone vector data 74 is the same as the arrangement order of the tone information in tone vector data 80.
[0345] In various image analysis systems, processing units 10A and 10B can be implemented as a single unit. However, processing units 10A and 10B do not necessarily have to be a single unit; the CPU 11, memory 12, storage unit 13, GPU 19, etc., can also be arranged in different locations and connected via a network. Processing units 10A and 10B, input unit 16, and output unit 17 also do not necessarily have to be located in one place; they can be arranged in different locations and connected in a manner that allows them to communicate with each other via a network. The same applies to processing units 20A, 20B, and 20C.
[0346] In the aforementioned disease analysis assistance system, the functional modules of the deep learning training data generation unit 101, machine learning training data generation units 101a and 101b, deep learning training data input unit 102, machine learning training data input units 102a and 102b, deep learning algorithm update unit 103, machine learning algorithm update units 103a and 103b, analysis data generation unit 201, analysis data input unit 202, and analysis unit 203 are executed on a single CPU 11 or a single CPU 21. However, these functional modules do not necessarily have to be executed on a single CPU; they can be distributed across multiple CPUs. Furthermore, these functional modules can be distributed across multiple GPUs, or across multiple CPUs and multiple GPUs.
[0347] In the aforementioned disease analysis support system, it will be used to perform Figure 9 and Figure 12 The processing procedures for each step described are pre-recorded in storage units 13 and 23. Alternatively, the program can be installed in processing units 10B and 20B from a computer-readable, non-transitory tangible recording medium 98 such as a DVD-ROM or USB memory. Alternatively, processing units 10B and 20B can be connected to a network 99, through which the program can be downloaded and installed from, for example, an external server (not shown).
[0348] In various disease analysis systems, input units 16 and 26 are input devices such as keyboards or mice, while output units 17 and 27 are display devices such as LCD displays. Alternatively, input units 16 and 26 and output units 17 and 27 can be integrated and implemented as a touch panel display device. Or, output units 17 and 27 can be constructed using a printer or similar device.
[0349] In the aforementioned disease analysis systems, the imaging device 300 is directly connected to the training device 100A or the disease analysis device 100B, but it can also be connected to the training device 100A or the disease analysis device 100B via network 99. Similarly, the imaging device 400 is directly connected to the disease analysis device 200A or the disease analysis device 200B, but it can also be connected to the disease analysis device 200A or the disease analysis device 200B via network 99.
[0350] [The effect of the recognizer]
[0351] <Training of Deep Learning and Machine Learning Algorithms>
[0352] In the evaluation, a total of 3,261 peripheral blood (PB) smears obtained from Juntendo University Hospital between 2017 and 2018 were used, including 1,165 PB smears from subjects with hematologic disorders (myelodysplastic syndrome (n=94), myeloproliferative neoplasms (n=127), acute myeloid leukemia (n=38), acute lymphoblastic leukemia (n=27), malignant lymphoma (n=324), multiple myeloma (n=82), and non-neoplastic hematologic disorders (n=473)). PB smear slides were prepared using a smear preparation device SP-10 (manufactured by Sysmex Corporation) and stained with May Grunwald-Giemsa. A total of 703,970 digitized cell images were obtained from the PB smear slides using an automated hematography analysis device DI-60 (manufactured by Sysmex Corporation). Following the method described above for generating training data for deep learning, training data 75 for deep learning is generated from images.
[0353] As the primary computer algorithm, a deep learning algorithm is used. The deep learning algorithm uses a Convolutional Neural Network (CNN) as the first neural network and a Fully Connected Neural Network (FCNN) as the second neural network to identify cell types and the types of abnormalities observed.
[0354] As a second computer algorithm, Extreme Gradient Boosting (EGB), a machine learning algorithm, is used to build an automated disease analysis assistance system.
[0355] Figure 22The configuration of the identifier used in the embodiment is shown. A deep learning algorithm is systematized to simultaneously detect cell types and types of anomalies.
[0356] This deep learning algorithm consists of two main modules: the "CNN module" and the "FCNN module." The CNN module extracts features represented by tone vector data from images captured by the DI-60. The FCNN module analyzes the features extracted by the CNN module, classifying cell images according to 97 abnormal features such as cell and nucleus size, morphology, and cytoplasmic image patterns, as well as 17 cell types.
[0357] The CNN module consists of two sub-modules. The initial (upstream) sub-module has three joint blocks, each with two parallel paths containing several convolutional network layers. This stacking of layers optimizes feature extraction from the input image data and output parameters to the next block. The second (downstream) sub-module contains eight consecutive blocks. Each block has a series of convolutional layers and two parallel paths consisting of one path without convolutional components. This is called a Residual Network (ResNet). The ResNet acts as a buffer to prevent system saturation.
[0358] Each layer—separable convolutions, exception-based convolutional layers (Conv 2D), batch normalization (BN), and activation layers (ACT)—serves a different purpose. Separable convolutions are a variant of convolution called Xception. Conv 2D is the main building block of the neural network, optimizing parameters during feature extraction, image processing, and the formation of a "feature map." The ACT following Conv 2D and BN layers is a Rectified Linear Unit (ReLU). The first submodule connects to a second submodule consisting of eight consecutive similar blocks to create the feature map. To avoid unpredictable saturation in deep layers caused by weight effects due to backpropagation, the second module acts as a bypass in Conv 2D. This deep convolutional neural network architecture is implemented using the backends of Keras and Tensorflow.
[0359] The identification results of 17 cell types identified by the first computer algorithm and the identification results of 97 abnormalities observed in each cell type were used to train the machine learning algorithm. When identifying the abnormalities observed in each cell type, neutrophils were not differentiated into segmented neutrophils and band neutrophils, and a association was established with the abnormalities observed. Furthermore, Figure 3The items "Other Abnormalities," "Pseudo-Chediak-Higashi Granularity," and "Other Abnormalities (including Agglutination)" for platelets shown in the abnormality report were excluded from the analysis. In XGBoost, the first information is generated from the identification results of the abnormality type for each cell type, and the second information is generated from the cell type identification results and then input.
[0360] To train the deep learning algorithm, the 703,970 digitized cell images were divided into a training dataset of 695,030 images and a testing dataset of 8,940 images.
[0361] To build the system, peripheral blood cells from 89 cases of myelodysplastic syndrome (MDS) and 43 cases of aplastic anemia (AA) were used for training. Then, images of PB smear specimens obtained from 26 MDS patients and 11 AA patients were used to test the EGB-based automated disease analysis assistance system.
[0362] The identification of cells used in training was performed using morphological criteria from the Clinical Examination Standards Institute (CLSI) H20-A2 guidelines and the 2016 revised WHO classification of myeloma and acute leukemia, by two committee-certified hematologists and one senior hematologist. The training dataset was categorized according to 17 cell types and 97 types of abnormalities observed.
[0363] Figure 23 The types and quantities of cell images used for training and testing are shown.
[0364] After training, the performance of the first computer algorithm was evaluated using the evaluation data set. Figure 24 The accuracy of cell species identification results obtained by the first computer algorithm after training is shown. The sensitivity and specificity calculated using ROC curves are good.
[0365] Figure 25 The accuracy of the anomaly identification results of the first computer algorithm, trained to this model, is shown. Sensitivity, specificity, and AUC, calculated using ROC curves, are all good.
[0366] This demonstrates that the first computer algorithm, after training, has good recognition accuracy.
[0367] Then, the identifier is used to distinguish between MDS and AA. Figure 26 A heatmap showing the contribution of each cell species to the observed abnormalities in terms of SHAP values. Figure 26The columns of the heatmap shown correspond to samples from one patient, and the rows correspond to abnormalities observed for each cell type. Patients in columns 1 to 26 on the left correspond to MDS patients, and patients in columns 27 to 37 on the right correspond to AA patients. The intensity of the heatmap indicates the detection rate. As shown in Figure 26, the detection rates of abnormal degranulation in neutrophils, abnormal granulation in eosinophils, and giant platelets are significantly higher in MDS patients than in AA patients.
[0368] Figure 27 The results of evaluating the accuracy of the identifier in the differential diagnosis of MDS and AA are shown. Evaluation was performed using sensitivity, specificity, and AUC calculated using ROC curves. The identifier's sensitivity and specificity were 96.2% and 100%, respectively, and the AUC of the ROC curve was 0.990, demonstrating high accuracy in the differential diagnosis of MDS and AA.
[0369] This indicates that the aforementioned identifier is useful for assisting in disease analysis.
[0370] [Symbol Explanation]
[0371] 200 Auxiliary devices for disease analysis
[0372] 20 Processing Department
[0373] 60. The first neural network
[0374] 61 The Second Neural Network
[0375] 62 The Second Neural Network
[0376] 67 Machine Learning Algorithms
[0377] 55 Disease Information
[0378] 53. Information related to the types of anomalies observed
[0379] 54. Information related to cell types
[0380] 81 Analyzing the data
Claims
1. A method for processing information used in disease analysis by a computer, comprising: acquiring first information by classifying each of a plurality of analysis target cells in an image of the analysis target cells obtained by photographing a smear sample of a blood specimen collected from a subject based on categories including at least eight of segmented neutrophil, band neutrophil, lymphocyte, monocyte, eosinophil, basophil, granuloblast, and platelet, acquiring second information by classifying abnormal findings possessed by a morphology of each of the analysis target cells based on a plurality of categories including nuclear morphology abnormality, cell size abnormality, cell deformity, cell destruction, vacuole, blast cell, inclusion body presence, Dohle body, satellite phenomenon, nuclear web abnormality, petaloid nucleus, large N / C ratio, bleb-like, fragmentation, and hairy cell-like morphology, using the first information and the second information corresponding to images of a plurality of cells contained in the blood specimen as information for analyzing aplastic anemia and / or myelodysplastic syndrome present in the subject using a computer algorithm.
2. The method of claim 1, wherein, The first information is related information of the number of cells of each cell category.
3. The method of claim 1, wherein, The second information is related information of the number of cells of each abnormality category.
4. The method of claim 1, wherein, In the process of classifying the morphology of each of the analysis target cells, analysis data including related information of each of the analysis target cells is input to a deep learning algorithm having a neural network structure, the morphology of each of the analysis target cells is classified using the deep learning algorithm.
5. The method according to claim 1, wherein the computer algorithm is a machine learning algorithm.
6. The method of claim 5, wherein, The machine learning algorithm is selected from a tree, a regression, a neural network, a Bayesian, a clustering, or an ensemble learning.
7. The method of claim 6, wherein, The machine learning algorithm is a gradient boosting tree.
8. The method of any one of claims 1 to 7, wherein, In the process of acquiring related information of the cell morphology classification, a probability that each of the analysis target cells belongs to each of a plurality of cell morphology classifications is calculated, a sum of the probabilities of each of the cell morphology classifications is calculated, the sum is acquired as the related information of the cell morphology classification.
9. The method according to claim 1, wherein the nuclear morphology abnormality includes at least one selected from the group consisting of hypersegmentation, hyposegmentation, pseudo Pelger-Huet abnormality, ring-shaped nucleus, spherical nucleus, oval-shaped nucleus, apoptosis, multinucleation, nuclear disruption, anucleated nucleus, naked nucleus, nuclear edge irregularity, nuclear fragmentation, intranuclear bridge, polynucleation, cleft nucleus, nuclear division, and nucleolus abnormality, the cell size abnormality includes a giant platelet.
10. The method of claim 1, wherein, The category of the cell further includes at least one selected from the group consisting of metamyelocyte, myelocyte, promyelocyte, granuloblast, plasma cell, atypical lymphocyte, immature eosinophil, immature basophil, erythroblast, and megakaryocyte.
11. An apparatus for assisting in disease analysis, comprising a processing unit that The first information is acquired by classifying each of the analysis target cells in images of a plurality of analysis target cells obtained by photographing a smear sample of a blood sample collected from a subject based on categories including at least 8 of segmented nucleus neutrophil, rod-shaped nucleus neutrophil, lymphocyte, monocyte, eosinophil, basophil, granuloblast, and platelet, The second information is acquired by classifying abnormal findings possessed by the form of each of the analysis target cells based on a plurality of categories including nuclear form abnormality, cell size abnormality, cell deformity, cell destruction, vacuole, immature cell, inclusion body presence, Dohle body, satellite phenomenon, nuclear web abnormality, petal-like nucleus, large N / C ratio, bleb-like, fragmentation, and hair-like cell-like form, The first information and the second information corresponding to images of a plurality of cells contained in the blood sample are used as information for analyzing aplastic anemia and / or myelodysplastic syndrome possessed by the subject using a computer algorithm.
12. A recording medium recording a program for assisting disease analysis, the program causing a computer to execute the steps of: The first information is acquired by classifying each of the analysis target cells in images of a plurality of analysis target cells obtained by photographing a smear sample of a blood sample collected from a subject based on categories including at least 8 of segmented nucleus neutrophil, rod-shaped nucleus neutrophil, lymphocyte, monocyte, eosinophil, basophil, granuloblast, and platelet, The second information is acquired by classifying abnormal findings possessed by the form of each of the analysis target cells based on a plurality of categories including nuclear form abnormality, cell size abnormality, cell deformity, cell destruction, vacuole, immature cell, inclusion body presence, Dohle body, satellite phenomenon, nuclear web abnormality, petal-like nucleus, large N / C ratio, bleb-like, fragmentation, and hair-like cell-like form, The first information and the second information corresponding to images of a plurality of cells contained in the blood sample are used as information for analyzing aplastic anemia and / or myelodysplastic syndrome possessed by the subject using a computer algorithm.
Citation Information
Patent Citations
Pathological tissue diagnosis support device
JP1998197522A
Method of using substance p analogs for treatment amelioration of myelodysplastic syndrome
US20090075903A1
Classifying biological samples using automated image analysis
US20180211380A1