Information processing method, information processing device, and method for creating machine learning model
The information processing method and device use a machine learning model to estimate multiple parameters from single cells, addressing the limitations of conventional analysis methods by enabling simultaneous estimation of various parameters.
Patent Information
- Application Number
- PCT/JP2025/001069
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-14
- Filing Date
- 2025-01-16
- Publication Date
- 2025-08-21
AI Technical Summary
Conventional methods for analyzing cellular parameters at the single-cell level are limited in the number of parameters that can be simultaneously analyzed, making it difficult to extract physiologically or pathologically important information due to technical and economic constraints.
An information processing method and device that utilizes a machine learning model to estimate additional parameters by learning the association between first and second specimen single-cell data, allowing for the estimation of multiple parameters from a single cell using a machine learning model trained on large datasets.
Enables the simultaneous estimation of a large number of parameters in a single cell, overcoming the limitations of conventional methods by estimating any combination of parameters that have been machine-learned, thereby facilitating comprehensive analysis.
Smart Images

Figure JP2025001069_21082025_PF_FP_ABST
Abstract
Description
Information processing method, information processing device, and machine learning model creation method
[0001] The present invention relates to an information processing method, an information processing device, and a method for creating a machine learning model.
[0002] Technologies for analyzing various cellular parameters at the single-cell level (e.g., single-cell RNA-seq, ATAC (Assay for Transposase-Accessible Chromatin)-seq, DNA methylation markers, chromosome conformation capture (Hi-C) method, mass cytometry analysis, etc.) have been developed. In recent years, technologies for simultaneously analyzing two or three of these parameters at the single-cell level have been disclosed. For example, Patent Document 1 discloses an image-based method for classifying human embryonic cells. Furthermore, Patent Document 2 discloses a method for barcoding multiple targets in cells using a barcode composition.
[0003] Japanese Patent Publication No. 2022-87297 Japanese Patent Publication No. 2023-506176
[0004] However, only specific limited parameters can be analyzed simultaneously with conventional techniques such as those disclosed in Patent Documents 1 and 2. In order to extract physiologically or pathologically important information from information obtained at the single cell level, it is required to use more information, but since there are many parameters related to cells, it has been difficult from technical and economical standpoints to simultaneously analyze many parameters for a single cell.
[0005] One aspect of the present invention is to provide an information processing method and the like that can estimate parameters that have not been obtained for a single cell from parameters that have been obtained at the single-cell level.
[0006] Therefore, one aspect of the present invention includes the following configuration.
[0007] an information processing method executed by an information processing device, comprising: an acquisition step of acquiring target single-cell data relating to a first item acquired from a target single cell; and an estimation step of generating estimated single-cell data relating to a second item of the target single cell from the target single-cell data using a machine learning model that has learned the association between first specimen single-cell data relating to the first item acquired from a specimen single cell collected from a specimen and second specimen single-cell data relating to a second item different from the first item.A machine learning model creation method comprising: a training data acquisition step of acquiring first specimen single-cell data relating to the first item acquired from a specimen single cell collected from a specimen and second specimen single-cell data relating to the second item different from the first item; and a generation step of generating a machine learning model by performing machine learning using the first specimen single-cell data as an explanatory variable and the second specimen single-cell data as a target variable. An information processing device comprising: an acquisition unit that acquires target single cell data relating to a first item acquired from a target single cell; and an estimation unit that generates estimated single cell data relating to the second item of the target single cell from the target single cell data using a machine learning model that has learned the relationship between first specimen single cell data relating to the first item acquired from a specimen single cell collected from a specimen, and second specimen single cell data relating to a second item different from the first item.
[0008] According to one aspect of the present invention, it is possible to provide an information processing method or the like that can estimate parameters that have not been obtained for a single cell from parameters obtained at the single-cell level.
[0009] FIG. 1 is a functional block diagram showing an example of the configuration of an estimation device according to one embodiment of the present invention. FIG. 2 is a flowchart showing an example of a process flow for estimating estimated single-cell data by an information processing method according to one embodiment of the present invention. FIG. 3 is a flowchart showing an example of a flow for constructing a machine learning model according to one embodiment of the present invention. FIG. 4 is a functional block diagram showing another example of the configuration of an estimation device according to one embodiment of the present invention. FIG. 5 is a diagram showing estimation results of RNA and chromatin expressed in mouse cells estimated by a machine learning model according to an example of the present invention. FIG. 6 is a diagram showing the sequences of primers used in qRT-PCR in mouse cells according to an example of the present invention. FIG. 7 is a diagram showing analysis results of RNA expression levels in mouse cells according to an example of the present invention. FIG. 8 is a diagram showing analysis results of protein expression levels in mouse cells according to an example of the present invention. FIG. 9 is a diagram showing measurement results of specific RNA expression levels in mouse cells according to an example of the present invention. FIG. 10 is a diagram showing the results of a comparison of the accuracy of results when using only actual measured values of each parameter and when using estimated single-cell data in addition to the actual measured values according to an example of the present invention.
[0010] [Embodiment 1] (Information Processing Device 10) The following describes an information processing device 10 according to one embodiment of the present invention. The information processing device 10 is a device that estimates estimated single cell data 23 from target single cell data 22. The information processing device 10 uses a machine learning model 21 created based on various types of data acquired from a large number of single cells, and therefore can simultaneously estimate a large number of parameters in single cells.
[0011] 1 is a functional block diagram showing an example of the configuration of an information processing device 10. The information processing device 10 includes a control unit 1 that controls each unit of the information processing device 10, a storage unit 2 that stores various data used by the information processing device 10, an input unit 11, and an output unit 16, but is not limited to this configuration. For example, the storage unit 2 may be an external device attached to the information processing device 10. The input unit 11 may be an input device such as a laptop computer or tablet connected to the information processing device 10 by wire or wirelessly. The output unit 16 and the output control unit 15 may be a communication unit and a communication control unit that transmit output results to another device or a network.
[0012] The input unit 11 is for accepting various input operations from a user and may be, for example, a keyboard, a mouse, a touch panel, etc. The input unit 11 may be used to input information including at least one of the biological species from which the target single cell or specimen single cell was collected, the type of the first item, the type of the second item, and the target single cell data 22.
[0013] The output unit 16 outputs the estimated single cell data 23. The output form of the output unit 16 is not particularly limited, and may be, for example, a display output, a print output, or an audio output.
[0014] <Controller 1> The controller 1 includes an acquirer 12, an extractor 13, an estimator 14, and an output controller 15. In addition, some of the blocks included in the controller 1 may be omitted from the controller 1 by assigning their functions to another device that can communicate with the information processing device 10.
[0015] The acquisition unit 12 acquires data input from the input unit 11. The acquisition unit 12 may acquire target single cell data 22, at least one type of first item, and at least one type of second item. The acquisition unit 12 may store the acquired data in the storage unit 2. The acquisition unit 12 may perform preprocessing on the acquired data.
[0016] The extraction unit 13 extracts features based on the target single-cell data 22 acquired by the acquisition unit 12. Examples of features include the above-mentioned types and expression levels of proteins; types and expression levels of RNA; types, modifications, and expression levels of DNA; types, modifications, open / closed states, amounts, and structures of chromatin; chromosome structures and interactions; epigenomic information; and binding of transcriptional regulators. Examples of methods for extracting features include a filter method using the ranking of expression levels of each type of protein or RNA, the relative expression levels of each type, and the presence or absence of expression of each type; an embedded method using information contributing to model accuracy obtained through machine learning model creation, and a dimensionality reduction method.
[0017] The estimation unit 14 inputs the target single cell data 22 into the trained machine learning model 21 to estimate estimated single cell data 23. The estimation unit 14 may estimate only one type of candidate for the estimated single cell data 23, or may estimate two or more types. The estimation unit 14 may store the estimated single cell data 23 in the storage unit 2.
[0018] The output control unit 15 causes the output unit 16 to output the estimation result output by the estimation unit 14 .
[0019] <Storage Unit 2> The storage unit 2 may store a machine learning model 21. Furthermore, target single cell data 22 and estimated single cell data 23 may be stored as necessary.
[0020] The machine learning model 21 learns the association between first specimen single-cell data relating to a first item obtained from a specimen single cell collected from a specimen, and second specimen single-cell data relating to a second item different from the first item. This enables the machine learning model 21 to estimate estimated single-cell data 23 relating to the second item from target single-cell data 22 relating to the first item.
[0021] As used herein, the term "specimen" refers to any organism from which a single cell is collected to be used to create the machine learning model 21. The specimen may be a living organism or a dead organism. There are no particular limitations on the organism that can serve as a specimen, as long as it is an organism from which the target single cell and specimen single cell described below can be derived. In other words, the specimen may be any organism, including a plant, an animal, or an insect.
[0022] In most cases, analyzing proteins and the like from a single cell requires destroying the cell. Therefore, the upper limit of the number of parameters that can be simultaneously analyzed for a single cell was approximately three. The information processing device 10 makes it possible to estimate many parameters that have not been obtained for the single cell from parameters obtained at the single-cell level, thereby enabling simultaneous analysis of various parameters.
[0023] Furthermore, conventional methods for analyzing parameters of single cells were only able to analyze specific combinations of parameters. On the other hand, the information processing device 10 can estimate any combination of all parameters that have been machine-learned in the machine learning model 21, not just specific combinations of parameters in a single cell, making it possible to simultaneously perform parameter analysis on combinations of parameters that were previously difficult to perform. (Information Processing Method)
[0024] An information processing method according to one embodiment of the present invention is an information processing method executed by an information processing device 10, and includes an acquisition step of acquiring target single cell data 22 relating to a first item acquired from a target single cell, and an estimation step of generating estimated single cell data 23 relating to a second item of the target single cell from the target single cell data 22 using a machine learning model 21 that has learned the association between first specimen single cell data relating to the first item acquired from a specimen single cell collected from a specimen, and second specimen single cell data relating to a second item different from the first item.
[0025] An information processing method according to one embodiment of the present invention will be described with reference to Fig. 2. Fig. 2 is a flowchart showing an outline of the information processing method according to one embodiment of the present invention.
[0026] In step S1, the acquisition unit 12 acquires target single cell data 22 from a target single cell (acquisition step). In this specification, single cell data means data obtained from a single cell. In other words, single cell data is data about an individual cell. In step S1, target single cell data 22 may be acquired from each of a plurality of single cells.
[0027] Here, the target single cell data 22 relates to a first item. The first item is not particularly limited as long as it can be commonly used as cell data. Examples of the first item include one or more items selected from the group consisting of: protein type and expression level; RNA type and expression level; DNA type, modification, and expression level; chromatin type, modification, open / closed state, amount, and structure; chromosome structure and interaction; epigenomic information; and binding of transcriptional regulators. Among these, the first item is preferably one or more items selected from data related to protein, RNA, DNA, and chromatin, as this makes it easier to estimate the state of the target. More preferably, the first item is one or more items selected from protein type, RNA type, DNA type, and chromatin type.
[0028] In one embodiment, after step S1, the acquisition unit 12 or the extraction unit 13 may perform preprocessing on the target single-cell data 22 related to the first item. Examples of preprocessing include, but are not limited to, averaging, standardization, and outlier removal. By performing preprocessing on the target single-cell data 22, the accuracy of the estimated single-cell data 23 can be improved.
[0029] In step S2, the extraction unit 13 extracts features based on the target single-cell data 22 acquired in step S1. From the viewpoint of improving the accuracy of the data, in step S2, it is preferable to perform further preprocessing on the extracted features, such as reduction of feature dimensions, undersampling, oversampling, averaging, noise reduction, and masking.
[0030] In step S3, the estimation unit 14 inputs the feature amounts extracted in step S2 into the machine learning model 21 to estimate estimated single cell data 23 of the target single cell. At this time, the machine learning model 21 is obtained by machine learning using specimen single cell data obtained for a specimen single cell collected from a specimen.
[0031] Here, the estimated single cell data 23 relates to a second item. As the second item, any data selectable as the first item can be selected as long as it is of a type different from that of the first item. Specifically, the second item may be one or more types selected from the group consisting of: protein type and expression level; RNA type and expression level; DNA type, modification, and expression level; chromatin type, modification, open / closed state, amount, and structure; chromosome structure and interaction; epigenomic information; and binding of transcriptional regulators. Among these, the second item is preferably a parameter related to protein, RNA, DNA, or chromatin, as this makes it easier to estimate the state of the target, and more preferably one or more types selected from the group consisting of protein type, RNA type, DNA type, and chromatin type.
[0032] The information processing method may include step S4, in which the output control unit 15 outputs the estimated single cell data 23. In step S4, the information processing device 10 causes the output device to output the estimated single cell data 23 as single cell data related to the target single cell. By including the output step in the information processing method, the estimated estimated single cell data 23 can be easily confirmed. In step S4, the target single cell data 22 may also be output.
[0033] The target single cell and the specimen single cell may be either a prokaryotic cell or a eukaryotic cell, but are preferably eukaryotic cells. Furthermore, the origin of the target single cell and the specimen single cell is not particularly limited, and they may be, for example, animal cells, plant cells, bacteria, archaea, or protist microbial cells. Among these, animal cells, plant cells, or insect cells are preferred, animal cells or plant cells are more preferred, and animal cells are even more preferred. Examples of animal cells include mammalian cells, avian cells, amphibian cells, fish cells, and insect cells, among which mammalian cells are preferred, and more preferably cells from humans, mice, rats, rabbits, dogs, or non-human primates.
[0034] The target single cell and the specimen single cell may be different cells or the same cell. The target single cell and the specimen single cell are preferably cells of organisms belonging to the same class, more preferably cells of organisms belonging to the same order, even more preferably cells of organisms belonging to the same family, even more preferably cells of organisms belonging to the same genus, and particularly preferably cells of organisms of the same species. With the above configuration, the accuracy of the obtained estimation result is improved.
[0035] Furthermore, the target single cell and the specimen single cell are preferably cells derived from the same tissue. Examples of the target single cell and the specimen single cell include mesenchymal stem cells, ES cells, keratinocytes, fibroblasts, bone marrow cells, endothelial cells, smooth muscle cells, Schwann cells, chondrocytes, adipocytes, osteoblasts, and vascular endothelial progenitor cells. If the single cell and the specimen single cell are derived from the same tissue, the accuracy of the obtained estimation result is improved.
[0036] In most cases, analyzing proteins and the like from single cells requires destroying the cells before measurement. Therefore, the upper limit of the number of parameters that can be measured simultaneously for a single cell is approximately three. According to an information processing method according to one embodiment of the present invention, a machine learning model 21 is created based on various types of data acquired from a large number of single cells, making it possible to simultaneously estimate a large number of parameters in a single cell.
[0037] Furthermore, while conventional methods for analyzing parameters of single cells can only measure specific combinations of parameters, the information processing method according to one embodiment of the present invention can estimate any combination of all parameters that have been machine-learned by the machine learning model 21, not just specific combinations of parameters in single cells.
[0038] <Method for Creating Machine Learning Model 21> A method for creating a machine learning model 21 according to one embodiment of the present invention includes a learning data acquisition step of acquiring first specimen single-cell data regarding a first item obtained for a specimen single cell collected from a specimen, and second specimen single-cell data regarding a second item different from the first item, and a generation step of generating the machine learning model 21 by performing machine learning using the first specimen single-cell data as an explanatory variable and the second specimen single-cell data as a target variable.
[0039] A method for creating the machine learning model 21 will be described with reference to Fig. 3. Fig. 3 is a flowchart showing an example of a method for creating the machine learning model 21.
[0040] Step S11 is a learning data acquisition step of acquiring first specimen single cell data and second specimen single cell data related to a second item for specimen single cells obtained from a specimen of a particular status.
[0041] In step S12, machine learning is performed using the first sample single-cell data as an explanatory variable and the second sample single-cell data as a response variable. Step S13 may be performed by a machine learning device including a learning unit for performing the machine learning.
[0042] The machine learning method in step S12 may be transfer learning such as HeMap, or may be deep learning such as Conditional Variational AutoEncoder (CVAE), a combination of CVAE and Generative Adversarial Networks (GAN), or a diffusion model. When deep learning or the like is used as the machine learning method, the first sample single-cell data and the second sample single-cell data are labeled with the sample status in advance. When the first and second single-cell data labeled with the sample status are used, the machine learning model 21 may perform machine learning for each sample status using the first sample single-cell data as an explanatory variable and the second sample single-cell data as a target variable. From the viewpoint of improving the accuracy of the obtained estimated single-cell data 23, the machine learning method is preferably transfer learning.
[0043] As used herein, the term "specimen status" refers to information related to the organism that serves as the specimen. Examples of such information include status information, information indicating the species of the specimen, and information on the type of a single cell in the specimen.
[0044] In one embodiment, the specimen status may be information indicating symptoms, diseases, conditions, etc. of the organism serving as the specimen. Specifically, if the specimen is an animal, the specimen status may include, for example, the presence or absence of infection, inflammation, bleeding, and disease. More specific specimen status may be, for example, information simply indicating a symptom, such as "inflammation," information on the site where the symptom occurs, or information on the specific disease name of the specimen. Furthermore, if the specimen is a plant, the specimen status may include, for example, the degree of dryness, temperature, salt concentration, and the presence or absence of disease in the growing environment. More specific specimen status may be, for example, information simply indicating "dry," information on the site where changes due to the dry state occur, or information on the specific growing environment to which the specimen is exposed.
[0045] Examples of inflammation include, but are not limited to, dermatitis (e.g., allergic dermatitis and atopic dermatitis), arthritis, vasculitis, glomerulonephritis, gastroenteritis, pneumonia, meningitis, obesity, and aging. Examples of bleeding include, but are not limited to, bleeding in the digestive organs and bleeding in other organs. Examples of diseases include, but are not limited to, autoimmune diseases, neurological and psychiatric diseases (e.g., Alzheimer's disease, stroke, Parkinson's disease, and schizophrenia), muscle degenerative diseases (e.g., muscular dystrophy), arteriosclerosis, and cancer.
[0046] The specimen status may include only one type of information, or may include multiple types of information.
[0047] The machine learning model 21 is preferably created by transfer learning of the association between the first specimen single-cell data and the second specimen single-cell data. When the machine learning model 21 is obtained by transfer learning, a latent space is generated by fusing the single-cell data from the common information potentially present among the measurement data of each parameter of multiple types of specimen single cells obtained from the specimen. With the above configuration, multiple types of data obtained from specimen single cells having a specific status are fused, making it possible to detect elements that are difficult to detect using data obtained by measuring one type of single cell.
[0048] Before step S13, preprocessing such as reduction of feature dimensions, undersampling, oversampling, averaging, noise reduction, and masking may be performed on the first sample single-cell data and the second sample single-cell data used to create the machine learning model 21. The preprocessing can be performed by any component including a learning unit, as long as it is performed before step S14.
[0049] In step S13, the machine learning model 21 is created using the data that was machine-learned in step S12. In step S14, the machine learning model 21 is stored, and the processing in Fig. 3 ends. The machine learning model 21 may be stored in a location that can be read by an information processing device that can execute an information processing method according to an embodiment of the present invention. The readable location may be, for example, the storage unit 2 of the information processing device 10.
[0050] In one embodiment, the machine learning model 21 is created by the following method. The following method is an example in which transfer learning using HeMap is used as the machine learning method and a mouse is used as the specimen organism. The method for creating the machine learning model 21 in the present invention is not limited to the following. 1. Collect multiple cells from a mouse. 2. Divide the collected cells into single cells and analyze the gene type, chromatin type, or protein type for each single cell to obtain training data. 3. Perform transfer learning using HeMap using the acquired training data to create the machine learning model 21, which is a latent space. 4. Create a model that estimates the feature quantities of each measurement data based on the machine learning model 21.
[0051] [Embodiment 2] (Information Processing Device 10a) An overview of an information processing device 10a according to another embodiment of the present invention will be described with reference to Fig. 4. Note that descriptions of matters that have already been described will be omitted.
[0052] The information processing device 10a includes a control unit 1a that controls each unit of the information processing device 10a, a storage unit 2a that stores various data used by the information processing device 10a, an input unit 11, and an output unit 16. The control unit 1a includes an acquisition unit 12a, an estimation unit 14a, and an accuracy verification unit 17. The storage unit 2a may store a machine learning model 21a, a third estimated data group 24 related to a third item, and third target data 25 related to the third item.
[0053] Here, the third item may be data that can be selected as the first item and the second item, or data that cannot be selected as the first item and the second item, as long as the third item is data that is different from the first item and the second item. In one embodiment, the third item may be the status of the subject and the specimen.
[0054] As used herein, "subject" refers to any organism from which a target single cell has been collected. In other words, a target single cell is a single cell collected from a "subject." The target may be the same organism as the specimen, or a different organism (e.g., a closely related species). Organisms that can be selected as targets include organisms of the same type as the specimen described above. Furthermore, as used herein, "subject status" refers to information related to the subject, and is the same type of information as the "specimen status" described above. In other words, if the specimen status is information indicating whether or not the specimen is in an inflammatory state, the subject status is information indicating whether or not the subject is in the inflammatory state.
[0055] <Control Unit 1a> The acquisition unit 12a may acquire third target data 25 in addition to the target single cell data 22. The estimation unit 14a estimates third estimated data in addition to the estimated single cell data 23. The estimation unit 14a inputs the estimated single cell data 23 into the machine learning model 21a to estimate a third estimated data group 24. The estimation unit 14a may store the third estimated data in the storage unit 2a. In one embodiment, when there are multiple machine learning models 21a, the third estimated data may be a third estimated data group 24 including multiple pieces of data. In other words, when there are multiple pieces of third estimated data, the third estimated data is referred to as the third estimated data group 24. When there is one machine learning model 21a, the third estimated data group 24 includes one piece of third estimated data.
[0056] In one embodiment, when a method such as HeMap is used as a machine learning method, the control unit 1a may be provided with an accuracy verification unit 17 that verifies the accuracy of the estimated single cell data 23 output from the machine learning model 21a.
[0057] The accuracy verification unit 17 may use the second to fifth machine learning models described in (1) to (4) below to verify the accuracy of the estimated single-cell data 23. When the accuracy verification unit 17 uses data on a third item, the third item is preferably data common to the specimen and the subject, and more preferably data that is useful or of interest as a measurement subject.
[0058] The second to fifth machine learning models may be included in the information processing device 10a as the machine learning model 21a. In an embodiment, the second to fifth machine learning models may be included in the estimation unit 14a. Alternatively, the second to fifth machine learning models may be included in a device different from the information processing device 10a. (1) A second machine learning model trained by machine learning using the first sample single-cell data as an explanatory variable and the third sample data related to the third item as a dependent variable. (2) A third machine learning model trained by machine learning using the second sample single-cell data as an explanatory variable and the third sample data related to the third item as a dependent variable. (3) A fourth machine learning model trained by machine learning using the second sample estimated single-cell data, estimated by inputting the first sample single-cell data and the first sample single-cell data into the machine learning model, as an explanatory variable and the third sample data related to the third item as a dependent variable. (4) A fifth machine learning model trained by machine learning using data obtained by randomly combining the first sample single-cell data and the second sample single-cell data as an explanatory variable and the third sample data related to the third item as a dependent variable.
[0059] First, the accuracy verification unit 17 performs the following steps (1') to (4') to generate a third estimated data group 24. The third estimated data group 24 includes first to fourth target estimated data. (1') First target estimated data obtained by inputting target single-cell data 22 for the first item into a second machine learning model. (2') Second target estimated data obtained by inputting target single-cell data measured for the second item into a third machine learning model. (3') Third target estimated data obtained by inputting data obtained by fusing the target single-cell data 22 and estimated single-cell data 23 in a predetermined format into a fourth machine learning model. (4') Fourth target estimated data obtained by inputting data obtained by randomly fusing the target single-cell data 22 for the first item and target single-cell data 23 measured for the second item into a fifth machine learning model.
[0060] Furthermore, the accuracy verification unit 17 compares each of the first to fourth target estimation data with the third target data 25, and calculates the matching rate between each of the first to fourth target estimation data and the third target data 25. Hereinafter, the calculated matching rates will be referred to as the first result accuracy, the second result accuracy, the third result accuracy, and the fourth result accuracy, respectively.
[0061] The accuracy verification unit 17 verifies the accuracy of the estimated single-cell data 23 based on the match rate calculated from each estimated data. <Case 1> For example, if the first result accuracy is less than the second result accuracy, the third result accuracy is greater than the fourth result accuracy, and the third result accuracy is greater than the first result accuracy, the accuracy verification unit 17 determines that the accuracy of the estimated single-cell data 23 is sufficiently high. The accuracy verification unit 17 then outputs the estimated single-cell data 23 to the output control unit 15. <Case 2> For example, if the first result accuracy is greater than the second result accuracy, the third result accuracy is greater than the fourth result accuracy, and the third result accuracy is greater than the second result accuracy, the accuracy verification unit 17 determines that the accuracy of the estimated single-cell data 23 is sufficiently high. The accuracy verification unit 17 then outputs the estimated single-cell data 23 to the output control unit 15. In Cases 1 and 2 above, the accuracy verification unit 17 may label the estimated single-cell data 23 to indicate that it has high accuracy. <Case 3> For example, if the third result accuracy is smaller than the fourth result accuracy, the accuracy verification unit 17 determines that the accuracy of the estimated single cell data 23 is low. In this case, the accuracy verification unit 17 may be configured not to output the estimated single cell data 23 to the output control unit 15. The accuracy verification unit 17 may label the estimated single cell data 23 that has been determined to have low accuracy, indicating that the data has low accuracy, and then output the data to the output control unit 15.
[0062] <Memory Unit 2a> In one embodiment, the machine learning model 21a may have learned associations between the first sample single-cell data, the second sample single-cell data, and third sample data regarding a third item different from both the first item and the second item, and in the estimation step, may further generate a third estimated data group 24 regarding the third item from the target single-cell data 22 and the estimated single-cell data 23. As another example, the machine learning model 21a may have learned associations between the first sample single-cell data, the second sample estimated single-cell data, and the third sample data.
[0063] The third estimated data group 24 may include first to fourth target estimated data output from the second to fifth machine learning models, respectively.
[0064] The third target data 25 is data relating to the third item of interest. The third target data 25 may be acquired by the acquisition unit 12a.
[0065] [Example of implementation by software] The functions of an estimation device (hereinafter referred to as "device") according to one embodiment of the present invention can be realized by a program that causes a computer to function as the device, and a program that causes a computer to function as each control block of the device (particularly each part included in the control unit 1).
[0066] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. The functions described in each of the above embodiments are realized by executing the program using the control device and storage device.
[0067] The program may be non-transitory and may be recorded on one or more computer-readable recording media. The recording media may or may not be included in the device. In the latter case, the program may be supplied to the device via any wired or wireless transmission medium.
[0068] Furthermore, some or all of the functions of the control blocks can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the control blocks is formed is also included in the scope of the present invention. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.
[0069] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI may run on the control device or on another device (for example, an edge computer or a cloud server).
[0070] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included within the technical scope of the present invention. [Summary] One aspect of the present invention includes the following configuration. <1> An information processing method executed by an information processing device, comprising: an acquisition step of acquiring target single-cell data related to a first item acquired from a target single cell; and an estimation step of generating estimated single-cell data related to the second item of the target single cell from the target single-cell data using a machine learning model that has learned an association between first specimen single-cell data related to the first item acquired from a specimen single cell collected from a specimen, and second specimen single-cell data related to a second item different from the first item. <2> The information processing method according to <1>, further comprising: an output step of causing an output device to output the target single-cell data and the estimated single-cell data as single-cell data related to the target single cell. <3> The information processing method described in <1> or <2>, wherein the first item is one or more types of data selected from the group consisting of: type and expression level of protein; type and expression level of RNA; type, modification and expression level of DNA; type, modification, opening / closing state, amount and structure of chromatin; structure and interaction of chromosome; epigenetic information; and binding of transcriptional regulatory factors; and the second item is data different from the first item and is one or more types of data selected from the group consisting of type and expression level of protein; type and expression level of RNA; type, modification and expression level of DNA; type, modification, opening / closing state, amount and structure of chromatin; structure and interaction of chromosome; epigenetic information; and binding of transcriptional regulatory factors.<4> The information processing method according to any one of <1> to <3>, wherein the machine learning model has further learned associations between the first specimen single-cell data, the second specimen single-cell data, and third specimen data relating to a third item different from both the first item and the second item, and wherein the estimation step further generates third estimated data relating to the third item from the target single-cell data and the estimated single-cell data. <5> The information processing method according to any one of <1> to <4>, wherein, after the acquisition step, one or more types of processing selected from the group consisting of averaging, standardization, and outlier removal are performed on the target single-cell data relating to the first item. <6> The information processing method according to any one of <1> to <5>, wherein the machine learning model is created by transfer learning of associations between the first specimen single-cell data and the second specimen single-cell data. <7> The information processing method according to any one of <1> to <6>, wherein the target single cell and the specimen single cell are cells of the same biological species. <8> The information processing method according to any one of <1> to <7>, wherein the target single cell and the specimen single cell are cells derived from the same tissue. <9> A method for creating a machine learning model, comprising: a learning data acquisition step of acquiring first specimen single-cell data related to a first item obtained from a specimen single cell collected from a specimen, and second specimen single-cell data related to a second item different from the first item, and a generation step of generating a machine learning model by performing machine learning using the first specimen single-cell data as an explanatory variable and the second specimen single-cell data as a target variable. <10> An information processing device comprising: an acquisition unit that acquires target single-cell data related to the first item obtained from a target single cell; and an estimation unit that generates estimated single-cell data related to the second item of the target single cell from the target single-cell data, using a machine learning model that has trained an association between the first specimen single-cell data related to the first item obtained from a specimen single cell collected from a specimen, and second specimen single-cell data related to a second item different from the first item.<11> A control program for causing a computer to function as the information processing device according to <10>, the control program causing the computer to function as the acquisition unit and the estimation unit. <12> A computer-readable recording medium having the control program according to <11> recorded thereon.
[0071] An embodiment of the present invention will now be described.
[0072] [Materials] <Mice> Hlf-tdTomato mice aged 7 to 13 weeks were used for the experiments. 100 μg of lipopolysaccharide (LPS) was intraperitoneally injected 16 hours before flow cytometry or RT-PCR. Mice that received an intraperitoneal injection of 100 μm of PBS were also used as controls.
[0073] [Creation of machine learning model] Two to three Hlf-tdTomato mice injected with LPS or PBS were used as specimens for each experiment, and two experiments were performed. RFP-positive hematopoietic stem cells (tdTomato) were isolated from each mouse. + CD150 + CD48 - Single cells of a mouse (LSK) were obtained and the types of proteins, RNAs, and chromatins were measured. The number of measurement samples was 1,514 proteins, 63 RNAs, and 288 chromatins, respectively. The measured data was used as the first or second specimen single-cell data, and machine learning was performed on the relationships between these measurement data to obtain a machine learning model. HeMap, a transfer learning method, was used for machine learning. The five proteins with the highest expression levels as a result of the actual measurement were input into the obtained machine learning model as the target single-cell data for the first item, and the RNA types and chromatin types were estimated as the estimated single-cell data for the second item. The results are shown in Figure 5.
[0074] [Analysis] To confirm the accuracy of the results of the RNAs predicted by the machine learning model, the expression levels of the predicted RNAs shown in Figure 5 and the proteins encoded by the RNAs were analyzed.
[0075] Each measurement was performed three times, and the results are expressed as the mean ± SEM of two independent experiments. Statistical analysis of the statistical data was performed by Student's t-test.
[0076] [RNA Expression Analysis] <Analysis Method> Analysis of RNA expression levels was performed by qRT-PCR. RNAs that appeared twice or more in the table in Figure 7 were selected for analysis. In addition, Mpeg2 / Il2rg was also analyzed because it is known to be closely related to inflammation. RFP-positive hematopoietic stem cells (tdTomato) were isolated from Hlf-tdTomato mice that had been intraperitoneally injected with LPS or PBS. + CD150 + CD48 - LSK) was added to a 96-well plate at a density of 100 cells per well. After adding 10 μL of cell lysis buffer to each well, the plate was immediately frozen (-80°C) and analyzed according to the RamDA-seq™ method (Hayashi, T. et al. Nature Communications 9, 619 (2018)). First-strand cDNA was synthesized using a Primer Script RT Reagent Kit (Takara Bio Inc.) and an NSR (not-so-random) primer. Second-strand cDNA was synthesized using Klenow Fragments (30-50 U / μL, New England Biolabs) and the complementary strand of the NSR primer. Quantitative PCR analysis was performed using SYBR Green (Applied Biosystems) on a Real-time PCR LightCycler 96 (Roche). The relative expression level of mRNA for each gene was normalized to the expression level of the Gapdh gene. Primers for detecting each gene are shown in Figure 6.
[0077] <Results> The analysis results are shown in Figure 7. In Figure 7, the vertical axis indicates the relative expression level of each gene. Sca1 is a positive control. Kit is a negative control. As shown in Figure 9, the expression levels of Lag3, IL2rg, Lfitm3, and Mpeg1 were increased. The above results demonstrate that the information processing method according to one embodiment of the present invention makes it possible to accurately estimate a large number of parameters related to a single cell.
[0078] [Protein Expression Analysis] <Analysis Method> Protein expression levels were analyzed for Lag3 and IL2rg, which showed increased expression levels in RNA expression level analysis. Protein expression level analysis was performed using flow cytometry. Femurs and tibias were collected from Hlf-tdTomato mice intraperitoneally injected with LPS or PBS. Bone marrow cells were extracted from the obtained bones using a syringe and suspended in FACS buffer (2% FBS and 2 mM EDTA in PBS). The suspension was centrifuged at 470G for 5 minutes, the supernatant was removed, and the cells were suspended in ACK lysis buffer to remove red blood cells from the bone marrow cells. The suspension was centrifuged at 470G for 5 minutes, the supernatant was removed, and the cells were resuspended in FACS buffer. Lineage (Lin) marker was used to label the bone marrow cells. Specifically, bone marrow cells were labeled using biotinylated antibodies, including anti-TER-119 antibody, anti-NK1.1 antibody, anti-CD11b (M1 / 70) antibody, anti-Gr-1 (RB6-8C5) antibody, anti-CD3e (145-2C11) antibody, anti-CD4 (GK1.5) antibody, anti-CD8a (53-6.7) antibody, anti-B220 (RA3-6B2) antibody, and anti-IL7RA (SB / 199) antibody (all manufactured by BioLegend). Lin marker-positive cell fractions were excluded using MACS streptavidin magnetic beads (manufactured by Miltenyi Biotec).
[0079] Fractions negative for the Lin marker were fluorescently labeled with antibodies: anti-c-Kit (2B8), anti-Sca-1 (D7), anti-CD48 (HM45-1), anti-CD150 (TC15-12F12.2), anti-Lag3 (C9B7W), and anti-IL2rg (TUGm2) (all manufactured by BioLegend). Hematopoietic stem cells (Lin) were then isolated from the fluorescently labeled cell population. - c-Kit + Sca-1 + CD48 ― CD150 + tdTomato + ) were separated using a FACS Aria III and analyzed using FlowJo (BD Biosciences). A separate population of fluorescently labeled cells was analyzed using a FACSSymphony™ without cell separation.
[0080] <Results> The results of protein analysis are shown in Figure 8. Figure 8 is a graph showing the fluorescence intensity of each marker in flow cytometry, and a graph showing the relative expression level of protein in each mouse. The graph showing fluorescence intensity in Figure 8 shows the results of measurements performed on two mice injected with PBS or LPS. Figure 8 shows that both Lag3 and Il2rg, whose gene expression levels were increased, were higher in mice injected with LPS than in mice injected with PBS. From the above, it can be seen that the expression level of RNA estimated by the information processing method according to one embodiment of the present invention is also increased at the protein level.
[0081] [RNA data] <Measurement method> RFP-positive hematopoietic stem cells (Lin - c-Kit + Sca-1 + CD48 ― CD150 + tdTomato +) were harvested and placed in a 96-well plate at a density of one cell per well. After adding 1 μL of cell lysis buffer to each well, the wells were immediately frozen (-80°C). RNA analysis was performed according to the RamDA-seq™ method (Hayashi, T. et al. Nature Communications 9, 619 (2018)). The quantity and quality of the isolated cDNA library were assessed using Multina (Shimadzu Corporation, Japan) and Bioanalyzer 2100 (Agilent, Waldbronn, Germany), followed by sequencing using the Next-Seq system (Illumina Inc.). Sequencing data were trimmed using the trim_galore (version 0.4.3) package (quality and length parameters: quality 30, length 30), and then filtered using SortMeRNA to remove mitochondrial mRNA. The filtered data were aligned to the mouse reference sequence (GRCm38, M17 version) using the RNA-seq aligner STAR (version 2.5.3a; Dobin et al., 2012). From the aligned BAM files, along with the mouse GFF annotation file (GRCm38, M17 version), the number of reads for each gene and sample was calculated using the featureCounts function in the Subreads program (version 1.6.1) to create a gene count matrix. Differences in expressed genes between the LPS and PBS samples were calculated using R (version 3.4) and the DESeq2 (version 1.20.0) package. Results: Figure 9 shows the results for RNA alone. As shown in Figure 9, Il2rg and Ifitm3 were expressed in both PBS- and LPS-injected mice, with increased expression levels in LPS-injected mice. On the other hand, Lag3 and Gpb4 were only expressed in LPS-injected mice.
[0082] In the data for RNA alone, Lag3 and Il2rg were ranked low. On the other hand, the data shown in Figures 7 and 8 show that the expression levels of Lag3 and Il2rg in mice injected with LPS are relatively increased compared to mice injected with PBS. This suggests that the information processing method according to one embodiment of the present invention may be capable of estimating factors that are difficult to detect from single data.
[0083] [Verification of Estimated Single Data] <Creation of Second Machine Learning Model> A second machine learning model was created by machine learning using the measurement data used when creating the machine learning model. The machine learning was performed using a combination of four types of explanatory variables and a target variable, and the second machine learning model was obtained. (1) Explanatory variable: data on 1514 measured proteins, Objective variable: PBS or LPS injection of the specimen mice (2) Explanatory variable: data on 63 measured RNAs, Objective variable: PBS or LPS injection of the specimen mice (3) Explanatory variable: data obtained by fusing the data on 1514 measured proteins with the data on 63 measured RNAs estimated by the machine learning model, Objective variable: PBS or LPS injection of the specimen mice (4) Explanatory variable: data obtained by randomly fusing the data on 1514 measured proteins with the data on 63 measured RNAs, Objective variable: PBS or LPS injection of the specimen mice In the second machine learning model, (1) was input with the protein measurement results, which were the target single cell data, (2) was input with the estimated RNA data, which were the estimated single cell data obtained by the machine learning model, and (3) and (4) were input with the fusion data of the target single cell data and the estimated single cell data, and a determination result was obtained as to whether the mice were injected with LPS or PBS. The obtained discrimination results were compared with the actual results of PBS injection or LPS injection for the mice from which the target single cell data was obtained, to verify the accuracy of the estimated single cell data.
[0084] <Results> The results are shown in Figure 10. In Figure 10, the measurement data is the result of inputting only the target single-cell data and determining whether it was a PBS injection or an LPS injection using the machine learning model (1) or (2), respectively. On the other hand, the estimated single-source data is the result of inputting, in addition to the target single-cell data, estimated single-source data estimated by the machine learning model and determining whether it was a PBS injection or an LPS injection using the machine learning model (3) or (4), respectively. Note that in the estimated single-source data, only the type of protein was used as the target single-cell data.
[0085] As shown in Figure 10, when estimated data on RNA type was used in addition to protein type, the accuracy of the determination was improved compared to when only measured data on protein type was used. Furthermore, the accuracy of the determination was similar to when only measured data on RNA type was used. Furthermore, when estimated data on chromatin type was used in addition, the accuracy of the determination was improved compared to when only measured data on RNA type was used. From the above, it was demonstrated that the estimated single-source data can estimate data with high accuracy.
[0086] The present invention can be used to estimate multiple parameters of a single cell, as well as to estimate the state of a subject.
[0087] 1, 1a, 1b Control unit 2, 2a, 2b Storage unit 10, 10a, 10b Information processing device 11 Input unit 12, 12a Acquisition unit 13 Extraction unit 14, 14a Estimation unit 15 Output control unit 16 Output unit 17 Accuracy verification unit 21, 21a Machine learning model 22 Target single cell data 23 Estimated single cell data 24 Third estimated data group 25 Third target data
Claims
1. An information processing method executed by an information processing device, comprising: an acquisition step of acquiring target single cell data relating to a first item acquired from a target single cell; and an estimation step of generating estimated single cell data relating to the second item of the target single cell from the target single cell data using a machine learning model that has learned the association between first specimen single cell data relating to the first item acquired from a specimen single cell collected from a specimen, and second specimen single cell data relating to a second item different from the first item.
2. The information processing method according to claim 1, further comprising an output step of causing an output device to output the target single cell data and the estimated single cell data as single cell data relating to the target single cell.
3. The information processing method of claim 1, wherein the first item is one or more types of data selected from the group consisting of: type and expression level of protein; type and expression level of RNA; type, modification and expression level of DNA; type, modification, open / closed state, amount and structure of chromatin; structure and interaction of chromosome; epigenetic information; and binding of transcriptional regulatory factors; and the second item is data different from the first item and is one or more types of data selected from the group consisting of: type and expression level of protein; type and expression level of RNA; type, modification and expression level of DNA; type, modification, open / closed state, amount and structure of chromatin; structure and interaction of chromosome; epigenetic information; and binding of transcriptional regulatory factors.
4. The method of claim 1, wherein the machine learning model has further learned associations between the first sample single-cell data, the second sample single-cell data, and third sample data regarding a third item different from both the first item and the second item, and wherein the estimation step further generates third estimated data regarding the third item from the target single-cell data and the estimated single-cell data.
5. The information processing method according to claim 1, wherein after the acquisition step, one or more processes selected from the group consisting of averaging, standardization, and outlier removal are performed on the target single-cell data relating to the first item.
6. The information processing method according to claim 1, wherein the machine learning model is created by transfer learning of the association between the first sample single-cell data and the second sample single-cell data.
7. The information processing method according to claim 1, wherein the target single cell and the specimen single cell are cells of the same biological species.
8. The information processing method according to claim 1, wherein the target single cell and the specimen single cell are cells derived from the same tissue.
9. A method for creating a machine learning model, comprising: a learning data acquisition step of acquiring first specimen single-cell data regarding a first item obtained from a specimen single cell collected from a specimen, and second specimen single-cell data regarding a second item different from the first item; and a generation step of generating a machine learning model by performing machine learning using the first specimen single-cell data as an explanatory variable and the second specimen single-cell data as a target variable.
10. An information processing device comprising: an acquisition unit that acquires target single cell data relating to a first item acquired from a target single cell; and an estimation unit that generates estimated single cell data relating to the second item of the target single cell from the target single cell data using a machine learning model that has learned the association between first specimen single cell data relating to the first item acquired from a specimen single cell collected from a specimen, and second specimen single cell data relating to a second item different from the first item.
11. A control program for causing a computer to function as the information processing device according to claim 10, the control program causing the computer to function as the acquisition unit and the estimation unit.
12. A computer-readable recording medium on which the control program according to claim 11 is recorded.
Citation Information
Patent Citations
Gene information estimation apparatus using topic model
JP2018139043A