Information processing method, information processing apparatus, method for creating machine learning model

The method uses a machine learning model to estimate additional parameters from single-cell data, addressing the limitations of conventional analysis by enabling simultaneous estimation of multiple parameters, thus improving cellular state understanding.

JP2025124502APending Publication Date: 2025-08-26NAT UNIV CORP KUMAMOTO UNIV +1

Patent Information

Application Number
JP2024020605
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-14
Publication Date
2025-08-26

Smart Images

  • Figure 2025124502000001_ABST
    Figure 2025124502000001_ABST
Patent Text Reader

Abstract

To provide an information processing method and the like with which a parameter not acquired for a single cell can be estimated from a parameter acquired at a single-cell level.SOLUTION: An information processing method includes: an acquisition step of acquiring target single cell data 22 related to a first item from a target cell; and an estimation step of estimating estimated single cell data 23 related to a second item of the target cell from the target single cell data 22 by using a machine learning model 21.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing method, an information processing device, and a method for creating a machine learning model. [Background technology]

[0002] Technologies for analyzing various cellular parameters at the single-cell level (e.g., single-cell RNA-seq, ATAC (Assay for Transposase-Accessible Chromatin)-seq, DNA methylation markers, chromosome conformation capture (Hi-C) method, mass cytometry analysis, etc.) have been developed. In recent years, technologies for simultaneously analyzing two or three of these parameters at the single-cell level have been disclosed. For example, Patent Document 1 discloses an image-based method for classifying human embryonic cells. Furthermore, Patent Document 2 discloses a method for barcoding multiple targets in cells using a barcode composition. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2022-87297 [Patent Document 2] Japanese Patent Application Publication No. 2023-506176 Summary of the Invention [Problem to be solved by the invention]

[0004] However, only specific limited parameters can be analyzed simultaneously with conventional technologies such as those disclosed in Patent Documents 1 and 2. In order to extract physiologically or pathologically important information from information obtained at the single cell level, it is necessary to use more information. However, since there are many parameters related to cells, it has been technically and economically difficult to analyze many parameters for a single cell simultaneously.

[0005] One aspect of the present invention is to provide an information processing method and the like that can estimate parameters that have not been obtained for a single cell from parameters that have been obtained at the single-cell level. [Means for solving the problem]

[0006] Therefore, one aspect of the present invention includes the following configuration.

[0007] <1> An information processing method executed by an information processing device, an acquisition step of acquiring target single cell data relating to a first item acquired from the target single cell; an estimation step of generating estimated single cell data on the second item of the target single cell from the target single cell data using a machine learning model that has learned the association between first specimen single cell data on the first item obtained from a specimen single cell collected from the specimen and second specimen single cell data on a second item different from the first item; An information processing method comprising: <2> an output step of outputting the target single cell data and the estimated single cell data to an output device as single cell data regarding the target single cell; <1> The information processing method described in <3> the first item is one or more types of data selected from the group consisting of: types and expression levels of proteins; types and expression levels of RNA; types, modifications, open / closed states, amounts, and structures of chromatin; structures and interactions of chromosomes; epigenome information; and binding of transcriptional regulatory factors; The second item is data different from the first item and is one or more types of data selected from the group consisting of: protein type and expression level; RNA type and expression level; DNA type, modification and expression level; chromatin type, modification, open / closed state, amount and structure; chromosome structure and interaction; epigenome information; and binding of transcriptional regulatory factors. <1> or <2> The information processing method described in <4> The machine learning model further comprises: The association between the first sample single-cell data, the second sample single-cell data, and third sample data relating to a third item different from both the first item and the second item has been learned; In the estimation step, third estimated data regarding a third item is further generated from the target single-cell data and the estimated single-cell data. <1> ~ <3> A method according to any one of the preceding claims. <5> After the acquisition step, one or more processes selected from the group consisting of an averaging process, a standardization process, and an outlier removal process are performed on the target single-cell data related to the first item. <1> ~ <4> 10. The information processing method according to claim 9, wherein <6> The machine learning model is created by transfer learning of the association between the first sample single-cell data and the second sample single-cell data. <1> ~ <5> 10. The information processing method according to claim 9, wherein <7> The target single cell and the specimen single cell are cells of the same biological species. <1> ~ <6> 10. The information processing method according to claim 9, wherein <8> The target single cell and the specimen single cell are cells derived from the same tissue. <1> ~ <7> 10. The information processing method according to claim 9, wherein <9> a learning data acquisition step of acquiring first specimen single cell data relating to a first item acquired from a specimen single cell collected from the specimen, and second specimen single cell data relating to a second item different from the first item; generating a machine learning model by performing machine learning using the first sample single-cell data as an explanatory variable and the second sample single-cell data as a response variable, How to create a machine learning model. <10> an acquisition unit that acquires target single cell data relating to a first item acquired from the target single cell; an estimation unit that generates estimated single cell data regarding the second item of the target single cell from the target single cell data using a machine learning model that has learned the association between first specimen single cell data regarding the first item obtained from a specimen single cell collected from a specimen and second specimen single cell data regarding a second item different from the first item; An information processing device comprising: <11> <10> 2. A control program for causing a computer to function as the information processing device according to claim 1, wherein the control program causes the computer to function as the acquisition unit and the estimation unit. <12> <11> A computer-readable recording medium having the control program described above recorded thereon. [Effects of the Invention]

[0008] According to one aspect of the present invention, it is possible to provide an information processing method or the like that can estimate parameters that have not been acquired for a single cell from parameters acquired at the single-cell level. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a functional block diagram showing an example of the configuration of an estimation device according to an embodiment of the present invention. [Figure 2] 10 is a flowchart showing an example of a processing flow for estimating estimated single-cell data using an information processing method according to one embodiment of the present invention. [Figure 3] 1 is a flowchart illustrating an example of a flow for constructing a machine learning model according to an embodiment of the present invention. [Figure 4] FIG. 2 is a functional block diagram showing another example of the configuration of an estimation device according to an embodiment of the present invention. [Figure 5] FIG. 1 shows the results of estimation of RNA and chromatin expressed in mouse cells estimated by a machine learning model according to an embodiment of the present invention. [Figure 6] FIG. 1 shows the sequences of primers used in qRT-PCR in mouse cells according to an example of the present invention. [Figure 7] FIG. 1 shows the results of analyzing RNA expression levels in mouse cells according to an example of the present invention. [Figure 8] FIG. 1 shows the analysis results of protein expression levels in mouse cells according to an example of the present invention. [Figure 9] FIG. 1 shows the results of measuring the expression level of a particular RNA in mouse cells according to an example of the present invention. [Figure 10] FIG. 10 is a diagram showing the results of a comparison of the accuracy of the results when using only the measured values ​​of each parameter according to an embodiment of the present invention and when using estimated single-cell data in addition to the measured values. DETAILED DESCRIPTION OF THE INVENTION

[0010] [Embodiment 1] (Information processing device) The following describes an information processing device 10 according to one embodiment of the present invention. The information processing device 10 is a device that estimates estimated single-cell data 23 from target single-cell data 22. The information processing device 10 uses a machine learning model created based on various types of data acquired from a large number of single cells, and therefore can simultaneously estimate a large number of parameters in single cells.

[0011] 1 is a functional block diagram showing an example of the configuration of an information processing device 10. The information processing device 10 includes a control unit 1 that controls each unit of the information processing device 10, a storage unit 2 that stores various data used by the information processing device 10, an input unit 11, and an output unit 16, but is not limited to this configuration. For example, the storage unit 2 may be an external device attached to the information processing device 10. The input unit 11 may be an input device such as a laptop computer or tablet connected to the information processing device 10 by wire or wirelessly. The output unit 16 and the output control unit 15 may be a communication unit and a communication control unit that transmit the output result to another device or a network.

[0012] The input unit 11 is for accepting various input operations from a user and may be, for example, a keyboard, a mouse, a touch panel, etc. The input unit 11 may be used to input information including at least one of the biological species from which the target single cell or specimen single cell was collected, the type of the first item, the type of the second item, and the target single-cell data.

[0013] The output unit 16 outputs the estimated single cell data. The output form of the output unit 16 is not particularly limited, and may be, for example, a display output, a print output, or an audio output.

[0014] <Control unit 1> The control unit 1 includes an acquisition unit 12, an extraction unit 13, an estimation unit 14, and an output control unit 15. In addition, some of the blocks included in the control unit 1 may be omitted from the control unit 1 by assigning their functions to another device that can communicate with the information processing device 10.

[0015] The acquisition unit 12 acquires data input from the input unit 11. The acquisition unit 12 may acquire target single cell data, at least one type of first item, and at least one type of second item data. The acquisition unit 12 may store the acquired data in the storage unit 2. The acquisition unit 12 may perform preprocessing on the acquired data.

[0016] The extraction unit 13 extracts features based on the target single-cell data 22 acquired by the acquisition unit 12. Examples of features include the types and expression levels of the above-mentioned proteins; types and expression levels of RNA; types, modifications, and expression levels of DNA; types, modifications, open / closed states, amounts, and structures of chromatin; chromosome structures and interactions; epigenomic information; and binding of transcriptional regulators. Methods for extracting features include, for example, a filter method using the ranking of expression levels of each type of protein or RNA, the relative expression levels of each type, and the presence or absence of expression of each type; an embedded method using information contributing to model accuracy obtained through machine learning model creation; and a dimensionality reduction method.

[0017] The estimation unit 14 inputs the target single cell data 22 into the trained machine learning model 21 to estimate estimated single cell data 23. The estimation unit 14 may estimate only one type of candidate for the estimated single cell data 23, or may estimate two or more types. The estimation unit 14 may store the estimated single cell data 23 in the storage unit 2.

[0018] The output control unit 15 causes the output unit 16 to output the estimation result output by the estimation unit 14.

[0019] <Storage section 2> The storage unit 2 may store a machine learning model 21. Furthermore, target single cell data 22 and estimated single cell data 23 may be stored as necessary.

[0020] The machine learning model 21 learns the association between first specimen single cell data relating to a first item obtained from a specimen single cell collected from a specimen, and second specimen single cell data relating to a second item different from the first item. This enables the machine learning model to estimate estimated single cell data 23 relating to the second item from target single cell data 22 relating to the first item.

[0021] As used herein, the term "specimen" refers to any organism from which a single cell is collected to be used in creating a machine learning model. The specimen may be a living organism or a dead organism. Organisms that can serve as specimens are not particularly limited as long as they are organisms from which the target single cell and specimen single cell described below can be derived. In other words, the specimen may be any organism, including plants, animals, and insects.

[0022] In most cases, analyzing proteins and the like from a single cell requires destroying the cell. Therefore, the upper limit of the number of parameters that can be simultaneously analyzed for a single cell was about three. The information processing device 10 makes it possible to estimate many parameters that have not been obtained for the single cell from parameters obtained at the single-cell level, thereby enabling simultaneous analysis of various parameters.

[0023] Furthermore, while conventional methods for analyzing parameters of single cells could only analyze specific combinations of parameters, the information processing device 10 can estimate any combination of all parameters that have been machine-learned in the machine learning model 21, rather than just specific combinations of parameters in a single cell, thereby enabling simultaneous parameter analysis of combinations of parameters that were previously difficult to perform. (Information processing method)

[0024] An information processing method according to one embodiment of the present invention is an information processing method executed by an information processing device 10, and includes an acquisition step of acquiring target single cell data 22 relating to a first item acquired from a target single cell, and an estimation step of generating estimated single cell data 23 relating to a second item of the target single cell from the target single cell data 22 using a machine learning model 21 that has learned the association between first specimen single cell data relating to the first item acquired from a specimen single cell collected from a specimen, and second specimen single cell data relating to a second item different from the first item.

[0025] An information processing method according to one embodiment of the present invention will be described with reference to Fig. 2. Fig. 2 is a flowchart showing an outline of the information processing method according to one embodiment of the present invention.

[0026] In step S1, the acquisition unit 12 acquires target single cell data 22 from a target single cell (acquisition step). In this specification, single cell data means data obtained from a single cell. In other words, single cell data is data about an individual cell. In step S1, target single cell data 22 may be acquired from each of a plurality of single cells.

[0027] Here, the target single cell data 22 relates to a first item. The first item is not particularly limited as long as it can be commonly used as cell data. Examples of the first item include one or more items selected from the group consisting of: protein type and expression level; RNA type and expression level; DNA type, modification, and expression level; chromatin type, modification, open / closed state, amount, and structure; chromosome structure and interaction; epigenomic information; and binding of transcriptional regulators. Among these, the first item is preferably one or more items selected from data related to protein, RNA, DNA, and chromatin, as this makes it easier to estimate the state of the target. More preferably, the first item is one or more items selected from protein type, RNA type, DNA type, and chromatin type.

[0028] In one embodiment, after step S1, the acquisition unit 12 or the extraction unit 13 may perform preprocessing on the target single-cell data 22 related to the first item. Examples of preprocessing include, but are not limited to, averaging, standardization, and outlier removal. By performing preprocessing on the target single-cell data, the accuracy of the estimated single-cell data 23 can be improved.

[0029] In step S2, the extraction unit 13 extracts features based on the target single-cell data 22 acquired in step S1. From the viewpoint of improving the accuracy of the data, in step S2, it is preferable to perform further preprocessing on the extracted features, such as reducing the feature dimension, undersampling, oversampling, averaging, noise reduction, and masking.

[0030] In step S3, the estimation unit 14 inputs the feature amount extracted in step S2 into the machine learning model 21 to estimate estimated single cell data 23 of the target single cell. At this time, the machine learning model 21 is obtained by machine learning using specimen single cell data obtained for a specimen single cell collected from a specimen.

[0031] Here, the estimated single cell data 23 relates to a second item. As the second item, any data that can be selected as the first item can be selected as long as it is of a different type from the first item. Specifically, the second item can be one or more types selected from the following: protein type and expression level; RNA type and expression level; DNA type, modification, and expression level; chromatin type, modification, open / closed state, amount, and structure; chromosome structure and interaction; epigenomic information; and binding of transcriptional regulators. Among these, the second item is preferably a parameter related to protein, RNA, DNA, or chromatin, as it is easier to estimate the state of the target, and more preferably one or more types selected from protein type, RNA type, DNA type, and chromatin type.

[0032] The information processing method may include step S4, in which the output control unit 15 outputs the estimated single cell data 23. In step S4, the information processing device outputs the estimated single cell data 23 to the output device as single cell data related to the target single cell. By including the output step in the information processing method, the estimated estimated single cell data 23 can be easily confirmed. In step S4, the target single cell data 22 may further be output.

[0033] The target single cell and the specimen single cell may be either a prokaryotic cell or a eukaryotic cell, but are preferably eukaryotic cells. Furthermore, the origin of the target single cell and the specimen single cell is not particularly limited, and they may be, for example, animal cells, plant cells, bacteria, archaea, or protist microbial cells. Among these, animal cells, plant cells, or insect cells are preferred, animal cells or plant cells are more preferred, and animal cells are even more preferred. Examples of animal cells include mammalian cells, avian cells, amphibian cells, fish cells, and insect cells, among which mammalian cells are preferred, and more preferably cells from humans, mice, rats, rabbits, dogs, or non-human primates.

[0034] The target single cell and the specimen single cell may be different cells or the same cell. The target single cell and the specimen single cell are preferably cells of organisms belonging to the same class, more preferably cells of organisms belonging to the same order, even more preferably cells of organisms belonging to the same family, even more preferably cells of organisms belonging to the same genus, and particularly preferably cells of organisms of the same species. With the above configuration, the accuracy of the obtained estimation result is improved.

[0035] Furthermore, the target single cell and the specimen single cell are preferably cells derived from the same tissue. Examples of the target single cell and the specimen single cell include mesenchymal stem cells, ES cells, keratinocytes, fibroblasts, bone marrow cells, endothelial cells, smooth muscle cells, Schwann cells, chondrocytes, adipocytes, osteoblasts, and endothelial progenitor cells. If the single cell and the specimen single cell are derived from the same tissue, the accuracy of the obtained estimation result is improved.

[0036] In most cases, analysis of proteins and the like in single cells requires destruction of the cells before measurement. Therefore, the upper limit of the number of parameters that can be simultaneously measured for a single cell is approximately three. According to an information processing method according to one embodiment of the present invention, a machine learning model 21 is created based on various types of data acquired from a large number of single cells, so that a large number of parameters in a single cell can be simultaneously estimated.

[0037] Furthermore, while conventional methods for analyzing parameters of single cells can only measure specific combinations of parameters, the information processing method according to one embodiment of the present invention can estimate any combination of all parameters that have been machine-learned by the machine learning model 21, not just specific combinations of parameters in a single cell.

[0038] <How to create machine learning model 21> A method for creating a machine learning model 21 according to one embodiment of the present invention includes a learning data acquisition step of acquiring first specimen single-cell data relating to a first item obtained from a specimen single cell collected from a specimen, and second specimen single-cell data relating to a second item different from the first item, and a generation step of generating the machine learning model by performing machine learning using the first specimen single-cell data as an explanatory variable and the second specimen single-cell data as a target variable.

[0039] A method for creating the machine learning model 21 will be described with reference to Fig. 3. Fig. 3 is a flowchart showing an example of a method for creating a machine learning model.

[0040] Step S11 is a learning data acquisition step of acquiring first specimen single cell data and second specimen single cell data relating to a second item for specimen single cells obtained from a specimen of a particular status.

[0041] In step S12, machine learning is performed using the first sample single-cell data as an explanatory variable and the second sample single-cell data as a response variable. Step S13 may be performed by a machine learning device including a learning unit for performing the machine learning.

[0042] The machine learning method in step S12 may be transfer learning such as HeMap, or deep learning such as Conditional Variational AutoEncoder (CVAE), a combination of CVAE and Generative Adversarial Networks (GAN), or a diffusion model. When deep learning or the like is used as the machine learning method, the first sample single-cell data and the second sample single-cell data are labeled with the sample status in advance. When the first and second single-cell data labeled with the sample status are used, the machine learning model may use the first sample single-cell data as an explanatory variable and machine learning using the second sample single-cell data as a target variable for each sample status. From the viewpoint of improving the accuracy of the obtained estimated single-cell data, transfer learning is preferably used as the machine learning method.

[0043] As used herein, the term "specimen status" refers to information related to the organism that serves as the specimen. Examples of such information include status information, information indicating the species of the specimen, and information on the type of a single cell in the specimen.

[0044] In one embodiment, the specimen status may be information indicating symptoms, diseases, conditions, etc. of the organism serving as the specimen. Specifically, if the specimen is an animal, the specimen status may include, for example, the presence or absence of infection, inflammation, bleeding, and disease. More specific specimen status may be, for example, information simply indicating a symptom, such as "inflammation," information on the site where the symptom occurs, or information on the specific disease name of the specimen. Furthermore, if the specimen is a plant, the specimen status may include, for example, the degree of dryness, temperature, salt concentration, and the presence or absence of disease in the growing environment. More specific specimen status may be, for example, information simply indicating "dry," information on the site where changes due to the dry state occur, or information on the specific growing environment to which the specimen is exposed.

[0045] Examples of inflammation include, but are not limited to, dermatitis (e.g., allergic dermatitis and atopic dermatitis), arthritis, vasculitis, glomerulonephritis, gastroenteritis, pneumonia, meningitis, obesity, and aging. Examples of bleeding include, but are not limited to, bleeding in the digestive organs and bleeding in other organs. Examples of diseases include, but are not limited to, autoimmune diseases, neurological and psychiatric disorders (e.g., Alzheimer's disease, stroke, Parkinson's disease, and schizophrenia), muscle degenerative diseases (e.g., muscular dystrophy), arteriosclerosis, and cancer.

[0046] The specimen status may include only one type of information, or may include multiple types of information.

[0047] The machine learning model is preferably created by transfer learning of the association between the single-cell data of the first specimen and the single-cell data of the second specimen. When the machine learning model is obtained by transfer learning, a latent space is generated by fusing the single-cell data from the common information potentially present among the measurement data of each parameter of multiple types of single cells of specimens obtained from the specimen. With the above configuration, multiple types of data obtained from single cells of specimens having a specific status are fused, making it possible to detect elements that are difficult to detect using data obtained by measuring a single type of single cell.

[0048] Before step S13, the first sample single-cell data and the second sample single-cell data used to create the machine learning model may be subjected to preprocessing such as reduction of feature dimensions, undersampling, oversampling, averaging, noise reduction, masking, etc. The preprocessing can be performed by any component including a learning unit, as long as it is performed before step S14.

[0049] In step S13, the machine learning model 21 is created using the data that was machine-learned in S12. In step S14, the machine learning model is stored, and the processing in FIG. 3 ends. The machine learning model may be stored in a location that can be read by an information processing device that can execute an information processing method according to an embodiment of the present invention. For example, it may be stored in the storage unit 2 of the information processing device 10.

[0050] In one embodiment, a machine learning model is created by the following method. The following method is an example in which transfer learning using HeMap is used as the machine learning method and a mouse is used as the specimen organism. The method for creating a machine learning model in the present invention is not limited to the following. 1. Collect multiple cells from each mouse. 2. The collected cells are divided into single cells, and the type of gene, type of chromatin, or type of protein is analyzed for each single cell to obtain training data. 3. Using the acquired training data, transfer learning is performed using HeMap to create a machine learning model, which is a latent space. 4. Create a model that estimates the features of each measurement data based on the machine learning model.

[0051] [Embodiment 2] (information processing device 10a) An overview of an information processing device 10a according to another embodiment of the present invention will be described with reference to Fig. 4. Note that descriptions of matters that have already been described will be omitted.

[0052] The information processing device 10a includes a control unit 1a that controls all the units of the information processing device 10a, a storage unit 2a that stores various data used by the information processing device 10a, an input unit 11, and an output unit 16. The control unit 1a includes an acquisition unit 12a, an estimation unit 14a, and an accuracy verification unit 17. The storage unit 2a may store a machine learning model 21a, a third estimated data group 24 related to a third item, and third target data 25 related to the third item.

[0053] Here, the third item may be data that can be selected as the first item and the second item, or may be data that cannot be selected as the first item and the second item, as long as the data is different from the first item and the second item. In one embodiment, the third item may be the status of the subject and the specimen.

[0054] As used herein, "subject" refers to any organism from which a target single cell has been collected. That is, a target single cell is a single cell collected from a "subject." The target may be the same organism as the specimen, or a different organism (e.g., a closely related species). Organisms that can be selected as targets include organisms of the same type as the specimen described above. Furthermore, as used herein, "subject status" refers to information related to the subject, and is the same type of information as the "specimen status" described above. That is, if the specimen status is information indicating whether or not the specimen is in an inflammatory state, the subject status is information indicating whether or not the subject is in the inflammatory state.

[0055] <Control unit 1a> The acquiring unit 12a may acquire third target data 25 in addition to the target single cell data 22. The estimation unit 14a estimates third estimated data in addition to the estimated single-cell data 23. The estimation unit 14a inputs the estimated single-cell data 23 into a machine learning model to estimate a third estimated data group 24. The estimation unit 14a may store the third estimated data in the storage unit 2a. In one embodiment, when there are multiple machine learning models, the third estimated data may be a third estimated data group 24 including multiple data. In other words, when there are multiple pieces of third estimated data, the third estimated data is referred to as the third estimated data group 24.

[0056] In one embodiment, when a method such as HeMap is used as the machine learning method, the control unit 1a may include an accuracy verification unit 17 that verifies the accuracy of the estimated single cell data 23 output from the machine learning model.

[0057] The accuracy verification unit 17 may verify the accuracy of the estimated single-cell data 23 using the second to fifth machine learning models described in (1) to (4) below. When the accuracy verification unit 17 uses data related to a third item, the third item is preferably data common to the specimen and the subject, and more preferably data that is useful or of interest as a measurement subject. In one embodiment, the estimation unit 14a may further include the second to fifth machine learning models. Alternatively, the second to fifth machine learning models may be included in a device different from the information processing device 10a. (1) A second machine learning model trained using the first sample single-cell data as the explanatory variable and the third sample data on the third item as the objective variable. (2) A third machine learning model trained using the second sample single-cell data as the explanatory variable and the third sample data on the third item as the objective variable. (3) A fourth machine learning model that is machine-trained using the first sample single-cell data and the second sample estimated single-cell data estimated by inputting the first sample single-cell data into the machine learning model as explanatory variables and the third sample data related to the third item as the target variable. (4) A fifth machine learning model was created using data obtained by randomly combining the first sample single-cell data and the second sample single-cell data as explanatory variables and the third sample data related to the third item as the objective variable.

[0058] First, the accuracy verification unit 17 performs the following steps (1') to (4') to generate a third estimated data group 24. The third estimated data group 24 includes first to fourth target estimated data. (1') First subject estimation data obtained by inputting subject single cell data 22 for the first item into the second machine learning model. (2') Second subject inferred data obtained by inputting single-cell data of the subject measured for the second item into a third machine learning model. (3') Third target estimated data obtained by inputting data obtained by fusing the target single cell data 22 and the estimated single cell data 23 in a predetermined format into the fourth machine learning model. (4') Fourth target estimated data obtained by inputting data that randomly combines target single cell data 22 for the first item and target single cell data 23 measured for the second item into the fifth machine learning model.

[0059] Furthermore, the accuracy verification unit 17 compares each of the first to fourth target estimation data with the third target data 25, and calculates the matching rates between the first to fourth target estimation data and the third target data 25. Hereinafter, the calculated matching rates will be referred to as the first result accuracy, the second result accuracy, the third result accuracy, and the fourth result accuracy, respectively.

[0060] The accuracy verification unit 17 verifies the accuracy of the estimated single cell data 23 based on the coincidence rate calculated from each estimated data. <Case 1> For example, if the first result accuracy is smaller than the second result accuracy, and the third result accuracy is larger than the fourth result accuracy and the third result accuracy is larger than the first result accuracy, the accuracy verification unit 17 determines that the accuracy of the estimated single-cell data 23 is sufficiently high. Then, the accuracy verification unit 17 outputs the estimated single-cell data 23 to the output control unit 15. <Case 2> For example, if the first result accuracy is greater than the second result accuracy, and the third result accuracy is greater than the fourth result accuracy, and the third result accuracy is greater than the second result accuracy, the accuracy verification unit 17 determines that the accuracy of the estimated single-cell data 23 is sufficiently high. Then, the accuracy verification unit 17 outputs the estimated single-cell data 23 to the output control unit 15. In the above cases 1 and 2, the accuracy verification unit 17 may label the estimated single-cell data 23 to indicate that the accuracy of the data is high. <Case 3> For example, if the third result accuracy is smaller than the fourth result accuracy, the accuracy verification unit 17 determines that the accuracy of the estimated single cell data 23 is low. In this case, the accuracy verification unit 17 may be configured not to output the estimated single cell data 23 to the output control unit 15. The accuracy verification unit 17 may label the estimated single cell data 23 that has been determined to have low accuracy, indicating that the data is low in accuracy, and then output the data to the output control unit 15.

[0061] <Storage section 2a> In one embodiment, the machine learning model 21a may have learned associations between the first sample single-cell data, the second sample single-cell data, and third sample data regarding a third item different from both the first item and the second item, and may further generate third estimated data regarding the third item from the target single-cell data and the estimated single-cell data in the estimation step.As another example, the machine learning model 21a may have learned associations between the first sample single-cell data, the second sample estimated single-cell data, and the third sample data.

[0062] The third estimated data group 24 may include first to fourth target estimated data output from the second to fifth machine learning models, respectively.

[0063] The third target data 25 is data relating to a third item of interest. The third target data 25 may be acquired by the acquisition unit 12a.

[0064] [Software implementation example] The functions of an estimation device (hereinafter referred to as the "device") according to one embodiment of the present invention can be realized by a program for causing a computer to function as the device, and a program for causing a computer to function as each control block of the device (particularly each part included in the control unit 1).

[0065] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. The control device and storage device execute the program, thereby realizing the functions described in each of the above embodiments.

[0066] The program may be non-transitory and may be recorded on one or more computer-readable recording media. The recording media may or may not be included in the device. In the latter case, the program may be supplied to the device via any wired or wireless transmission medium.

[0067] Furthermore, some or all of the functions of the control blocks can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the control blocks is formed is also included in the scope of the present invention. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.

[0068] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI ​​may run on the control device or on another device (for example, an edge computer or a cloud server).

[0069] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Example]

[0070] An embodiment of the present invention will now be described.

[0071] 〔material〕 <Mouse> Experiments were performed using 7- to 13-week-old Hlf-tdTomato mice. 100 μg of lipopolysaccharide (LPS) was injected intraperitoneally 16 hours before flow cytometry or RT-PCR. For comparison, mice were injected intraperitoneally with 100 μm of PBS.

[0072] [Creating a machine learning model] Two to three Hlf-tdTomato mice injected with LPS or PBS were used as specimens, and experiments were performed twice. RFP-positive hematopoietic stem cells (tdTomato) were isolated from each mouse. + CD150 + CD48 -Single cells from a mouse (LSK) were acquired and the protein, RNA, and chromatin types were measured. The number of measurement samples was 1,514 proteins, 63 RNAs, and 288 chromatins, respectively. The measured data were used as the first or second specimen single-cell data, and machine learning was performed to determine the relationships between these measurement data to obtain a machine learning model. HeMap, a transfer learning algorithm, was used for machine learning. The five proteins with the highest expression levels as measured in the obtained machine learning model were input as the target single-cell data for the first item, and the RNA and chromatin types were estimated as the inferred single-cell data for the second item. The results are shown in Figure 5.

[0073] 〔analysis〕 To confirm the accuracy of the RNA results predicted by the machine learning model, we analyzed the expression levels of the predicted RNAs shown in Figure 5 and the proteins encoded by the RNAs.

[0074] Each measurement was performed three times, and the results are expressed as the mean ± SEM of two independent experiments. Statistical analysis of the statistical data was performed using Student's t-test.

[0075] [RNA expression analysis] <Analysis method> Analysis of RNA expression levels was performed by qRT-PCR. RNAs that appeared twice or more in the table in Figure 7 were selected for analysis. In addition, Mpeg2 / Il2rg was also analyzed because it is known to be closely related to inflammation. RFP-positive hematopoietic stem cells (tdTomato) were isolated from Hlf-tdTomato mice that had been intraperitoneally injected with LPS or PBS. + CD150 + CD48 -LSK (Likely a marker) was added to a 96-well plate at a density of 100 cells per well. After adding 10 μL of cell lysis buffer to each well, the cells were immediately frozen (-80°C) and analyzed according to the RamDA-seq™ method (Hayashi, T. et al. Nature Communications 9, 619 (2018)). First-strand cDNA was synthesized using the Primer Script RT Reagent Kit (Takara Bio Inc.) and NSR (not-so-random) primers. Second-strand cDNA was synthesized using Klenow Fragments (30–50 U / μL, New England Biolabs) and the complementary strand of the NSR primer. Quantitative PCR analysis was performed using SYBR Green (Applied Biosystems) on a Real-time PCR LightCycler 96 (Roche). The relative mRNA expression levels of each gene were normalized to the expression level of the Gapdh gene. The primers used to detect each gene are shown in Figure 6.

[0076] <Result> The analysis results are shown in Figure 7. In Figure 7, the vertical axis shows the relative expression level of each gene. Sca1 is the positive control, and Kit is the negative control. As shown in Figure 9, the expression levels of Lag3, IL2rg, Lfitm3, and Mpeg1 were increased. The above results demonstrate that the information processing method according to one embodiment of the present invention makes it possible to accurately estimate a large number of parameters related to a single cell.

[0077] [Protein expression analysis] <Analysis method> Protein expression levels of Lag3 and IL2rg, which showed increased expression levels in the RNA expression analysis, were analyzed by flow cytometry. Femurs and tibias were collected from Hlf-tdTomato mice injected intraperitoneally with LPS or PBS. Bone marrow cells were extracted from the bones using a syringe and suspended in FACS buffer (2% FBS and 2 mM EDTA in PBS). The suspension was centrifuged at 470G for 5 minutes, the supernatant was removed, and the cells were suspended in ACK lysis buffer to remove red blood cells from the bone marrow cells. The suspension was centrifuged at 470G for 5 minutes, the supernatant was removed, and the cells were resuspended in FACS buffer. Lineage (Lin) marker was used to label the bone marrow cells. Specifically, bone marrow cells were labeled using biotinylated antibodies: anti-TER-119 antibody, anti-NK1.1 antibody, anti-CD11b (M1 / 70) antibody, anti-Gr-1 (RB6-8C5) antibody, anti-CD3e (145-2C11) antibody, anti-CD4 (GK1.5) antibody, anti-CD8a (53-6.7) antibody, anti-B220 (RA3-6B2) antibody, and anti-IL7RA (SB / 199) antibody (all manufactured by BioLegend), and cell fractions positive for the Lin marker were excluded using MACS streptavidin magnetic beads (manufactured by Miltenyi biotec).

[0078] The Lin marker-negative fraction was fluorescently labeled with antibodies: anti-c-Kit (2B8), anti-Sca-1 (D7), anti-CD48 (HM45-1), anti-CD150 (TC15-12F12.2), anti-Lag3 (C9B7W), and anti-IL2rg (TUGm2) (all manufactured by BioLegend). Hematopoietic stem cells (Lin) were then isolated from the fluorescently labeled cell population. - c-Kit + Sca-1 + CD48 ― CD150 + tdTomato +) were separated using a FACS Aria III and analyzed using FlowJo (BD Biosciences). A separate population of fluorescently labeled cells was analyzed using a FACSymphony™ without cell separation.

[0079] <Result> The results of the protein analysis are shown in Figure 8. Figure 8 is a graph showing the fluorescence intensity of each marker in flow cytometry, and a graph showing the relative expression level of the protein in each mouse. The graph showing fluorescence intensity in Figure 8 shows the results of measurements performed on two mice injected with PBS or LPS. Figure 8 shows that both Lag3 and Il2rg, whose gene expression levels were increased, were higher in mice injected with LPS than in mice injected with PBS. From the above, it can be seen that the RNA estimated by the information processing method according to one embodiment of the present invention also has an increased expression level at the protein level.

[0080] [RNA-only data] <Measurement method> RFP-positive hematopoietic stem cells (Lin) were isolated from Hlf-tdTomato mice injected intraperitoneally with LPS or PBS. - c-Kit + Sca-1 + CD48 ― CD150 + tdTomato +) were harvested and placed in a 96-well plate at a density of one cell per well. After adding 1 μL of cell lysis buffer to each well, the wells were immediately frozen (-80°C). RNA analysis was performed according to the RamDA-seq™ method (Hayashi, T. et al. Nature Communications 9, 619 (2018)). The quantity and quality of the isolated cDNA libraries were assessed using Multina (Shimadzu Corporation, Japan) and Bioanalyzer 2100 (Agilent, Waldbronn, Germany), followed by sequencing using the Next-Seq system (Illumina Inc.). Sequencing data were trimmed using the trim_galore (version 0.4.3) package (quality and length parameters set to 30 and 30), and then filtered using SortMeRNA to remove mitochondrial mRNA. The filtered data were aligned to the mouse reference sequence (GRCm38, M17 version) using the RNA-seq aligner STAR (version 2.5.3a; Dobin et al., 2012). From the aligned bam files, along with the mouse GFF annotation file (GRCm38, M17 version), the number of reads for each gene and sample was calculated using featureCounts in the Subreads program (version 1.6.1) to create a gene count matrix. Calculation of differentially expressed genes between LPS and PBS samples was performed using R (version 3.4) and the DESeq2 (version 1.20.0) package. <Result> The results of measuring RNA alone are shown in Figure 9. As shown in Figure 9, Il2rg and Ifitm3 were expressed in both PBS- and LPS-injected mice, with expression levels increased in LPS-injected mice. On the other hand, Lag3 and Gpb4 were expressed only in LPS-injected mice.

[0081] In the data for RNA alone, Lag3 and Il2rg are ranked low. On the other hand, the data shown in Figures 7 and 8 show that the expression levels of Lag3 and Il2rg are relatively increased in mice injected with LPS compared to mice injected with PBS. This suggests that the information processing method according to one embodiment of the present invention may be capable of estimating factors that are difficult to detect from single data.

[0082] [Verification of estimated single data] <Creating a second machine learning model> A second machine learning model was created by machine learning using the measurement data used to create the first machine learning model. The machine learning was performed using a combination of four types of explanatory variables and a target variable, and the second machine learning model was obtained. (1) Explanatory variables: data on 1,514 measured proteins, Response variables: PBS injection or LPS injection of the sample mice (2) Explanatory variables: data from 63 measured RNAs, Objective variables: PBS or LPS injection of the sample mice (3) Explanatory variables: Data obtained by combining data on 1,514 measured proteins and data on 63 RNAs estimated by the machine learning model; Objective variables: PBS or LPS injection of the sample mice (4) Explanatory variables: Data obtained by randomly combining data from 1,514 measured proteins and data from 63 measured RNAs. Response variables: PBS or LPS injections of the sample mice. In the second machine learning model, (1) was input with the protein measurement results (target single cell data), (2) with the estimated RNA data (estimated single cell data obtained by the machine learning model), and (3) and (4) with the fusion data of the target single cell data and the estimated single cell data, and a discrimination result was obtained as to whether the mouse was injected with LPS or PBS. The obtained discrimination result was compared with the actual results of PBS injection or LPS injection for the mouse from which the target single cell data was obtained, and the accuracy of the estimated single cell data was verified.

[0083] <Result> The results are shown in Figure 10. In Figure 10, the measured data is the result of inputting only the target single-cell data and using the machine learning models (1) or (2) to determine whether it was a PBS injection or an LPS injection. On the other hand, the estimated single-source data is the result of inputting the target single-cell data as well as the estimated single-source data estimated by the machine learning model and using the machine learning models (3) or (4) to determine whether it was a PBS injection or an LPS injection. Note that for the estimated single-source data, only the protein type was used as the target single-cell data.

[0084] As shown in Figure 10, when estimated data on RNA types was used in addition to protein types, the accuracy of the determination was improved compared to when only measured data on protein types was used. Also, the accuracy of the determination was similar to when only measured data on RNA types was used. Furthermore, when estimated data on chromatin types was used in addition, the accuracy of the determination was improved compared to when only measured data on RNA types was used. These results demonstrate that estimated single-source data can estimate data with high accuracy. [Industrial Applicability]

[0085] The present invention can be used to estimate multiple parameters of a single cell, as well as to estimate the state of a subject. [Explanation of symbols]

[0086] 1, 1a, 1b Control unit 2, 2a, 2b storage section 10, 10a, 10b Information processing device 11 Input section 12, 12a Acquisition part 13 Extraction part 14, 14a Estimation section 15 Output control section 16 Output section 17 Accuracy Verification Department 21 Machine Learning Models 22 target single cell data 23 Estimated single cell data 24 Third Estimation Data Set 25 Third Target Data

Claims

1. An information processing method executed by an information processing device, an acquisition step of acquiring target single cell data relating to a first item acquired from the target single cell; an estimation step of generating estimated single cell data regarding the second item of the target single cell from the target single cell data using a machine learning model that has learned the association between first specimen single cell data regarding the first item obtained from a specimen single cell collected from the specimen and second specimen single cell data regarding a second item different from the first item; An information processing method comprising:

2. an output step of causing an output device to output the target single cell data and the estimated single cell data as single cell data regarding the target single cell; The information processing method according to claim 1 .

3. the first item is one or more types of data selected from the group consisting of types and expression levels of proteins; types and expression levels of RNA; types, modifications, and expression levels of DNA; types, modifications, open / closed states, amounts, and structures of chromatin; structures and interactions of chromosomes; epigenome information; and binding of transcriptional regulatory factors; The information processing method of claim 1, wherein the second item is data different from the first item and is one or more types of data selected from the group consisting of: protein type and expression level; RNA type and expression level; DNA type, modification and expression level; chromatin type, modification, open / closed state, amount and structure; chromosome structure and interaction; epigenomic information; and binding of transcriptional regulators.

4. The machine learning model further comprises: The system has already learned associations between the first sample single-cell data, the second sample single-cell data, and third sample data relating to a third item different from both the first item and the second item; In the estimation step, third estimated data regarding a third item is further generated from the target single-cell data and the estimated single-cell data. The method of claim 1.

5. The information processing method according to claim 1 , wherein after the acquisition step, one or more processes selected from the group consisting of averaging, standardization, and outlier removal are performed on the target single-cell data relating to the first item.

6. The information processing method according to claim 1 , wherein the machine learning model is created by transfer learning of an association between the first sample single-cell data and the second sample single-cell data.

7. The information processing method according to claim 1 , wherein the target single cell and the specimen single cell are cells of the same biological species.

8. The information processing method according to claim 1 , wherein the target single cell and the specimen single cell are cells derived from the same tissue.

9. a learning data acquisition step of acquiring first specimen single cell data relating to a first item acquired from a specimen single cell collected from the specimen, and second specimen single cell data relating to a second item different from the first item; generating a machine learning model by performing machine learning using the first sample single-cell data as an explanatory variable and the second sample single-cell data as a response variable, How to create a machine learning model.

10. an acquisition unit that acquires target single cell data related to a first item acquired from the target single cell; an estimation unit that generates estimated single cell data regarding the second item of the target single cell from the target single cell data using a machine learning model that has learned the association between first specimen single cell data regarding the first item obtained from a specimen single cell collected from a specimen and second specimen single cell data regarding a second item different from the first item; An information processing device comprising:

11. A control program for causing a computer to function as the information processing device according to claim 10, the control program causing a computer to function as the acquisition unit and the estimation unit.

12. A computer-readable recording medium on which the control program according to claim 11 is recorded.

Citation Information

Patent Citations

  • Apparatus, method, and system for image-based human germ cell classification

    JP2022087297A

  • Compositions and methods for light-induced biomolecular barcoding

    JP2023506176A

Cited By

  • Game machine

    JP2026056552A