Batch effect correction system, control program, and batch effect correction method
The batch effect correction system uses a variational autoencoder to separate and correct measurement environment-induced fluctuations from cellular differences, ensuring accurate reconstruction of biological signals and conditions in omics analysis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-03-26
AI Technical Summary
Existing batch effect correction methods using variational autoencoders risk losing biologically meaningful differences between cellular conditions while attempting to correct measurement protocol and environment-induced fluctuations, limiting their applicability.
A batch effect correction system utilizing a variational autoencoder separates latent variables into those indicating cellular differences and measurement environment differences, using machine learning to minimize prediction errors and preserve biological signals during correction.
The system effectively corrects batch effects while maintaining accurate reconstruction of cellular conditions and types, enhancing the reliability of omics analysis results.
Smart Images

Figure JP2025032850_26032026_PF_FP_ABST
Abstract
Description
Batch effect correction system, control program, batch effect correction method
[0001] The present invention relates to a batch effect correction system, a control program, and a batch effect correction method.
[0002] In measurement data obtained from omics analysis such as transcriptomics, it is known that even when the conditions of the target cells are the same, unintended fluctuations in the measurement values can occur due to differences in the measurement protocol and the measurement environment, such as the experimenter. These fluctuations are called batch effects, and methods have been proposed to remove them from measurement data of omics analysis.
[0003] For example, Non-Patent Document 1 discloses a method for correcting batch effects using a variational autoencoder.
[0004] R. Lopez et al., Deep generative modeling for single-cell transcriptomics., Nature Methods 15, 1053-1058, 2018
[0005] However, in the prior art disclosed in Non-Patent Document 1, there was a risk that the batch effect correction would also correct information indicating biologically meaningful differences between cellular conditions.
[0006] One aspect of the present invention aims to realize a batch effect correction system that can specifically correct batch effects.
[0007] To solve the aforementioned problems, a batch effect correction system according to one aspect of the present invention includes: an encoding unit that compresses information of a plurality of input cells into a multi-dimensional latent variable for each cell, which includes a first latent variable that varies due to differences in the cells and a second latent variable that varies due to differences in the measurement environment of the cells; a correction unit that corrects the second latent variable so that the difference in the second latent variable between the cells becomes smaller; a decoding unit that reconstructs the cell information using the corrected second latent variable; a first prediction unit that predicts the difference in the measurement environment using the second latent variable; and a first prediction unit that uses the first latent variable to determine the difference between the cells. The information processing device includes a second prediction unit that makes predictions, and the encoding unit is machine-learned to obtain a second latent variable that minimizes the prediction error between the prediction result of the first prediction unit and the first label information, using training data in which information of a plurality of sample cells and first label information indicating the differences in the measurement environment of the sample cells are associated, and the first latent variable that minimizes the prediction error between the prediction result of the second prediction unit and the second label information, using training data in which information of a plurality of sample cells and second label information indicating the differences between the sample cells are associated, and the first latent variable that minimizes the prediction error between the prediction result of the second prediction unit and the second label information, is obtained through the compression.
[0008] To solve the aforementioned problems, a batch effect correction method according to one aspect of the present invention is a method performed by a computer, comprising the steps of: compressing information of a plurality of input cells into a multidimensional latent variable for each cell, including a first latent variable that varies due to differences in the cells and a second latent variable that varies due to differences in the measurement environment of the cells; correcting the second latent variable so that the difference in the second latent variable between the cells becomes small; reconstructing the cell information using the corrected latent variable; predicting the differences in the measurement environment using the second latent variable; and using the first latent variable to... The compression step includes a step of predicting differences, wherein the compression step uses training data in which information of multiple sample cells and first label information indicating differences in the measurement environment of the sample cells are associated, and machine learning is performed so that the second latent variable that minimizes the prediction error between the prediction result of the step of predicting differences in the measurement environment and the first label information is obtained by the compression, and machine learning is performed using training data in which information of multiple sample cells and second label information indicating differences in the sample cells are associated, so that the first latent variable that minimizes the prediction error between the prediction result of the step of predicting differences in the cells and the second label information is obtained by the compression.
[0009] According to one aspect of the present invention, a batch effect correction system that can specifically correct batch effects can be realized.
[0010] This is a block diagram showing an example of the schematic configuration of the correction system according to Embodiment 1. This is a flowchart showing an example of the flow of the learning process performed by the correction device according to Embodiment 1. This is a flowchart showing an example of the flow of the correction process performed by the correction device according to Embodiment 1. This is a diagram showing an example of the schematic configuration of the correction system according to Embodiment 2. This is a block diagram showing an example of the schematic configuration of the correction device according to Embodiment 2. In the condition variables of Embodiment 1, the left figure shows the distribution of values for each condition, and the right figure shows the distribution of values for each batch. In the cell type variables of Embodiment 1, the upper figure shows the distribution for each batch, and the lower figure shows the distribution for each cell type. This is a diagram showing the distribution for each batch in the batch variables of Embodiment 1. This is a diagram showing the correlation between the actual gene expression level and the reconstructed gene expression level in Embodiment 1. This is a diagram showing that in the dimensionality reduction space before and after correction, the difference between batches is corrected, while at the same time, the difference between conditions and between cell types is maintained.
[0011] [Embodiment 1] <Summary of the present invention> Conventionally, in omics analysis, it is known that even when measuring the same object, unintended variations in the measured values, known as batch effects, can occur due to differences in various factors such as the protocol or the experimenter. Batch effects can confound biologically meaningful signals included in the measured values, making it difficult to interpret the measurement results. Therefore, removing batch effects from measurement data in omics analysis is important for performing omics analysis efficiently and accurately.
[0012] For this reason, various batch effect correction methods, such as scVI disclosed in Non-Patent Document 1, have been proposed. scVI is known as a method for dimensionality reduction and batch effect correction using a variational autoencoder, which is a type of generative model.
[0013] The scVI batch correction method corrects the latent variable space obtained by dimensionality reduction to eliminate differences between batches. However, conventional methods that correct for information caused by batch effects also cause at least some of the information caused by cellular differences that are not due to batch effects (i.e., biologically meaningful signals, hereinafter also referred to as "biological signals"). Therefore, the scope of application of conventional methods was limited. Biological signals are, for example, cellular information that fluctuates depending on the presence or absence of disease, such as information showing an increase or decrease in the expression level of a certain gene depending on the presence or absence of disease.
[0014] The problem with conventional correction methods stems from the difficulty in separating confounded batch effects from biological signals, and the attempt to remove batch effects without separating the two. That is, for example, the method described in Non-Patent Document 1 uses a common latent space to correct batch effects and reconstruct biological signals, which leads to the disappearance of biological signals. One embodiment of the present invention aims to provide a method that uses a variational autoencoder, which is a generative model, to separate information about confounded batch effects from information about biological signals before correcting batch effects.
[0015] Specifically, one embodiment of the present invention uses a variational autoencoder trained through machine learning to separate latent variables in the compression of cellular information into latent variables that indicate differences due to differences between cells and latent variables that indicate batch effects. This machine learning can be performed by setting up prediction units that predict differences between cells and differences in the measurement environment, respectively. By using this configuration in which latent variables are separated, a batch effect correction method can be realized that minimizes the loss of biological signals while correcting only the latent variables that indicate batch effects.
[0016] <Outline Configuration of Correction System 100> Below, the outline configuration of a batch effect correction system 100 according to one embodiment of the present invention will be described with reference to Figure 1. Figure 1 is a block diagram showing an example of the outline configuration of the correction system 100.
[0017] As shown in Figure 1, the correction system 100 may include a correction device 1 (information processing device) and a display device 2. Figure 1 shows a correction system 100 that includes one correction device 1 and one display device 2. However, the configuration of the correction system 100 is not limited to this. For example, the correction system 100 does not have to include a display device 2, or it may have multiple display devices 2.
[0018] In the correction system 100, the correction device 1 and the display device 2 are connected to each other in a way that allows them to communicate with one another. The correction device 1 and the display device 2 may be connected directly by wire or wireless connection, or they may be connected via a communication network. The type of communication network is not limited and may be a local area network (LAN) or the internet.
[0019] The correction device 1 is a computer that corrects and outputs data containing information on multiple cells that has been input, in order to reduce batch effects. The details of the functions of the correction device 1 will be described later, particularly in the description of the control unit 10.
[0020] The display device 2 may be a computer, smartphone, or tablet terminal used by a user of the correction system 100. The display device 2 may acquire and display data summarizing the correction results of the cell information from the correction device 1.
[0021] Figure 1 shows a correction system 100 in which the display device 2 is separate from the correction device 1. However, the configuration of the correction system 100 is not limited to this. For example, the display device 2 may be an integrated device with the correction device 1, in which case the display device 2 may be a display unit (display, etc.) provided by the correction device 1.
[0022] <Configuration of Correction Device 1> Next, the configuration of the correction device 1 will be described. The correction device 1 comprises a control unit 10, a storage unit 30, and an input unit 40.
[0023] The storage unit 30 may be, for example, an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The storage unit 30 stores the information transmitted from the control unit 10 and may also read the stored information by the control unit 10.
[0024] The input unit 40 is configured such that the user of the correction device 1 inputs information. The input unit 40 may be, for example, at least any one of a keyboard, a mouse, or a touch pad. Further, the input unit 40 may be an interface that receives information input from an external storage device such as a USB memory.
[0025] The control unit 10 is a control device that comprehensively controls each part of the correction device 1. The control unit 10 may be, for example, a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). The control unit 10 includes an acquisition unit 11 and a VAE 20 (Variational Autoencoder). The VAE 20 includes an encoding unit 21, a prediction unit 22, a correction unit 23, and a decoding unit 24.
[0026] The control unit 10 may read a control program, which is software stored in the storage unit 30, and expand it in a memory such as a RAM (Random Access Memory) to execute the functions of each part.
[0027] (Acquisition Unit 11) The acquisition unit 11 acquires input data including cell information. The cell information included in the input data is information about a plurality of cells.
[0028] The input data including cell information is, for example, data indicating the measurement results of cells. The input data is preferably omics data obtained by omics analysis, and more preferably omics data obtained by single-cell analysis. If the data including cell information is in such a form, the correction device 1 can easily acquire a sufficient amount of data for proceeding with machine learning by the VAE 20.
[0029] The types of omics analysis are not particularly limited, but examples include transcriptomics analysis, genomics analysis, epigenomics analysis, proteomics analysis, metabolomics analysis, and lipidomics analysis. Furthermore, multi-omics analysis, which is a combination of these, can also be exemplified as an example of omics analysis as described above.
[0030] The acquisition unit 11 may further acquire various other information used by the correction device 1. The acquisition unit 11 may acquire various information from the storage unit 30. The acquisition unit 11 may also acquire at least a portion of the various information as input information from the input unit 40. The acquisition unit 11 may store the various input information acquired from the input unit 40 in the storage unit 30.
[0031] (Encoding Unit 21) The encoding unit 21 is an encoder that compresses the information of multiple cells contained in the input data into multi-dimensional latent variables for each cell. This compression is also called dimensionality reduction, projection into latent space, or feature vectorization. The encoding unit 21 can be realized by a regression algorithm capable of such compression into latent variables. The regression algorithm is not particularly limited, but examples include machine learning algorithms such as deep neural networks (DNNs).
[0032] The compression performed by the encoding unit 21 is carried out for each cell in the case of information from multiple cells. That is, the input data stores information about the cell, such as measurement data, for each cell. For example, if the input data is data obtained by single-cell RNA-Seq, the input data includes gene expression measurement data for each cell. The encoding unit 21 compresses the data into latent variables so that the gene expression information is retained as features for each cell.
[0033] The multidimensional latent variables include at least two-dimensional latent variables: a first latent variable and a second latent variable. The first latent variable is optimized to vary due to differences in cells. The second latent variable is optimized to vary due to differences in the measurement environment of cells.
[0034] As mentioned above, the measurement data (cellular information) obtained through omics analysis can include both biological signals and batch effects. The "cellular difference" in the first latent variable refers to the difference in biological signals between cells.
[0035] The first latent variable includes at least a condition variable that varies depending on differences in cell conditions. In this specification, "cell conditions" refer to the conditions set for each group when analyzing the differences between different cell groups. These conditions include, for example, the individuals from which the cells originate when comparing cells obtained from different individuals. Examples of different individuals include, but are not limited to, the presence or absence of disease and drug administration in each individual. In this case, the comparison may be between two groups or between three or more groups.
[0036] Furthermore, the first latent variable may include a cell type variable, for example, which varies depending on differences in cell types. In other words, the first latent variable may include a cell type variable (third latent variable) that varies depending on differences in cell types. Cell types can be defined as categories classified by differences in the tissue from which they originate and differences in cell lines, etc.
[0037] The first latent variable may further include variables defined separately from the conditional variable and the cell type variable, insofar as they are variables that vary with respect to differences between cells. For example, if the cell information is single-cell RNA-seq measurement data obtained by transcriptomics analysis, information on the proportion of mitochondrial-derived reads may be used for accuracy control. The first latent variable may further include variables that vary with such a proportion of mitochondrial-derived reads.
[0038] The second latent variable is defined as a latent variable that fluctuates due to batch effects. In other words, the "difference in the measurement environment of cells" in the second latent variable refers to the difference caused by batch effects.
[0039] The first latent variable may be one-dimensional or may contain multiple-dimensional latent variables. In other words, the VAE 20 may use one-dimensional latent variables as the first latent variable from among the multiple-dimensional latent variables compressed by the encoding unit 21, or it may use two or more dimensions as the first latent variable. Similarly, the second latent variable may be one-dimensional or may contain multiple-dimensional latent variables.
[0040] Furthermore, the multi-dimensional latent variables generated by the encoding unit 21 may include other latent variables besides the first and second latent variables. For example, if the encoding unit 21 compresses cell information into 10-dimensional latent variables, a 4-dimensional latent variable may be assigned to the first latent variable and a 5-dimensional latent variable may be assigned to the second latent variable. The remaining 1-dimensional latent variable may, for example, be used as a scale factor variable in the calculation of a scale factor used to normalize the measurement data, if the cell information is single-cell RNA-seq measurement data.
[0041] Furthermore, when using the conditional variable and cell type variable described above as the first latent variable, the conditional variable and cell type variable may each be assigned a one-dimensional latent variable, or multiple-dimensional latent variables. For example, when using a four-dimensional latent variable as the first latent variable, one dimension may be assigned to the conditional variable and three dimensions to the cell type variable. When the number of dimensions of a latent variable is set excessively high, the posterior distribution of the latent variable will have a mean close to 0 and a variance close to 1. In this case, the difference between the posterior distribution of the latent variable and its prior distribution (which follows a normal distribution with a mean of 0 and a variance of 1) becomes small. Such a latent variable cannot be said to appropriately capture the information of the cell before compression. In other words, the number of dimensions of the latent variable can be adjusted from the perspective of whether the latent variable of each dimension appropriately captures the information of the cell. The number of dimensions exemplified above is an example of a suitable number of dimensions from this perspective.
[0042] Furthermore, while a one-dimensional latent variable may be assigned to the second latent variable, it is preferable to assign a multi-dimensional latent variable. The second latent variable is a variable that fluctuates due to differences caused by batch effects, and therefore has a significant impact on the accuracy of batch effect correction. By assigning a latent variable of two or more dimensions to such a second latent variable, the accuracy of batch effect correction by VAE20 can be improved.
[0043] The assignment of multiple-dimensional latent variables is not limited to these, and can be increased or decreased as appropriate based on the final batch effect correction results.
[0044] Furthermore, the latent variables of each dimension are used mutually exclusively in VAE20. That is, latent variables assigned as batch variables are not used as other conditional variables, etc. With this configuration, the VAE20's machine learning process makes it easier for the encoding unit 21 to compress the latent variables so that the batch variables retain features related to the batch. The same applies to other latent variables such as conditional variables.
[0045] The encoding unit 21 maintains the latent variable partitioning pattern through subsequent machine learning and batch effect correction processes. For example, if 10-dimensional latent variables are generated as a 10x1 matrix, the encoding unit 21 performs compression processing so that rows 1 through 5 are used for batch variables, row 6 for condition variables, rows 7 through 9 for cell type variables, and row 10 for scale factor variables. The encoding unit 21 stores the same latent variables for the same purpose at the same position in the matrix for all input cell information. Similarly, the prediction unit 22 always retrieves and uses the corresponding latent variables from the same position in the matrix.
[0046] When the encoding unit 21 compresses cell information into multi-dimensional latent variables, the VAE 20 performs machine learning, including the prediction unit 22 described below, so that the first and second latent variables become latent variables that retain features according to the definition described above. In this specification, the following example will be described when the encoding unit 21 compresses cell information into 10-dimensional latent variables, with 5 dimensions for batch variables, 1 dimension for condition variables, 3 dimensions for cell type variables, and 1 dimension for scale factor variables.
[0047] (Prediction Unit 22) The prediction unit 22 is a functional unit that uses the multi-dimensional latent variables compressed by the encoding unit 21 individually to predict the batch, conditions, and cell type for each cell. The prediction unit 22 has a batch prediction unit 22a (first prediction unit), and may further have a condition prediction unit 22b (second prediction unit) and a cell type prediction unit 22c (third prediction unit). The prediction unit 22 can be implemented by a regression algorithm that uses latent variables as explanatory variables and the target of prediction, such as a batch, as the dependent variable. The regression algorithm is not particularly limited, but examples include machine learning algorithms such as deep neural networks (DNNs).
[0048] The batch prediction unit 22a predicts the batch size for each cell using batch variables. In other words, the batch prediction unit 22a predicts the differences in the measurement environment of cells using batch variables. Here, the input data for cell information includes first label information indicating differences in the measurement environment as training data. The first label information is batch information for each cell. The batch prediction unit 22a is trained using learning data in which the batch variables, which are explanatory variables, and the first label information, which is the target variable, are associated for each cell, and has the function of predicting differences in the measurement environment (batch) from the input batch variables.
[0049] The condition prediction unit 22b predicts differences in conditions for each cell using condition variables. Here, the input data for cell information includes second label information indicating differences between cells as training data. The condition prediction unit 22b is trained using training data in which the explanatory variable (condition variable) and the target variable (second label information) are associated for each cell, and has the function of predicting differences in cell conditions from the input condition variables. Here, the condition prediction unit 22b performs the prediction using information indicating differences in cell conditions included in the second label information.
[0050] The cell type prediction unit 22c uses a cell type variable to predict the difference in cell type for each cell. Here, the input data for cell information includes a third label information indicating the difference in cell type as training data. The third label information is information included in the second label information. The cell type prediction unit 22c is trained using learning data in which the cell type variable, which is an explanatory piece, and the third label information, which is the target variable, are associated for each cell, and has the function of predicting the difference in cell type from the input of the cell type variable.
[0051] Each prediction result from the prediction unit 22 is incorporated as a prediction error into the loss function described later and used for machine learning of the VAE 20. As a result, the encoding unit 21 is trained to perform compression into latent variables so that the batch variable, condition variable, and cell type variable retain features according to their respective definitions.
[0052] To improve the accuracy of batch effect correction, it is important that the encoding unit 21 can accurately compress feature quantities indicating differences in the measurement environment from cell information into batch variables, and compress other cell-specific condition information into condition variables. For this purpose, it is preferable that the first label information and second label information are extracted from input data containing information on as many cells as possible. In other words, the accuracy of batch effect correction by VAE 20 depends on the amount and accuracy of cell information included in the input data. The accuracy of cell information refers to, for example, the accuracy of the combination of cells and the tag information assigned to them. In this respect, it is preferable that the input data containing cell information is omics data, as described above.
[0053] These prediction units 22 function in the machine learning of the VAE 20, but are not used in the batch effect correction process performed by the VAE 20.
[0054] (Correction Unit 23) The correction unit 23 is a functional unit that corrects batch variables so that the differences in batch variables between cells are reduced. Specifically, the correction unit 23 may equalize the values of the batch variables to a value such as 0, or it may multiply them by a coefficient less than 1, such as 0.1, so that the batch variables become smaller. If the configuration is such that the batch variables are equalized to a predetermined value, the VAE 20 can effectively correct or remove the batch effect from the cell information. Also, if the configuration is such that the batch variables are multiplied by a predetermined coefficient, the VAE 20 can reduce the loss of feature quantities used for reconstruction by the decoding unit 24 while correcting the batch effect.
[0055] The correction unit 23 functions in the batch effect correction process performed by the VAE 20, but is not used in the machine learning of the VAE 20.
[0056] (Decoding Unit 24) The decoding unit 24 is a decoder that reconstructs cell information using latent variables corrected for batch variables. The decoding unit 24 can be realized by a regression algorithm capable of reconstruction using such latent variables. The regression algorithm is not particularly limited, but examples include machine learning algorithms such as deep neural networks (DNNs). Furthermore, it is preferable from the viewpoint of reconstruction accuracy that the regression algorithm constituting the decoding unit 24 corresponds to the regression algorithm constituting the encoding unit 21.
[0057] The latent variables after correcting the batch variables may include all 10-dimensional latent variables, including the batch variables corrected by the correction unit 23. Alternatively, the decoding unit 24 may be configured to reconstruct cell information using only latent variables other than the batch variables. In this case, the correction by the correction unit 23 is performed by removing the batch variables from the 10-dimensional latent variables, resulting in 5-dimensional latent variables.
[0058] <Machine Learning Method and Correction Method> A batch effect correction method according to one embodiment of the present invention will be explained using the case in which the correction device 1 is made to execute the correction method as an example. First, an example of a machine learning method executed by the correction device 1 will be described with reference to Figure 2, and then an example of a batch effect correction method executed by the correction device 1 will be described with reference to Figure 3. Note that the contents that have already been explained in the above-mentioned section on the correction system 100 will be omitted here.
[0059] (Machine Learning Method) Figure 2 is a flowchart showing an example of the machine learning flow performed by the correction device 1. Figure 2 also shows the processing flow performed by the correction system 100 equipped with the correction device 1.
[0060] First, the acquisition unit 11 acquires information about the sample cells (S11). In this specification, the information about cells used for machine learning of the VAE 20 may be referred to as "information about the sample cells." Here, the information about the sample cells will be explained using the example of measurement data (count values) obtained in transcriptomics analysis using single-cell RNA-seq. The acquisition unit 11 may output at least a portion of the input measurement data to the encoding unit 21 as training data and the remainder as verification data.
[0061] The encoding unit 21 compresses the information of sample cells contained in the training data into a 10-dimensional latent variable. Of the 10-dimensional latent variable, 5 dimensions are used as batch variables, 1 dimension as a condition variable, 3 dimensions as a cell type variable, and 1 dimension as a scale factor variable.
[0062] The batch prediction unit 22a predicts the batch using a five-dimensional batch variable (S3). The condition prediction unit 22b predicts the cell conditions using a one-dimensional condition variable (S4). The cell type prediction unit 22c predicts the cell type using a three-dimensional cell type variable (S5).
[0063] The decoding unit 24 calculates the scale factor from the scale factor variables (S6). The decoding unit 24 reconstructs the information of the sample cells using the 9-dimensional latent variables, which have been normalized by removing the scale factor variables (S7).
[0064] The VAE 20 calculates a loss function based on the prediction results from the prediction unit 22 and the reconstruction results from the decoding unit 24 (S8). The loss function calculated here includes a reconstruction loss function and a regularization loss function used in machine learning of a typical variational autoencoder, as well as a prediction loss function calculated from the prediction error of the prediction unit 22. The prediction error of the prediction unit 22 is the error between the prediction result of each prediction unit 22 and the first label information, second label information, or third label information. The reconstruction loss function and the regularization loss function can be calculated by conventionally known methods.
[0065] The prediction loss function may be calculated as the sum of the prediction errors of each prediction unit 22. Each prediction error of each prediction unit 22 can be calculated by a conventionally known method, for example, as the cross-entropy error.
[0066] The loss function calculated by VAE20 can be calculated, for example, by summing the above-mentioned reconstruction loss function, regularization loss function, and prediction loss function. VAE20 is trained by optimizing each parameter in the encoding unit 21, prediction unit 22, and decoding unit 24 to minimize the calculated loss function (S9).
[0067] In other words, the encoding unit 21 is trained using training data that associates information about sample cells with first label information indicating differences in the measurement environment (batch) of the sample cells, so that batch variables that minimize the prediction error of the batch prediction unit 22a can be obtained by compression. The encoding unit 21 is also trained using training data that associates information about sample cells with second label information including differences in the conditions of the sample cells, so that condition variables that minimize the prediction error of the condition prediction unit 22b can be obtained by compression. Furthermore, the encoding unit 21 is also trained using training data that associates information about sample cells with third label information indicating differences in the types of sample cells, so that cell type variables that minimize the prediction error of the cell type prediction unit 22c can be obtained by compression.
[0068] The information of sample cells includes first label information and second label information. Third label information is an example of the information included in second label information. Generally, cell measurement data such as omics data includes metadata in addition to the measurement results, such as information about the type of cells being measured (second label information) and batch information such as measurement conditions (first label information). First and second label information can be obtained from this metadata included in the sample cell information.
[0069] In this manner, the encoding unit 21 is trained by machine learning to generate latent variables that optimize not only the reconstruction results by the decoding unit 24 but also the prediction results of the prediction unit 22. Specifically, the encoding unit 21 is trained by machine learning to compress the latent variables of each dimension assigned to each prediction unit 22 so that they retain features favorable to the predictions of each prediction unit 22.
[0070] Furthermore, the encoding unit 21 is machine-learned to generate latent variables so that features related to conditions other than batch and cell type are preferentially included in the condition variables and cell type variables. This is because the machine learning described above is performed using the condition prediction unit 22b and the cell type prediction unit 22c in addition to the batch prediction unit 22a. Therefore, even if the batch variables are corrected in the batch effect correction described later, the cell conditions and cell types can be accurately reconstructed by the decoding unit 24.
[0071] (Batch Effect Correction Method) Figure 3 is a flowchart showing an example of the batch effect correction method performed by the correction device 1. Figure 3 also shows the processing flow performed by the correction system 100 equipped with the correction device 1.
[0072] The acquisition unit 11 acquires cell information (S11). It is preferable that after the VAE 20 executes the machine learning method described above, it uses the information of the sample cells used for machine learning as input data again to correct for batch effects. At least a portion of the information of the sample cells and the information of the cells to be corrected for batch effects are the same. With this configuration, batch effect correction can be performed on input data having the same batch, conditions and cell type as the data used for machine learning by the VAE 20. Therefore, the VAE 20 can correct for batch effects and reconstruct the information of conditions and cell types with high accuracy. In this case, the acquisition unit 11 may skip the process in S11.
[0073] Furthermore, machine learning of VAE20 and batch effect correction using a trained VAE20 may be performed using different input data. In this case, it is preferable that the training data used for machine learning and the input data used for batch effect correction be as similar as possible in terms of batch, conditions, and cell type. An example of this is when performing transcriptomics analysis using cells derived from patients with the same disease over a long period of time, and using a VAE20 that was previously trained and used to correct batch effects for newly acquired measurement data.
[0074] Next, the encoding unit 21 compresses the cell information contained in the input data into 10-dimensional latent variables (S12). The correction unit 23 corrects the batch variables among the 10-dimensional latent variables (S13). This correction is performed, for example, by uniformly setting the values of all batch variables to 0.
[0075] The decoding unit 24 calculates the scale factor from the scale factor variables (S14). The decoding unit 24 reconstructs the cell information using a nine-dimensional latent variable that includes the corrected batch variable, which has been normalized by removing the scale factor variables (S15).
[0076] Through the above processing, VAE20 can generate cell information that accurately preserves cell conditions and cell type information while correcting for batch effects.
[0077] [Embodiment 2] Another embodiment of the present invention will be described below. For the sake of convenience of explanation, components having the same function as those described in the above embodiment will be denoted by the same reference numerals, and their descriptions will not be repeated.
[0078] The correction system 100 shown in Figure 1 is configured such that the correction device 1 includes an input unit 40 that accepts cell information input from the user and outputs the correction results to the display device 2, but it is not limited to this configuration. For example, as shown in Figure 4, the correction system 100a according to this embodiment may include a correction device 1a that is connected to communication terminals 5a and 5b used by each user via a communication network 9.
[0079] In the correction system 100a shown in Figure 4, the correction device 1a receives data containing cell information from each of the communication terminals 5a and 5b. The correction device 1a then corrects the batch effect from the data containing cell information received from the communication terminals 5a and 5b and transmits the reconstructed data to the communication terminals 5a and 5b.
[0080] Figure 4 shows a correction system 100a including communication terminals 5a and 5b and a correction device 1a, but is not limited to this. In the correction system 100a, the correction device 1a may be able to communicate with, for example, three or more communication terminals.
[0081] (Configuration of Correction Device 1a) The configuration of the correction device 1a will be explained with reference to Figure 5. Figure 5 is a functional block diagram showing an example of the configuration of a correction system 100a according to one embodiment of the present invention. As shown in Figure 5, the correction device 1a includes a communication unit 50 that functions as a communication interface with communication terminals 5a and 5b. The acquisition unit 11 receives data including cell information via the communication unit 50.
[0082] The correction device 1a transmits the reconstructed data, corrected for batch effects, to the communication terminals 5a and 5b via the communication unit 50. The correction device 1a may also generate a web page on which the reconstructed data corresponding to the received information can be downloaded, and provide the user who sent the information with information to access the web page.
[0083] [Example of implementation by software] The function of the correction device 1 (hereinafter referred to as "device") can be implemented by a program that causes a computer to function as the device, and by a program that causes a computer to function as each control block of the device (particularly each part included in the control unit 10).
[0084] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., memory) as hardware for executing the program. By executing the program using this control device and storage device, the functions described in each of the embodiments are realized.
[0085] The above program may be recorded on one or more computer-readable recording media, not temporary ones. These recording media may or may not be provided by the above device. In the latter case, the program may be supplied to the above device via any wired or wireless transmission medium.
[0086] Furthermore, some or all of the functions of each of the above control blocks can also be realized by logic circuits. For example, an integrated circuit in which logic circuits functioning as each of the above control blocks are formed is also included in the scope of the present invention. In addition, it is also possible to realize the functions of each of the above control blocks by, for example, a quantum computer.
[0087] Furthermore, each process described in the above embodiments may be performed by AI (Artificial Intelligence). In this case, the AI may operate on the control device described above, or it may operate on other devices (for example, an edge computer or a cloud server).
[0088] [Summary] The batch effect correction system according to embodiment 1 of the present invention comprises an information processing device having: an encoding unit that compresses information of a plurality of input cells into a multidimensional latent variable for each cell, which includes a first latent variable that varies due to differences in the cells and a second latent variable that varies due to differences in the measurement environment of the cells; a correction unit that corrects the latent variable so that the difference in the second latent variable between the cells becomes small; and a decoding unit that reconstructs the cell information using the corrected latent variable.
[0089] A batch effect correction system according to embodiment 2 of the present invention further comprises, in embodiment 1, an information processing device which further includes a first prediction unit which predicts the difference in the measurement environment using the second latent variable, and the encoding unit which is machine-learned using training data in which information of a plurality of sample cells and first label information indicating the difference in the measurement environment of the sample cells are associated, so that the second latent variable which minimizes the prediction error between the prediction result of the first prediction unit and the first label information is obtained by compression.
[0090] A batch effect correction system according to embodiment 3 of the present invention, in embodiment 1 or 2, further comprises an information processing device which further comprises a second prediction unit which predicts the differences between the cells using the first latent variable, and the encoding unit which is machine-learned using training data in which information of a plurality of sample cells and second label information indicating the differences between the sample cells are associated, such that the first latent variable which minimizes the prediction error between the prediction result of the second prediction unit and the second label information is obtained by compression.
[0091] A batch effect correction system according to aspect 4 of the present invention, in any of aspects 1 to 3, wherein the first latent variable includes a third latent variable that varies depending on the difference in the type of cell, the second label information includes a third label information indicating the difference in the type of cell, the information processing device further includes a third prediction unit that predicts the difference in the type of cell, and the encoding unit may be machine-trained to obtain the third latent variable that minimizes the prediction error between the prediction result of the third prediction unit and the third label information by compression, using training data to which information of a plurality of sample cells and the third label information are associated.
[0092] In the batch effect correction system according to embodiment 5 of the present invention, the cell information may be information obtained by single-cell analysis in any of embodiments 1 to 4.
[0093] In the batch effect correction system according to embodiment 6 of the present invention, in any of embodiments 1 to 5, the cell information may be information obtained by transcriptomics analysis.
[0094] A control program according to aspect 7 of the present invention is a control program for causing a computer to function as an information processing device as described in aspect 1, wherein the computer functions as the encoding unit, the correction unit, and the decoding unit.
[0095] A batch effect correction method according to aspect 8 of the present invention is a method performed by a computer and includes the steps of: compressing information of a plurality of input cells into a multidimensional latent variable for each cell, which includes a first latent variable that varies due to differences in the cells and a second latent variable that varies due to differences in the measurement environment of the cells; correcting the latent variable so that the difference in the second latent variable between the cells becomes small; and reconstructing the cell information using the corrected latent variable.
[0096] [Additional Notes] The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
[0097] One embodiment of the present invention is described below.
[0098] [1. Generation of Test Data] Test data for use in testing the correction system according to one embodiment of the present invention was created using a Python program that modified the method described in Splatter (Zappia et al., 2017). The test data was created assuming count data obtained by transcriptomics analysis.
[0099] (Assumed Settings) Three cell types, A, B, and C, were assumed to exist, and 1000 cells were assumed to be obtained as measurement samples from each type. Each cell type was assumed to specifically express different marker genes. Hereafter, the types of marker genes in each cell type will be referred to as mA, mB, and mC.
[0100] The study assumed two conditions: 'WT' (healthy individuals without disease) and 'Disease' (patients with disease). For each cell type, the gene types showing lower expression in 'Disease' compared to 'WT' were designated as ddA, ddB, and ddC, respectively. Conversely, the gene types showing higher expression in 'Disease' compared to 'WT' were designated as duA, duB, and duC, respectively.
[0101] For the sake of simplicity, it was assumed that there are no common genes among mA, mB, mC, ddA, ddB, ddC, duA, duB, and duC. Hereafter, mA, mB, mC, ddA, ddB, ddC, duA, duB, and duC will be collectively referred to as variable genes. For the 15,000 genes excluding the bottom 5,000 genes based on the basic average expression levels described below, gene types were assigned using a multinomial distribution such that mA, mB, and mC each account for 4%, and ddA, ddB, ddC, duA, duB, and duC each account for 1%. The final number of genes is shown in Table 1 below.
[0102]
[0103] Furthermore, three batches were created for both 'WT' and 'Disease', and it was assumed that a batch effect would be applied to approximately 2% of the genes randomly. A control dataset without the batch effect was also created under the same conditions.
[0104] (Creation of Basic Average Expression Level) First, the basic average expression levels of 20,000 genes were sampled from a gamma distribution (shape=0.6, rate=0.3). Genes were sampled using a binomial distribution with a mean of 0.05, and for each gene i, the basic average expression level was determined by the expression level g determined by the following formula (1). i,out By replacing it with this, the gene was made to take the average expression level of outliers.
[0105] However, here g median μ is the median of the basic mean expression levels of all genes before substitution. out = 4, σ out We set it to = 0.5.
[0106] (Variation in expression levels dependent on cell type) For variable gene i, the variation in expression levels dependent on cell type was sampled using the following formula (2).
[0107] ν type,i When < 1.1, the gene expression level was replaced with 1.1. Here, μ m = 2, σ m = 0.5.
[0108] When gene i is mX or ddX (X indicates cell type, either A, B, or C), the expression level in cell type X was multiplied by ν type,i times, and the expression levels in other cell types were multiplied by 1 / ν type,i times. When gene i is duX (X indicates cell type, either A, B, or C), the average expression level in all cells was multiplied by 1 / ν type,i times.
[0109] (Variation in expression levels dependent on conditions) For condition-dependent genes (ddA, ddB, ddC, duA, duB, duC), the variation in expression levels dependent on conditions was sampled using the following formula (3).
[0110] ν disease,i When < 1.1, the gene expression level was replaced with 1.1. Here, μ d = 2, σ d = 0.5.
[0111] When gene i is duX (X indicates cell type, either A, B, or C), the expression level in cell type X under the 'Disease' condition was multiplied by ν disease,i times, and when gene i is ddX (X indicates cell type, either A, B, or C), the average expression level in cell type X under the 'Disease' condition was multiplied by 1 / ν disease,i times.
[0112] (Adding batch effect) In batch b, for the expression level of gene i, c b,i was randomly set to 0 with probability 0.8, 1 with probability 0.1, and -1 with probability 0.1. Furthermore, the magnitude of the batch effect was sampled using the following formula (4).
[0113] ν batch,i <In the case of 1.1, the gene expression level was replaced with 1.1. Note that μ b = ln2, σ b = 0.12. Using these values, the average expression level of gene i was calculated using ν batch,i cb, i I doubled it.
[0114] (Setting fluctuations for each cell) The average expression level of gene i in cell type X under condition D, obtained by applying the fluctuations set as described above, is g i D,X Let's assume that, under condition D, the library size of cells k belonging to cell type X is l. k The sample was then sampled using the following formula (5).
[0115] Note that μ l =ln(5000),σ l = 0.2. Using this value, the expression level g, taking into account the library size for each cell, was calculated. i,k D,X This was calculated using the following formula (6).
[0116] Furthermore, cell-specific fluctuations are introduced using the following formula (7), and the true expression level of each cell, gT, is determined. i,k D,X I obtained it.
[0117] Here, df 0 The values were set to 60 and Φ to 0.1. This process was performed for each batch and cell type.
[0118] (Simulation using single-cell RNA-seq) Based on the gene expression levels set as described above, a simulation using single-cell RNA-seq was performed. The count value of gene i in cell k belonging to condition D and cell type X was sampled as shown in equation (8) below.
[0119] Normally, in single-cell RNA-seq, a dropout occurs where the gene count value becomes 0. Under condition D, in cell k belonging to cell type X, the probability of gene i dropping out is π. i,kCalculate y using the formula (9) below, and with that probability y i,k Replaced with 0.
[0120] Here, α = -1, x 0 We set it to = 0.
[0121] [2. Construction of a Variational Autoencoder According to an Embodiment] A variational autoencoder according to one embodiment of the present invention was constructed using Python.
[0122] (Encoding section) The encoding section was implemented as an encoder based on a Deep Neural Network (DNN). Using count value vector data of 20,000 genes as input, the DNN was used to obtain the mean μ and log variance ln(σ) of the probability distribution of the 10-dimensional latent variables. 2 The model calculates the latent variables according to a normal distribution and uses this to sample them. The DNN portion consists of three layers, each stacked with a linear coupling layer, a batch normalization layer, a ReLU function layer, and a dropout layer, with the number of nodes in each layer being 2000, 400, and 200, respectively. For the latent variable sampling portion, calculations were performed for each latent variable (each dimension) as shown in equation (1) below.
[0123]
[0124] (Partitioning of Latent Space) For a 10-dimensional latent variable, one dimension was assigned to the condition prediction unit, three dimensions to the cell type prediction unit, and five dimensions to the batch prediction unit. The remaining one dimension was used as a scale factor variable and was not assigned to the prediction unit.
[0125] (Prediction Unit) Each prediction unit was implemented by combining a linear coupling layer with a softmax activation layer.
[0126] (Decoding Unit) The decoding unit was implemented as a decoder based on a Deep Neural Network (DNN). The decoding unit calculates the parameters of a zero-expanded negative binomial distribution (ZINB) from the latent variables, and reconstructs the count value of single-cell RNA-seq by sampling from the distribution determined by those parameters.
[0127] The scale factor s was calculated by transforming the constant multiple of the aforementioned scaling factor variable using an exponential function. From the nine-dimensional latent variables excluding the scaling factor variable, values corresponding to the expression level and zero-expansion probability of each gene were obtained using a DNN, and these were scaled using the softmax function and softplus function, respectively, to obtain the relative expression level ρ and zero-expansion probability π. 0 The following was obtained. The DNN was constructed by combining a linear coupling layer, a batch normalization layer, a ReLU function layer, and a dropout layer in that order, with three layers stacked on top of each other, and the number of nodes in the three layers being 200, 400, and 2000, respectively.
[0128] Furthermore, the variance control parameter θ was calculated by applying a soft plus function to the linear sum of the nine-dimensional latent variables, excluding the variable used for the scale factor. From these values, the mean was defined as the product of the relative expression level ρ and the scale factor s, ρs as the variance control parameter, and π as the zero expansion probability. 0 A zero-expansion negative binomial distribution was created, and the count values were reconstructed.
[0129] [3. Machine Learning of Variational Autoencoder According to the Example] Machine learning of the variational autoencoder was performed using a portion of the test data as training data to minimize the loss function shown below.
[0130] (Loss Function) In machine learning, the loss function shown in equation (11) below was used.
[0131]
[0132] Here, L reconst L is the reconstruction loss function, and is defined as the log-likelihood of the zero-expanded negative binomial distribution in the decoding section for the input to the variational autoencoder. KL L is the regularization loss function, which is the Kullback-Leibler divergence between the posterior distribution of the latent variable calculated from the input to the variational autoencoder and the prior distribution of the latent variable, a normal distribution with mean 0 and variance 1. predThis is the prediction loss function, which is calculated by summing the prediction errors in each prediction unit using cross-entropy. The hyperparameters were set to β = 5 and γ = 100.
[0133] (Machine Learning Execution) The test data was randomly split into training data and confirmation data in a 2:1 ratio, and machine learning was performed using the training data. The Adam machine learning algorithm was used, with a learning rate of 0.001. The machine learning batch size was 128, and the number of epochs was 25. The confirmation data was used only to check the progress of the machine learning.
[0134] [4. Correction of batch effects using variational autoencoders] We attempted to correct batch effects using test data and a pre-trained variational autoencoder.
[0135] (Example 1: Correction of batch effect) Using the trained encoding unit, all test data was compressed again into 10-dimensional latent variables. Of the obtained latent variables, only the batch variables, which were used for batch prediction by the batch prediction unit, were replaced with 0 to correct the latent variables. The corrected latent variables were input into the decoding unit to obtain the parameters of a zero-expanded negative binomial distribution. The scale factor calculated based on the scale factor latent variable was adjusted to the median, and then the count value was reconstructed from the zero-expanded negative binomial distribution.
[0136] Seurat v4 (Hao et al., 2021) was used to analyze the reconstructed data. The reconstructed count values using the corrected latent variables were loaded into a Seurat object. Highly variable genes (HVGs) were detected using the FindVariableFeatures function. Gene expression differential analysis was also performed using the FindMarkers function. In this analysis, only the parameter 'test.use=MAST' was specified, and the other parameters were left as default parameters. The number of genes with pval_adj (Bonferroni correction) < 0.05 was counted for each gene type.
[0137] (Comparative Examples 1 and 2: Cases without batch effect correction) Seurat v4 was used to analyze the test data. After loading the count values of the test data into the Seurat object, normalization and HVG detection were performed using the SCTransform function. The number of HVGs was set to 3000 genes. Using the HVGs, Principal component analysis (PCA) was performed using the RunPCA function to compress the data into 20-dimensional latent variables. Furthermore, gene expression variation analysis was performed using the FindMarkers function on the values corrected by the SCTransform function. At this time, only the parameter 'test.use=MAST' was specified, and the other parameters were left as default parameters, and the number of genes for which pval_adj (Bonferroni correction) < 0.05 was counted for each gene type. These analyses were performed on both the data with and without batch effect (Comparative Example 1).
[0138] (Comparative Example 3: Correction of Batch Effect using Conventional Harmony Technology) Seurat v4 was used to analyze the test data. After loading the count values of the test data into a Seurat object, the objects were separated by batch, and the SCTransform function was applied to each to normalize them. Furthermore, the FindIntegrationFeatures function was applied to the separated object groups to obtain 3000 HVGs. The separated objects were reintegrated, and PCA was performed using the RunPCA function with the HVGs to compress them into 20-dimensional latent variables. The RunHarmony function from the Harmony package (Korsunsky et al., 2019) was applied to this PCA data to correct for batch effects.
[0139] (Calculation of LISI) The Local Invers Simpson Index (LISI) was calculated using the lisi package (Korsunsky et al., 2019). LISI is a data diversity index, and the more diverse the data, the larger the value. For each comparative example described above, LISI was calculated for each cell for the PCA values of the data without batch effects, the PCA values without batch effect correction, and the PCA values after batch effect correction using Harmony. In addition, LISI was calculated for each cell for the conditional variable, cell type variable, and batch variable among the latent variables obtained by the variational autoencoder in the example. The average values of LISI obtained in the example and each comparative example were compared.
[0140] (Evaluation of Results) The reconstruction results for Example 1 were evaluated. First, the latent variables compressed by the encoding unit after training were analyzed. Figure 6 shows the conditional variables for Example 1, with the left figure showing the distribution of values for each condition and the right figure showing the distribution of values for each batch. As shown in Figure 6, the distribution of values for the conditional variables was separated for each condition, while the distribution of values for each batch under the same condition was not separated, and the values were mixed. In other words, it was shown that the conditional variables retained features related to the cell conditions independently of the batch.
[0141] Figure 7 shows the distribution of the cell type variable in Example 1. The upper figure shows the distribution per batch, and the lower figure shows the distribution per cell type. As shown in Figure 7, the distribution of values for the cell type variable was separated by cell type, while the distribution of values per batch was not separated, and the values were mixed. In other words, it was shown that the cell type variable retains features related to cell type independently of the batch.
[0142] Figure 8 shows the batch distribution of the batch variable in Example 1. As shown in Figure 8, the distribution of values for the batch variable was separated for each batch. In other words, it was shown that the batch variable holds features related to the batch.
[0143] Next, the average values of LISI calculated in Example 1 and Comparative Examples 1 to 3 are shown in Table 2 below.
[0144]
[0145] As shown in Table 2, in Example 1, the LISI values were small for the condition variable for the condition, for the cell type variable for the cell type, and for the batch variable for the batch, indicating almost no mixing of the value distributions. On the other hand, the LISI values were large for the batch for the condition variable and cell type variable, indicating that the value distributions were well mixed for the batch. This indicates that in Example 1, the features related to biological signals were appropriately compressed as condition variables and cell type variables, rather than as batch variables. Even after correcting for such batch variables, the loss of biological signals included in the overall latent variables can be minimized. In other words, it is suggested that by correcting for batch variables and then performing reconstruction by the decoding unit mainly using the condition variables and cell type variables, condition and cell type information can be accurately reconstructed while correcting for batch effects.
[0146] On the other hand, as shown in Comparative Example 3, when batch effect correction was applied to the PCA data related to Comparative Example 1 using the conventional Harmony method, it was shown that the batch was mixed and the effect was reduced, but it was also shown that mixing occurred in the cell conditions. This problem did not occur with the condition variables in Example 1.
[0147] Next, the reconstructed data obtained in Example 1 was evaluated. Figure 9 shows the correlation between the actual gene expression level and the reconstructed gene expression level in Example 1. As shown in Figure 9, in cell type A, the reconstructed gene expression level (Mean corrected count) correlated well with the actually measured gene expression level (True mean) under both 'WT' and 'Disease' conditions. This was true for variable genes and other genes as well. In other words, it was demonstrated that the variational autoencoder according to one embodiment of the present invention could accurately reconstruct gene expression levels.
[0148] Furthermore, the extent to which variable genes could be identified was evaluated based on the results of Example 1 and Comparative Examples 1 and 2. Cell marker genes (mA, mB, mC) and condition-dependent gene expression levels (ddA, ddB, ddC, duA, duB, duC) were used for evaluation. The frequency of misidentification of genes other than these variable genes (non-variables) was also evaluated. Table 3 below shows the number of genes that Seurat picked as HVGs in Example 1 and Comparative Examples 1 and 2.
[0149]
[0150] As shown in Table 3, Example 1 showed that more variable genes were found compared to Comparative Examples 1 and 2. Furthermore, Example 1 showed that fewer non-variable genes were mistakenly identified as variable genes compared to Comparative Examples 1 and 2. These results suggest that batch effects were effectively corrected in Example 1.
[0151] Figure 10 shows that in Example 1, the dimensionality reduction space before and after correction corrects the differences between batches while simultaneously preserving the differences between conditions and cell types. In Figure 10, "Corrected" shows the dimensionality reduction space using the reconstructed gene expression levels in Example 1, and "Uncorrected" shows the dimensionality reduction space using the uncorrected gene expression levels of the batch variables in Example 1. As shown in Figure 10, in "Uncorrected," batch-by-batch separation is observed under the same conditions or cell types, whereas in "Corrected," batches are mixed under the same conditions or cell types.
[0152] Furthermore, the results of the gene expression change analysis in Example 1 and Comparative Examples 1 and 2 are shown in Table 4 below. In Table 4, "TP" indicates True Positive and "FP" indicates False Positive.
[0153]
[0154] As shown in Table 4, the proportion of FP was reduced in Example 1 compared to Comparative Example 1. In other words, even though batch effect correction was performed in Example 1, the reconstruction accuracy was improved compared to Comparative Example 1.
[0155] 1, 1a Correction device (information processing device) 21 Encoding unit 22 Prediction unit 22a Batch prediction unit (first prediction unit) 22b Condition prediction unit (second prediction unit) 22c Cell type prediction unit (third prediction unit) 23 Correction unit 24 Decoding unit 100, 100a Correction system
Claims
1. An information processing device comprising: an encoding unit that compresses information of multiple input cells into a multi-dimensional latent variable for each cell, including a first latent variable that varies due to differences in the cells and a second latent variable that varies due to differences in the measurement environment of the cells; a correction unit that corrects the second latent variable so that the difference in the second latent variable between the cells becomes small; a decoding unit that reconstructs the cell information using the corrected second latent variable; a first prediction unit that predicts the differences in the measurement environment using the second latent variable; and a second prediction unit that predicts the differences between the cells using the first latent variable. The encoding unit is machine-learned using training data in which information of multiple sample cells and first label information indicating differences in the measurement environment of the sample cells are associated, so that the second latent variable that minimizes the prediction error between the prediction result of the first prediction unit and the first label information can be obtained by the compression, and the batch effect correction system is machine-learned using training data in which information of multiple sample cells and second label information indicating differences between the sample cells are associated, so that the first latent variable that minimizes the prediction error between the prediction result of the second prediction unit and the second label information can be obtained by the compression.
2. The batch effect correction system according to claim 1, wherein the first latent variable includes a third latent variable that varies depending on the difference in the type of cell, the second label information includes a third label information indicating the difference in the type of cell, the information processing device further includes a third prediction unit that predicts the difference in the type of cell, and the encoding unit is machine-trained to obtain the third latent variable that minimizes the prediction error between the prediction result of the third prediction unit and the third label information by compression, using training data to which information of a plurality of sample cells and the third label information are associated.
3. The batch effect correction system according to claim 1, wherein the information of the cells is information obtained by single-cell analysis.
4. The batch effect correction system according to any one of claims 1 to 3, wherein the information of the cells is information obtained by transcriptomics analysis.
5. A control program for causing a computer to function as an information processing device according to claim 1, wherein the computer functions as the encoding unit, the correction unit, the decoding unit, the first prediction unit, and the second prediction unit.
6. The process includes: compressing the input information of multiple cells into a multi-dimensional latent variable for each cell, which includes a first latent variable that varies due to differences in the cells and a second latent variable that varies due to differences in the measurement environment of the cells; correcting the second latent variable so that the difference in the second latent variable between the cells is reduced; reconstructing the cell information using the corrected second latent variable; predicting the differences in the measurement environment using the second latent variable; and predicting the differences between the cells using the first latent variable. A computer-based batch effect correction method, comprising the following steps: in the compression step, a training data system is used in which information on multiple sample cells is associated with first label information indicating differences in the measurement environment of the sample cells. The training system is machine-learned to obtain a second latent variable that minimizes the prediction error between the prediction result of the step for predicting differences in the measurement environment and the first label information through the compression; and the training data system is used in which information on multiple sample cells is associated with second label information indicating differences in the sample cells. The training system is machine-learned to obtain a first latent variable that minimizes the prediction error between the prediction result of the step for predicting differences in the cells and the second label information through the compression.
Citation Information
Patent Citations
A pre-processing system for harmonizing heterogeneous omics datasets by addressing bias and / or batch effects
JP2024529403A
Systems and Methods for Homogenization of Disparate Datasets
US20220101952A1