Single-cell transcriptome batch correction method based on mutual nearest neighbors

By using the autoencoder model for feature extraction and preliminary correction, and searching for MNN pairs in low-dimensional space, the problems of poor batch correction effect and high time complexity in the prior art are solved, and more efficient and accurate batch correction effect is achieved.

CN116312792BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310372606.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-10
Publication Date
2025-07-01
Estimated Expiration
2043-04-10

AI Technical Summary

Technical Problem

The existing single-cell transcriptome batch correction methods have shortcomings in improving correction effect and reducing time complexity, especially in terms of the efficiency and accuracy of searching MNN pairs.

Method used

The autoencoder model was used to perform feature extraction and preliminary batch correction on single-cell transcriptome data, and MNN pairs were searched for in the low-dimensional embedding space, and the calculated correction vectors were iteratively corrected using the searched MNN.

Benefits of technology

It improves the accuracy and number of MNN pairs, enhances the effect of batch correction, reduces time complexity, and can efficiently process large-scale data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312792B_ABST
    Figure CN116312792B_ABST
Patent Text Reader

Abstract

The present invention discloses a mutual nearest neighbor-based single-cell transcriptome batch correction method, which mainly solves the problems of few MNN pairs found by the existing mutual nearest neighbor-based batch correction method and poor correction effect. The implementation solution is as follows: screening the gene features of single-cell transcriptome data to select highly expressed genes; constructing an autoencoder composed of an encoder and a decoder, and using the highly expressed genes selected from the single-cell transcriptome data to cross-train the encoder and the decoder; using the trained autoencoder to extract the features of the highly expressed genes selected from the single-cell transcriptome data to obtain single-cell low-dimensional embedded batch-free information data; searching for mutual nearest neighbor pairs in the single-cell low-dimensional embedded batch-free information data, and calculating a correction vector using the same; using the correction vector to perform batch correction on the data; the present invention finds many MNN pairs and has a good batch correction effect, and can be used for the preprocessing of single-cell transcriptome data in bioinformatics experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data mining, and particularly relates to a method for correcting batches of single-cell transcriptomes, which can be used for preprocessing single-cell transcriptome data in bioinformatics experiments. Background Art

[0002] With the development of single-cell sequencing technology and the decline in sequencing costs, more and more single-cell data are generated. In bioinformatics, the integration of multi-source scRNA-seq data sets is crucial for interpreting the heterogeneity and interactions between cells in complex biological systems. However, there are often batch effects between multi-source data, and such effects are difficult to remove but can be reduced. If the effect is relatively small, it is acceptable, but if the batch effect is severe, it may be confused with real biological differences.

[0003] To solve this problem, many methods have been proposed, and many batch correction methods are based on the mutual nearest neighbor strategy. The methods based on the mutual nearest neighbor strategy have significant advantages in terms of time efficiency, and the batch correction results of many mutual nearest neighbor-based methods are very good. However, the mutual nearest neighbor strategy uses local matching MNN pairs for global correction, so the batch correction effect of the methods based on the mutual nearest neighbor strategy depends on the number and accuracy of MNN pair matching.

[0004] In 2018, the article published by Haghverdi et al. in the journal Nature biotechnology proposed the MNNcorrect method. This method is a batch correction method based on the mutual nearest neighbor strategy, and its implementation scheme is to search for MNN pairs in the original expression space of single-cell transcriptome data, and then calculate a correction vector using the searched MNN pairs, and add this correction vector to each data of the batch to be corrected to complete the batch correction of single-cell transcriptome data. The efficiency of searching for MNN pairs by this method is very low, the accuracy is not high, and the number is also small.

[0005] In 2019, Haghverdi et al. proposed the FastMNN method based on MNNcorrect. This method first performs PCA dimensionality reduction on single-cell transcriptome data, and then searches for MNN pairs in the PCA space. The efficiency of searching for MNN pairs by this method is relatively high, and the number of searched MNN pairs is also much larger than that searched in the original space. However, the disadvantage is that the accuracy of the searched MNN pairs is not high enough.

[0006] In 2021, Yang et al. proposed the iSMNN method in the journal Briefings in bioinformatics. To improve the accuracy of the searched MNN pairs, this method first clusters each batch of single-cell transcriptome data, then searches for MNN pairs among cells of the same type, and finally uses the searched MNN pairs for batch correction. This method completes the batch correction of single-cell transcriptome data through iterative correction until the evaluation threshold converges locally, so there may be an overcorrection problem.

[0007] In 2021, Wang et al. proposed iMAP in the journal Genome biology. This method searches for MNN pairs in a low-dimensional space, then trains a GAN network using the searched MNN pairs, and finally uses the trained GAN network for batch correction. This method does not have the problem of overcorrecting the data, but its time complexity is relatively high. Summary of the Invention

[0008] The object of the present invention is to overcome the deficiencies of the above-mentioned prior art, and propose a method for batch correction of single-cell transcriptome based on mutual nearest neighbors, so as to more accurately find MNN pairs, improve the effect of batch correction, reduce the time complexity, and improve the batch correction efficiency.

[0009] The technical solution of the present invention is as follows: First, perform feature selection on the genes in the single-cell transcriptome data, then train an autoencoder to perform feature extraction and preliminary batch correction on the single-cell transcriptome data, and finally search for MNN pairs on the basis of the previous step, and use the searched MNN pairs to calculate the correction vector to perform batch correction on the data again. Its implementation steps are as follows:

[0010] (1) Use single-cell transcriptome sequencing technology to measure the gene expression values of each cell in the sample, generating multi-batch single-cell transcriptome data;

[0011] (2) Perform feature selection on the genes in the single-cell transcriptome data, that is, screen out the top 2000 genes with the largest variance from the genes in the data as highly expressed genes;

[0012] (3) Construct an autoencoder composed of an encoder and a decoder;

[0013] (4) Construct the loss function of the autoencoder:

[0014] (4a) Randomly select a cell from each batch of data to form a training sample set x:

[0015] x = (x1, x2,..., x i ,..., x m ), where x iDenote the cell data from the $i$-th batch, where $i$ ranges from 1 to $m$, and $m$ is the number of batches;

[0016] (4b) Each time, input a training sample set into the autoencoder. The encoder in the autoencoder encodes $x$ i into a low-dimensional embedded batch-free information $z$ i_bio and batch noise $z$ i_nio :

[0017] $z$ bio $=$ ($z$ 1_bio , $z$ 2_bio , $\cdots$, $z$ i_bio , $\cdots$, $z$ m_bio ),

[0018] $z$ nio $=$ ($z$ 1_nio , $z$ 2_nio , $\cdots$, $z$ i_nio , $\cdots$, $z$ m_nio );

[0019] Among them, $z$ i_nio is represented by a one-hot-vector, with a dimension of the number of batches $m$; the dimension of $z$ i_nio is 1 and 0, that is, the $i$-th dimension is 1, and the other dimensions are 0;

[0020] (4c) Take $z$ bio and $z$ nio obtained from the encoder as the input of the decoder part. Through the decoder, restore $z$ bio to the data of the original dimension which is called the original pseudo-cell;

[0021] (4d) Calculate the reconstruction loss of the autoencoder according to the original sample set $x$ and the pseudo-cell :

[0022] (4e) Add random noise $z$ bio to $z$ nio_ran obtained from the encoder, and input $z$ bio and $z$ nio_ran into the decoder to get the random-noise pseudo-cell Then input into the encoder to obtain the low-dimensional embedded batch-free information $z$ bio_c after removing the random noise;

[0023] (4f) Calculate the content loss of the autoencoder using $z$ bio and $z$ bio_c : $L$ c $=$ $\|z$ bio $- z$bio_c ||;

[0024] (4g) According to the reconstruction loss L of the autoencoder r and the content loss L of the autoencoder c , construct the autoencoder loss function as

[0025] (5) cross-training the encoder and decoder in the autoencoder using m batches of single-cell transcriptome data until the loss function L converges to obtain a trained autoencoder;

[0026] (6) Input the m batches of single-cell transcriptome data into the trained autoencoder for feature extraction and preliminary batch correction to obtain a low-dimensional embedded batch-free dataset Z = (Z1, Z2, ..., Z i , ..., Z m ), where Z i It represents the dimension embedding of the i-th batch data after the encoder feature extraction without batch information.

[0027] (7) Iteratively correct the low-dimensional embedding dataset Z without batch information:

[0028] (7a) Select the batch of data with the largest number of cells in data set Z as the reference data set Z ref ;

[0029] (7b) In the remaining dataset ZZ ref Select the dataset Z with the largest number of cells max ;

[0030] (7c) In the Z ref and the Z max Search for mutually nearest neighbor MNN pairs (x ref , x max ), use the searched MNN pair to calculate the correction vector V, and Z max All the data in Z are added to vector V to complete Z max Correction;

[0031] (7d) The corrected Z max With Z ref Merge as a new reference dataset and convert Z max Remove from the remaining batch dataset;

[0032] (7e) Repeat (7b)-(7d) until all datasets are merged into the reference dataset to obtain the corrected low-dimensional embedded dataset without batch information Z′=(Z′1, Z′2, ..., Z′ i , ..., Z′ m ), where Z′i It represents the result after correcting the data of the i-th batch.

[0033] Compared with the prior art, the present invention has the following advantages:

[0034] First, since the present invention designs an autoencoder model structure to extract features from data and adds random noise during the training of the autoencoder, it can perform preliminary batch correction during the feature extraction of data, improving the batch correction effect.

[0035] Second, since the present invention performs preliminary batch correction on data in the feature extraction stage, it can efficiently search for MNN pairs, and the number of searched MNN pairs is larger and the accuracy is higher, improving the batch correction effect.

[0036] Third, since the present invention cross-trains the encoder and the decoder and adds random noise during the training process, the running time is not affected by the data scale, and the running time only increases slightly as the data scale grows, enabling batch correction of large-scale data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is the implementation flowchart of the present invention;

[0038] Figure 2 is the autoencoder structure diagram in the present invention;

[0039] Figure 3 is the simulation result diagram of batch correction of the real data set DC by the present invention.

[0040] Figure 4 is the simulation result diagram of batch correction of the real data set Cell line by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The following further elaborates in detail on the embodiments and effects of the present invention with reference to the accompanying drawings.

[0042] This embodiment takes the simulated data set generated by the splatter tool as an example. There are three batches of data in this data set, with 3000 cells in each batch and 20000 genes in each cell.

[0043] Refer to Figure 1 This example is based on the mutual nearest neighbor-based single-cell transcriptome batch correction method, and the implementation steps are as follows:

[0044] Step 1, obtain the single-cell transcriptome data set.

[0045] Generate a single-cell transcriptome simulation dataset using the existing open-source splatter software package. There are three batches of data in this dataset, with 3000 cells in each batch and 20000 genes in each cell.

[0046] Step 2: Perform feature screening on the 20000 genes of each cell in the single-cell transcriptome dataset.

[0047] Calculate the variance of each gene in the dataset using the pp.highly_variable_genes function in the open-source scanpy package, and select the top 2000 highly expressed genes ranked by variance.

[0048] Step 3: Construct an autoencoder.

[0049] Refer to Figure 2 The specific implementation of this step is as follows:

[0050] 3.1) Build an encoder consisting of three fully connected layers. The number of nodes in the first layer is 1024, the number of nodes in the second layer is 512, and the number of nodes in the third layer is 259. Initialize the weight of the edge between the first layer nodes and the second layer nodes to 0.025, and initialize the weight of the edge between the second layer nodes and the third layer nodes to 0.036;

[0051] 3.2) Build a decoder consisting of 3 fully connected layers. The number of nodes in the first layer is 259, the number of nodes in the second layer is 512, and the number of nodes in the third layer is 1024; Initialize the weight of the edge between the first layer nodes and the second layer nodes to 0.036, and initialize the weight of the edge between the second layer nodes and the third layer nodes to 0.025;

[0052] 3.3) Combine the above-built encoder and decoder to form an autoencoder.

[0053] Step 4: Construct the loss function of the autoencoder.

[0054] 4.1) Randomly select one cell from each batch of data to form a training sample set x:

[0055] x = (x1, x2, x3),

[0056] where x1 represents the cell data from the first batch, x2 represents the cell data from the second batch, and x3 represents the cell data from the third batch;

[0057] 4.2) Each time, input a training sample set x into the autoencoder. The encoder in the autoencoder encodes x into a low-dimensional embedded batch-free z bio and batch noise z nio :

[0058] zbio =(z 1_bio , z 2_bio , z 3_bio )

[0059] z nio =(z 1_nio , z 2_nio , z 3_nio )

[0060] where z 1_nio represents the noise of the first batch of data, with its first dimension being 1 and other dimensions being 0; z 2_nio represents the noise of the second batch of data, with its second dimension being 1 and other dimensions being 0; z 3_nio represents the noise of the third batch of data, with its third dimension being 1 and other dimensions being 0;

[0061] 4.3) Take the z bio and z nio obtained from the encoder as the input of the decoder part, and restore z bio to the data of the original dimension Call the original pseudo-cell;

[0062] 4.4) Calculate the reconstruction loss of the autoencoder according to the original sample set x and the pseudo-cell :

[0063] 4.5) Add random noise z bio to the z nio_ran obtained from the encoder, and input z bio and z nio_ran into the decoder to get the random-noise pseudo-cell Then input it into the encoder to obtain the low-dimensional embedding without batch information z bio_c after removing the random noise;

[0064] 4.6) Use z bio and z bio_c to calculate the content loss of the autoencoder: L c =||z bio -z bio_c ||;

[0065] 4.7) According to the reconstruction loss L r of the autoencoder and the content loss L c of the autoencoder, construct the autoencoder loss function as

[0066] Step 5: Use the single-cell transcriptome dataset to cross-train the encoder and decoder in the autoencoder.

[0067] 5.1) Fix the parameters of the decoder, calculate the loss function L for the training data, and use the loss function to update the parameters of the three-layer fully connected network of the encoder until the loss function L converges locally.

[0068] 5.2) Fix the parameters of the encoder, calculate the loss function L for the training data, and use the loss function to update the parameters of the three-layer fully connected network of the decoder until the loss function L converges locally.

[0069] 5.3) Repeat (5.1)-(5.2) until the loss function reaches the convergence threshold set by the user to obtain the trained autoencoder.

[0070] Step 6: Extract features from the single-cell transcriptome dataset to obtain a low-dimensional embedded batch-free information dataset.

[0071] Input the single-cell transcriptome dataset into the trained autoencoder for feature extraction and preliminary batch correction to obtain a low-dimensional embedded batch-free information dataset Z:

[0072] Z = (Z1, Z2, Z3)

[0073] where Z1 represents the low-dimensional embedded batch-free information data after feature extraction of the first batch of data by the encoder, Z2 represents the low-dimensional embedded batch-free information data after feature extraction of the second batch of data by the encoder, and Z3 represents the low-dimensional embedded batch-free information dataset after feature extraction of the third batch of data by the encoder.

[0074] Step 7: Correct the low-dimensional embedded batch-free information dataset Z.

[0075] 7.1) Select Z1 in the low-dimensional embedded batch-free information dataset Z as the reference dataset Z ref ;

[0076] 7.2) Select Z2 as the dataset to be corrected Z in the remaining datasets (Z2, Z3) max ;

[0077] 7.3) Search for mutual nearest neighbor MNN pairs (x ref and x max ) between the said Z ref and the said Z max , calculate the correction vector V using the searched MNN pairs, and add the vector V to all the data in Z max to complete the correction of Z max ;

[0078] 7.4) After the correction of Zmax Combine with Z ref to form a new reference dataset, and remove Z max from the remaining dataset;

[0079] 7.5) Select Z3 from the remaining dataset (Z3) as the dataset Z to be corrected max ;

[0080] 7.6) Search for mutual nearest neighbor MNN pairs (x ref , x max ) between Z ref and Z max . Calculate the correction vector V using the searched MNN pairs, and add the vector V to all the data in Z max to complete the correction of Z max ;

[0081] 7.7) Combine the corrected Z max with Z ref to form a new reference dataset, and remove Z max from the remaining dataset to obtain the corrected low-dimensional embedded batch-free information dataset.

[0082] In this example, there are 3 batches, each batch has 3000 cells, and the single-cell transcriptome simulation data of each cell with 20000 genes has completed batch correction.

[0083] The technical effects of the present invention are described below in combination with simulation experiments.

[0084] I. Simulation conditions:

[0085] The computer hardware CPU for the simulation experiment is Intel Core (TM) i7-7700, and the computer hardware memory is 16G;

[0086] Computer software: Conda integrated development software on the WINDOWS 10 system.

[0087] II. Simulation content:

[0088] Simulation 1: The present invention and 7 existing methods, namely LIGER, iMAP, DESC, Harmony, Scanorama, FastMNN, and Seurat, are respectively used for batch correction in the human peripheral blood dendritic cell dataset DC, and the LISI, KBET, and ARI are used as evaluation indicators to compare their batch correction effects. The results are shown in Table 1:

[0089] Table 1 Evaluation of batch correction of the present invention and 7 existing methods in the DC dataset

[0090]

[0091] All the evaluation indicators in Table 1 are scores after quantification. The higher the score, the better the batch correction result. It can be seen from Table 1 that the correction result of the present invention has the best score in all batch correction evaluation indicators, achieving a good purpose of batch removal.

[0092] Simulation 2: The present invention and the existing 7 methods LIGER, iMAP, DESC, Harmony, Scanorama, FastMNN, and Seurat were respectively subjected to batch correction in the Cell line dataset, and the LISI, KBET, and ARI were used as evaluation indicators to compare their batch correction effects. The results are shown in Table 2:

[0093] Table 2 Evaluation of batch correction of the present invention and the existing 7 methods in the Cell line dataset

[0094]

[0095] It can be seen from Table 2 that the present invention has the highest score in the evaluation indicator LISI and is second only to LISI in KBET. There is not much difference among all methods in the evaluation indicator ARI. In summary, the present invention has a good correction effect on the Cell line dataset, and the result is better than other methods.

[0096] Simulation 3: The UMAP visualization tool was used to display the batch correction effect of the present invention in the dataset DC. The results are as Figure 3 shown, where Figure 3 a is the visualization result of the dataset DC before correction, Figure 3 b is the visualization result of the dataset DC after correction.

[0097] From Figure 3 it can be seen that there is an obvious batch effect in the data before correction. After batch correction by the present invention, the data of different batches can be well aggregated together, the batch effect is removed, and single-cell transcriptome data without batch effect is obtained.

[0098] Simulation 4: The UMAP visualization tool was used to display the batch correction effect of the present invention in the dataset Cell line. The results are as Figure 4 shown, where Figure 4 a is the visualization result of the dataset Cell line before correction, Figure 4 b is the visualization result of the dataset Cell line after correction.

[0099] From Figure 4It can be seen that there is an obvious batch effect in the data before correction. After batch correction by the present invention, the data of different batches can be well aggregated together, the batch effect is removed, and single-cell transcriptome data without batch effect is obtained.

Claims

1. A batch correction method for single-cell transcriptomes based on mutual nearest neighbors, characterized in that, It includes the following steps: (1) Measure the gene expression values of each cell in the sample using single-cell transcriptome sequencing technology to generate multi-batch single-cell transcriptome data; (2) Perform feature selection on the genes in the single-cell transcriptome data, that is, screen out the top 2000 genes with the largest variance from the genes in the data as highly expressed genes; (3) Construct an autoencoder composed of an encoder and a decoder; (4) Construct the loss function of the autoencoder: (4a) Randomly select a cell from each batch of data to form a training sample set x; x = (x1, x2, …, x i , …, x m ), where x i represents the cell data from the i-th batch, and the value range of i is from 1 to m, where m is the number of batches; (4b) Each time a training sample set is input to the autoencoder, the encoder in the autoencoder encodes x i into a low-dimensional embedding without batch information z i_bio and batch noise z i_nio : z bio =(z 1_bio , z 2_bio , …, z i_bio , …, z m_bio ), z nio = (z 1_nio , z 2_nio , …, z i_nio , …, z m_nio ); Among them, z i_nio is represented by a one-hot vector with a dimension of the batch size m; z i_nio has dimensions of 1 and 0, that is, the i-th dimension is 1 and the other dimensions are 0; (4c) Obtain z from the encoder bio and z nio as the input of the decoder part, and through the decoder, restore z bio to the data of the original dimension which is called the original pseudo-cell; (4d) Calculate the reconstruction loss of the autoencoder based on the original sample set x and the pseudo-cells Calculate the reconstruction loss of the autoencoder: (4e) z obtained from the encoder bio Add random noise z nio_ran and input z bio and z nio_ran into the decoder to obtain random noise pseudo-cells Then input it into the encoder to obtain the low-dimensional embedding z without batch information after removing the random noise bio_c ; (4f) Use z bio and z bio_c to calculate the content loss of the autoencoder: L c = ||z bio - z bio_c ||; (4g) According to the reconstruction loss L of the autoencoder r and the content loss L of the autoencoder c , construct the autoencoder loss function as (5) Use m batches of single-cell transcriptome data to cross-train the encoder and decoder in the autoencoder until the loss function L converges to obtain a trained autoencoder; (6) Input the single-cell transcriptome data of m batches into the trained autoencoder for feature extraction and preliminary batch correction to obtain a low-dimensional embedded batch-free information dataset Z = (Z1, Z2, …, Z i , …, Z m ), where Z i represents the batch-free information data of the i-th batch of data after encoder feature extraction; (7) Iteratively correct the low-dimensional embedded batch-free information dataset Z; (7a) Select the batch data with the largest number of cells in the dataset Z as the reference dataset Z ref ; (7b) Select the dataset Z with the largest number of cells from the remaining dataset Z-Z ref from the remaining dataset Z-Z max ; (7c) Search for mutual nearest neighbor (MNN) pairs (x ref , x max ) between the said Z ref and the said Z max . Calculate the correction vector V using the searched MNN pairs, and add all the data in Z max with the vector V to complete the correction of Z max . (7d) The corrected Z max is merged with Z ref to form a new reference data set, and Z max is deleted from the remaining batch data set; (7e) Repeat (7b)-(7d) until all data sets are merged into the reference data set to obtain the corrected low-dimensional embedded batch-free information data set Z ′ =(Z ′ 1, Z ′ 2, …, Z ′ i , …, Z ′ m ), where Z ′ i represents the result after correction of the i-th batch of data.

2. The method according to claim 1, characterized in that, The top 2000 highly variable genes selected from the transcriptome data in (2) are the top 2000 highly expressed genes extracted from the single-cell transcriptome data through the pp.highly_variable_genes function in the scanpy package.

3. The method according to claim 1, characterized in that, In step (4), the encoder in the autoencoder is constructed, which consists of a three-layer fully connected network. The number of nodes in the first layer is 1024, the number of nodes in the second layer is 512, and the number of nodes in the third layer is 256 + m, where m is the number of batches of input data. The weight of the edge between the first layer nodes and the second layer nodes is initialized to 0.025, and the weight of the edge between the second layer nodes and the third layer nodes is initialized to 0.

036.

4. The method according to claim 1, wherein In step (4), the decoder in the autoencoder is constructed, which consists of a three-layer fully connected network. The number of nodes in the first layer is 256 + m, where m is the number of batches of input data, the number of nodes in the second layer is 512, and the number of nodes in the third layer is 1024. The weight of the edge between the first layer nodes and the second layer nodes is initialized to 0.036, and the weight of the edge between the second layer nodes and the third layer nodes is initialized to 0.

025.

5. The method according to claim 1, wherein In step (5), using m batches of single-cell transcriptome data to cross-train the encoder and decoder in the autoencoder is realized as follows: (5a) Fix the parameters of the decoder, calculate the loss function L for the training data, and use the loss function to update the parameters of the encoder network until the loss function L converges locally; (5b) Fix the parameters of the encoder, calculate the loss function L for the training data, and use the loss function to update the parameters of the decoder network until the loss function L converges locally; (5c) Repeat (5a)-(5b) until the loss function reaches the convergence threshold set by the user to obtain a trained autoencoder.

6. The method according to claim 1, characterized in that, In step (7), search for MNN pairs between the reference dataset Z ref and the dataset Z to be corrected max as follows: (7a) For the reference data set Z ref each data a i in the data set Z to be corrected max search for the 20 data with the closest distance to it to form i the nearest neighbor set of a; (7b) Treat the correction dataset Z max for each data b i in the reference dataset Z ref search for the 20 data with the closest distance to it to form i the nearest neighbor set of b; (7c) Traverse the nearest neighbor sets of all data. If a i exists in the nearest neighbor set of b i , and b i exists in the nearest neighbor set of a i , then obtain a set of MNN pairs.

Citation Information

Patent Citations

  • Data batch effect correction method based on connected graph and generative adversarial network

    CN114400047A

  • Neural network for cell image analysis for identification of abnormal cells

    US6463438B1