Method and system for integrated analysis of multi-omics single-cell transcriptome data

By standardizing gene names, eliminating background RNA contamination, and correcting sequencing depth and batch effects, the accuracy and quality issues in the integrated analysis of multi-source single-cell transcriptome data were resolved, enabling more efficient data integration and analysis.

CN121545583BActive Publication Date: 2026-05-08HANGZHOU NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU NORMAL UNIVERSITY
Filing Date
2026-01-19
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

The integration analysis of multi-source single-cell transcriptome data suffers from background RNA contamination, differences in reference genomes and gene annotations, uneven sequencing depth, and batch effects, resulting in low accuracy and poor quality of data integration.

Method used

By unifying gene names and filtering source-specific genes, a specific background gene pool was constructed to eliminate background RNA contamination, correcting sequencing depth differences, and anchor cells were constructed for batch effect correction. The k-means algorithm and metacell method were used for cell type identification and correction.

Benefits of technology

This improved the accuracy of integrated analysis of single-cell transcriptome data, reduced the impact of batch effects caused by multiple sources on subsequent analyses, and ensured data quality and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545583B_ABST
    Figure CN121545583B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data analysis, in particular to a multi-source single-cell transcriptome data integration analysis method and system. The method comprises the following steps: converting all gene names of different source single-cell transcriptome data into a unified version, filtering source-specific genes; constructing a specific background gene pool and identifying contaminant genes to eliminate background RNA contamination; calculating the average sequencing depth of each sample cell and correcting the sequencing depth of each sample cell; based on different source single-cell transcriptome data, cell type identification is carried out through cell clustering and characteristic gene expression analysis, and each cell type is obtained; for each cell type, an anchor cell is constructed, and the expression profile of the depth-balanced cell of the same cell type is corrected to obtain the integrated cell expression profile. The accuracy of single-cell transcriptome data integration analysis is improved, and the influence of batch effect caused by multi-source on subsequent other analysis is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis technology, and in particular to a method and system for the integrated analysis of multi-source single-cell transcriptome data. Background Technology

[0002] The advent of single-cell transcriptome sequencing (scRNA-seq) technology has overcome the limitation of traditional bulk RNA sequencing (bulk RNA-seq) in distinguishing transcriptional heterogeneity within individual cells. It enables precise capture of gene expression characteristics at single-cell resolution, providing robust data support for cell subpopulation classification, rare cell type identification, analysis of dynamic changes in cell state, and the elucidation of intercellular regulatory networks. Single-cell transcriptome sequencing technology has been widely applied in medicine, biology, agriculture, and other fields, covering research on various human tissues and organs, model organisms, and crops. With the rapid popularization and promotion of single-cell transcriptome sequencing technology, research teams both domestically and internationally have accumulated massive amounts of multi-source single-cell transcriptome data. Integrating and analyzing multi-source single-cell transcriptome data has significant scientific value and practical application implications. On the one hand, obtaining many biological samples is difficult and costly. Integrating sample data from different sources can effectively lower the research threshold and reduce the waste of resources caused by repeated experiments. On the other hand, single-cell transcriptome sequencing experiments and data analysis are costly. Integrating data from multiple sources can maximize data utilization, uncover more comprehensive and reliable biological laws behind the data, and provide the possibility for comparative analysis across research, species, and platforms.

[0003] However, the integrated analysis of multi-source single-cell transcriptome data faces a series of challenges, including background RNA contamination, differences in reference genomes and gene annotations, uneven sequencing depth, and significant batch effects. Specifically, the processes of single-cell isolation and library construction easily generate large amounts of free RNA, leading to high levels of background RNA contamination, which distorts gene expression signals and interferes with the identification of true biological characteristics. Differences exist between different versions of the genome sequence, and gene names, gene numbers, and gene function annotations may also change or have aliases, making it impossible to accurately match the same gene in data from different sources. Factors such as sequencing throughput of different sequencing platforms, experimental batch technical parameter settings, and sample RNA quality can lead to significant differences in sequencing depth among data from different sources. Furthermore, different experimental batches, sequencing platforms, and sample processing conditions can introduce batch effects caused by non-biological factors, resulting in technical variations in the data that are irrelevant to the research objectives.

[0004] Therefore, traditional integration analysis of multi-source single-cell transcriptome data often results in low accuracy and poor quality due to background RNA contamination, differences in reference genomes and gene annotations, uneven sequencing depth, and significant batch effects. Summary of the Invention

[0005] Therefore, in order to solve the above-mentioned technical problems, a method and system for integrating and analyzing multi-source single-cell transcriptome data is provided, which can eliminate technical variations in multi-source data and improve the quality and accuracy of data integration.

[0006] An integrated analysis method for multi-source single-cell transcriptome data, the method comprising:

[0007] The gene names of all genes from single-cell transcriptome data from different sources were counted, and each gene name was converted into a uniform version. Genes expressed only in a single source were recorded as source-specific genes and filtered to obtain the initial cellular expression profile of the gene set.

[0008] A specific background gene pool was constructed, and contaminating genes in each sample cell of the initial cell expression profile were identified by expression ratio threshold and average expression level threshold. Background RNA contamination was eliminated based on the identification results to obtain the cell purification expression profile.

[0009] Based on the cell purification expression profile, the average sequencing depth of each sample cell is calculated, and the sequencing depth of each sample cell is corrected to obtain the depth-balanced cell expression profile after correction.

[0010] Based on single-cell transcriptome data from different sources, cell types were identified by cell clustering and characteristic gene expression analysis to obtain various cell types.

[0011] Anchor cells are constructed for each cell type. Based on the anchor cells, the centroid and source correction factor of the depth-balanced cell expression profile are calculated. The depth-balanced cell expression profile of the same cell type is then corrected to obtain the integrated cell expression profile.

[0012] In one embodiment, the individual gene names are converted to a uniform version, and genes expressed only in a single source are designated as source-specific genes and filtered, including:

[0013] Based on the stable IDs and gene names of the NCBI and Ensembl reference genomes, all gene names from single-cell transcriptome data from different sources are converted into a unified version;

[0014] The gene expression levels of all cells within a single source are merged, the expressed genes of each source are counted, and genes expressed only in a single source are recorded as source-specific genes and deleted.

[0015] In one embodiment, a specific background gene pool is constructed, and contamination genes in each sample cell of the initial cell expression profile are identified by using an expression ratio threshold and an average expression level threshold, including:

[0016] A specific background gene pool is constructed, and unsupervised clustering is performed on individual sample cells. The expression ratio and average expression level of each gene in the specific background gene pool in each cell population are calculated.

[0017] Genes whose expression ratio exceeds the expression ratio threshold and whose average expression level in each group exceeds the average expression level threshold are identified as polluting genes.

[0018] In one embodiment, background RNA contamination is eliminated based on the identification results to obtain the cell-purified expression profile, including:

[0019] The minimum expression level of the polluting gene in each cell population was taken as the polluting expression level, and the polluting expression level was removed.

[0020] All sample cells in the initial cell expression profile were subjected to contamination gene identification and removal to obtain the purified cell expression profile.

[0021] In one embodiment, based on the cell purification expression profile, the average sequencing depth of each sample cell is calculated, including:

[0022] Based on the cell purification expression profile, the total number of transcripts in all cells of the sample and the number of cells in the sample were determined.

[0023] The average sequencing depth of each sample cell was calculated based on the total number of transcripts in all cells of the sample and the number of cells in the sample.

[0024] In one embodiment, sequencing depth correction is performed on each sample cell to obtain a depth-balanced cell expression profile after correction, including:

[0025] Compare the average sequencing depth of each sample cell, select the minimum sequencing depth from the various average sequencing depths as the benchmark, and calculate the maximum correction depth;

[0026] If the average sequencing depth of the sample cells is greater than the maximum correction depth, the number of transcripts to be reduced in the sample cells is calculated and the gene expression level is reduced to complete the sequencing depth correction and obtain the depth-balanced cell expression profile after correction.

[0027] In one embodiment, anchor cells are constructed for each cell type, including:

[0028] The k-means algorithm is used to classify cells of the same origin and cell type, and seed cells are randomly selected from the classified cells.

[0029] Identify similar cells corresponding to each seed cell, and merge the expression profiles of similar cells by averaging them to generate anchor cells.

[0030] In one embodiment, the centroid and source correction factor of the depth-balanced cell expression profile are calculated based on the anchor cells, and the depth-balanced cell expression profile of the same cell type is corrected, including:

[0031] Calculate the centroid of the expression profile corresponding to the cell type based on anchor cells of the same cell type from all sources;

[0032] The expression profile difference between each source anchor cell and the centroid is calculated, and the average value of each expression profile difference is calculated. The average value is used as the source correction factor.

[0033] The correction is completed by subtracting the source correction factor from the expression profile of the cell type.

[0034] An integrated analysis system for multi-source single-cell transcriptome data, the system comprising:

[0035] A unified naming and gene filtering module is used to count all gene names in single-cell transcriptome data from different sources, convert each gene name into a unified version, and record genes expressed only in a single source as source-specific genes and filter them to obtain the initial cell expression profile of the gene set.

[0036] The contamination elimination module is used to construct a specific background gene pool, identify contaminating genes in each sample cell of the initial cell expression profile by means of expression ratio threshold and average expression level threshold, and eliminate background RNA contamination based on the identification results to obtain the cell purified expression profile.

[0037] The difference correction module is used to calculate the average sequencing depth of each sample cell based on the cell purification expression profile, and to correct the sequencing depth of each sample cell to obtain a depth-balanced cell expression profile after correction.

[0038] The cell type identification module is used to identify cell types based on single-cell transcriptome data from different sources by analyzing cell clustering and characteristic gene expression levels, thereby obtaining various cell types.

[0039] The cell expression profile correction and integration module is used to construct anchor cells for each cell type, calculate the centroid and source correction factor of the deep balanced cell expression profile based on the anchor cells, and correct the deep balanced cell expression profile of the same cell type to obtain the integrated cell expression profile.

[0040] In one embodiment, the naming and gene filtering module is also used to convert all gene names of single-cell transcriptome data from different sources into a unified version based on the stable IDs and gene names of the NCBI and Ensembl reference genomes; merge the gene expression levels of all cells within a single source; count the expressed genes from each source; and deleting genes that are expressed only in a single source as source-specific genes.

[0041] The aforementioned integrated analysis method and system for multi-source single-cell transcriptome data reduces the influence of the reference genome by standardizing gene names and removing source-specific genes; it reduces the impact of cell-free RNA contamination by identifying and removing background RNA; and it eliminates batch effects between samples by correcting for sequencing depth differences, constructing anchor cells, and using anchor correction methods. This improves the accuracy of integrated analysis of single-cell transcriptome data and reduces the impact of batch effects caused by multiple sources on subsequent analyses. Attached Figure Description

[0042] Figure 1 This is a diagram illustrating the application environment of an integrated analysis method for multi-source single-cell transcriptome data in one embodiment.

[0043] Figure 2 This is a flowchart illustrating a method for integrating and analyzing multi-source single-cell transcriptome data in one embodiment;

[0044] Figure 3 This is a block diagram of an integrated analysis system for multi-source single-cell transcriptome data in one embodiment;

[0045] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0047] The method for integrating and analyzing multi-source single-cell transcriptome data provided in this application can be applied to, for example... Figure 1 The application environment shown. For example... Figure 1As shown, the application environment includes computer device 110. Computer device 110 can statistically analyze all gene names from single-cell transcriptome data from different sources, convert each gene name to a unified version, and identify genes expressed only in a single source as source-specific genes and filter them to obtain an initial cell expression profile of the gene set. Computer device 110 can construct a specific background gene pool, identify contaminating genes in each sample cell of the initial cell expression profile using expression ratio thresholds and average expression level thresholds, and eliminate background RNA contamination based on the identification results to obtain a purified cell expression profile. Based on the purified cell expression profile, computer device 110 can calculate the average sequencing depth of each sample cell and perform sequencing depth correction on each sample cell to obtain a depth-balanced cell expression profile after correction. Computer device 110 can identify cell types based on single-cell transcriptome data from different sources through cell clustering and characteristic gene expression level analysis to obtain various cell types. Computer device 110 can construct anchor cells for each cell type, calculate the centroid and source correction factor of the depth-balanced cell expression profile based on the anchor cells, and correct the depth-balanced cell expression profile of the same cell type to obtain an integrated cell expression profile. Among them, computer equipment 110 may include, but is not limited to, various personal computers, laptops, smartphones, robots, tablets and other devices.

[0048] In one embodiment, such as Figure 2 As shown, an integrated analysis method for multi-source single-cell transcriptome data is provided, including the following steps:

[0049] Step 202: Count all gene names in single-cell transcriptome data from different sources, convert each gene name to a unified version, and record genes expressed only in a single source as source-specific genes and filter them to obtain the initial cell expression profile of the gene set.

[0050] Computer equipment can standardize gene names and filter source-specific genes, providing a unified identification benchmark for subsequent background RNA contamination removal. Filtering source-specific genes can also eliminate noise genes that are expressed only in a single source, preventing such genes from being misidentified as contaminant genes or interfering with the calculation of the expression ratio and average expression level of real contaminant genes, thus ensuring the accuracy of contamination removal.

[0051] In one embodiment, the integrated analysis method for multi-source single-cell transcriptome data may further include a gene name unification and gene filtering process. The specific process includes: converting all gene names of single-cell transcriptome data from different sources into a unified version based on the stable IDs and gene names of the NCBI and Ensembl reference genomes; merging the gene expression levels of all cells within a single source, counting the expressed genes from each source, and deleting genes that are expressed only in a single source as source-specific genes.

[0052] The computer device can statistically analyze all gene names from single-cell transcriptomes from different sources. Based on stable IDs and gene names from NCBI and Ensembl reference genomes, it converts the original gene names into a unified version to avoid potential gene name variations or aliases between different versions. Next, the device can merge gene expression levels from all cells within a single source, marking genes with non-zero expression levels as expressed genes from that source. Then, it statistically analyzes the expressed genes from each source individually, marking genes expressed only in a single source as source-specific genes, and finally removing all source-specific genes from the expression profile.

[0053] Step 204: Construct a specific background gene pool, identify contaminating genes in each sample cell of the initial cell expression profile by means of expression ratio threshold and average expression level threshold, and eliminate background RNA contamination based on the identification results to obtain the cell purified expression profile.

[0054] Computer equipment can identify, calculate, and remove background genes and contaminant expression levels in each sample, based on a specific background gene pool.

[0055] In one embodiment, the integrated analysis method for multi-source single-cell transcriptome data may further include a process for identifying contaminating genes. The specific process includes: constructing a specific background gene pool, performing unsupervised clustering of individual sample cells, calculating the expression ratio and average expression level of each gene in the specific background gene pool in each cell population, and identifying genes whose expression ratio exceeds the expression ratio threshold and whose average expression level in each population exceeds the average expression threshold as contaminating genes.

[0056] Computer equipment can construct a specific background gene pool corresponding to a target tissue or cell, such as a common background gene pool for the liver, and then identify contaminating genes in a single sample based on unsupervised clustering and calculate the contamination expression level.

[0057] Specifically, the computer equipment can use the standard analysis workflow of Seurat software to perform data standardization, PCA dimensionality reduction, and clustering on individual cell samples, calculating the expression proportion of each gene in the background gene pool across all cells in the sample and its average expression level in each cell population. Specifically, when calculating the expression proportion of each gene across all cells in the sample, it can be based on... Calculations show that This represents the percentage of cells expressing gene A across all cells in a single sample. This represents the number of cells expressing gene A in the sample, and N represents all cells in the sample; the average expression level of each gene in each cell population is calculated and labeled as follows. ,like This represents the average expression level of gene A in group 1 cells.

[0058] The computer equipment can be pre-set with an expression ratio threshold. and average expression threshold This is used to identify contaminating genes. Specifically, if the gene expression ratio in a specific background gene pool in a single sample is greater than an expression ratio threshold... Meanwhile, the average expression level in each cell population was greater than the average expression threshold. If a gene is identified as a contaminant gene, it is labeled as such, and the minimum expression level of this gene in each cell population is recorded as the contaminant expression level. In this embodiment, the specific background gene pool can be increased or decreased according to the type of target tissue or cell and research needs.

[0059] In one embodiment, the integrated analysis method for multi-source single-cell transcriptome data may further include a process for eliminating background RNA contamination. The specific process includes: taking the minimum expression level of contaminating genes in each cell population as the contamination expression level and removing the contaminating expression level; identifying and removing contaminating genes in all sample cells of the initial cell expression profile to obtain a purified cell expression profile.

[0060] Computer equipment can subtract the contamination expression level P of all contaminating genes from the expression profile of a single sample. Exp This generates a cell expression profile with background removed. Then, the computer can repeatedly identify contaminating genes in the cell expression profiles of all samples and subtract the contaminating gene expression levels to obtain the background-removed expression profile for each sample, i.e., the expression profile after cell purification.

[0061] Step 206: Based on the expression profile after cell purification, calculate the average sequencing depth of each sample cell, and perform sequencing depth correction on each sample cell to obtain a depth-balanced cell expression profile after correction.

[0062] Computer equipment can calculate the average sequencing depth of each sample cell based on the expression profile after cell purification, and flatten the transcripts to ensure that the average sequencing depth of the sample with the maximum depth is no more than 130% of that of the sample with the minimum depth.

[0063] In one embodiment, the integrated analysis method for multi-source single-cell transcriptome data may further include a process for calculating the average sequencing depth. The specific process includes: determining the total number of transcripts in all cells of the sample and the number of cells in the sample based on the expression profile after cell purification; and calculating the average sequencing depth of each sample cell based on the total number of transcripts in all cells of the sample and the number of cells in the sample.

[0064] Computer devices can, based on the background-removed cell expression profile, according to a formula Calculate the average sequencing depth for each sample, where The total number of transcripts in all cells of the sample is represented by N, which is the number of cells in the sample. This allows us to calculate the average sequencing depth (Depth) for a single sample.

[0065] In one embodiment, the integrated analysis method for multi-source single-cell transcriptome data may further include a sequencing depth correction process, which specifically includes: comparing the average sequencing depth of each sample cell, selecting the minimum sequencing depth from each average sequencing depth as a benchmark, and calculating the maximum correction depth; if the average sequencing depth of the sample cells is greater than the maximum correction depth, then calculating the number of transcripts to be reduced in the sample cells and reducing gene expression levels to complete the sequencing depth correction and obtain a depth-balanced cell expression profile after correction.

[0066] Computer devices can compare the Depth values ​​of each sample and mark the minimum value as... Then based on the formula Calculate the maximum correction depth If the average sequencing depth of a single sample is greater than Then, sequencing depth correction is performed on the sample.

[0067] Specifically, in this embodiment, the computer device can be configured according to the formula. The number of transcripts to be reduced, i.e., the corrected transcript number, is calculated; where Depth represents the average sequencing depth of the sample, and N represents the number of cells in the sample. Then, a computer can randomly select transcripts from a cell-transcript pool. The number of transcripts is reduced, and the expression level of the corresponding genes is decreased in the cell expression profile to obtain the cell expression profile after depth correction.

[0068] The cell-transcript pool consists of transcripts of different genes from each cell, arranged according to their actual expression levels. Transcripts are reduced through random sampling to achieve sequencing depth flattening. For example, if cell 1 has 2 transcripts of gene A and 3 transcripts of gene B, and cell 2 has 3 transcripts of gene A and 1 transcript of gene B, then the cell-transcript pool would contain {A1, A1, B1, B1, B1, A2, A2, A2, B2}.

[0069] Computer equipment can sequentially analyze all samples with an average sequencing depth greater than [missing information]. The samples were subjected to sequencing depth correction to obtain a depth-balanced cell expression profile after correction.

[0070] Step 208: Based on single-cell transcriptome data from different sources, cell types are identified through cell clustering and characteristic gene expression analysis to obtain various cell types.

[0071] Computer equipment can merge all samples from a single source and annotate cell types in samples from each source through cell clustering and characteristic gene expression analysis.

[0072] Specifically, computer equipment can merge all samples from a single source and perform cell type identification using conventional methods. These conventional methods typically involve grouping cells and identifying cell types based on the expression levels of characteristic genes for each cell type. The computer equipment can then annotate all cells with their cell type; for example, in a liver sample, annotations would typically include hepatocytes, bile duct cells, fibroblasts, B cells, T cells, and myeloid cells.

[0073] Step 210: Construct anchor cells for each cell type, calculate the centroid and source correction factor of the depth-balanced cell expression profile based on the anchor cells, and correct the depth-balanced cell expression profile of the same cell type to obtain the integrated cell expression profile.

[0074] Within each cell type, the computer device can use the metacell method to obtain cell type anchors and correct batch effects in the data based on these anchors.

[0075] In one embodiment, the integrated analysis method for multi-source single-cell transcriptome data may further include the process of constructing anchor cells. The specific process may include: classifying cells of the same source and cell type using the k-means algorithm, and randomly selecting seed cells from the classified cells; identifying similar cells corresponding to each seed cell, and merging the expression profiles of similar cells by averaging them to generate anchor cells.

[0076] Computer equipment can use the metacell method to obtain several anchor cells that can represent the expression profile characteristics of that cell type from all cells of the same source and cell type. In this embodiment, the number of anchor cells is 50.

[0077] The implementation process of metacell can be summarized as follows: using the k-means algorithm to divide cells of the same origin and cell type into 10 categories, extracting at least one seed cell from each category to obtain a total of 50 seed cells; for a single seed cell, merging the expression profiles of each seed cell with the average of its 25 most similar cells to generate anchor cells; performing the same processing on each seed cell to obtain the expression profiles of 50 anchor cells.

[0078] In one embodiment, the integrated analysis method for multi-source single-cell transcriptome data may further include a correction process, specifically including: calculating the centroid of the expression profile corresponding to the cell type based on anchor cells of the same cell type from all sources; calculating the expression profile difference between each source anchor cell and the centroid, and calculating the average value of each expression profile difference, using the average value as the source correction factor; subtracting the source correction factor from the expression profile of the cell type to complete the correction.

[0079] The computer device can calculate the centroid of the expression profile of the cell type based on anchor cells of the same cell type from all sources, calculate the difference in expression profile between each source anchor cell and the centroid, and take the average value as the correction factor for that source. The correction factor is then subtracted from the expression profile of the corresponding cell type from that source. This process is repeated for each cell type to obtain integrated single-cell transcriptome data.

[0080] In this embodiment, taking hepatocytes (Hep) as an example, the centroid of hepatocytes is calculated based on all anchor cells from different sources, and then the expression spectrum difference between each anchor cell and the centroid is calculated. The average of the differences from the same source is taken as the hepatocyte correction factor for that source. Subtracting the hepatocyte correction factor from the hepatocyte expression spectrum of that source yields the corrected hepatocyte expression spectrum. This correction process is repeated for hepatocytes from other sources to obtain the complete corrected multi-source data cell expression spectrum.

[0081] This application provides a method for integrating and analyzing multi-source single-cell transcriptome data. By standardizing gene names and removing source-specific genes, the influence of the reference genome can be reduced; by identifying and removing background RNA, the impact of cell-free RNA contamination can be reduced; and by correcting for sequencing depth differences, constructing anchor cells, and using anchor correction methods, batch effects between samples are eliminated. This improves the accuracy of single-cell transcriptome data integration and analysis and reduces the impact of batch effects caused by multiple sources on subsequent analyses.

[0082] It should be understood that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart above may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0083] In one embodiment, such as Figure 3 As shown, an integrated analysis system for multi-source single-cell transcriptome data is provided, including: a naming and gene filtering module 310, a contamination elimination module 320, a difference correction module 330, a type identification module 340, and a cell expression profile correction and integration module 350, wherein:

[0084] The naming and gene filtering module 310 is used to count all gene names in single-cell transcriptome data from different sources, convert each gene name into a unified version, and record genes expressed only in a single source as source-specific genes and filter them to obtain the initial cell expression profile of the gene set.

[0085] The contamination elimination module 320 is used to construct a specific background gene pool. It identifies contaminating genes in each sample cell in the initial cell expression profile by using expression ratio thresholds and average expression level thresholds, and eliminates background RNA contamination based on the identification results to obtain the cell purification expression profile.

[0086] The difference correction module 330 is used to calculate the average sequencing depth of each sample cell based on the expression profile after cell purification, and to correct the sequencing depth of each sample cell to obtain a depth-balanced cell expression profile after correction.

[0087] The cell type identification module 340 is used to identify cell types based on single-cell transcriptome data from different sources by analyzing cell clustering and characteristic gene expression levels, thereby obtaining various cell types.

[0088] The cell expression profile correction and integration module 350 is used to construct anchor cells for each cell type, calculate the centroid and source correction factor of the depth-balanced cell expression profile based on the anchor cells, and correct the depth-balanced cell expression profile of the same cell type to obtain the integrated cell expression profile.

[0089] In one embodiment, the naming and gene filtering module 310 is also used to convert all gene names of single-cell transcriptome data from different sources into a unified version based on the stable ID and gene name of the NCBI and Ensembl reference genome; merge the gene expression levels of all cells within a single source; count the expressed genes from each source; and deleting genes that are expressed only in a single source as source-specific genes.

[0090] In one embodiment, the contamination elimination module 320 is further configured to construct a specific background gene pool, perform unsupervised grouping of individual sample cells, calculate the expression ratio and average expression level of each gene in the specific background gene pool in each cell group, and identify genes whose expression ratio exceeds the expression ratio threshold and whose average expression level in each group exceeds the average expression level threshold as contamination genes.

[0091] In one embodiment, the contamination removal module 320 is further configured to use the minimum expression level of contaminating genes in each cell population as the contamination expression level and remove the contamination expression level; and to determine and remove contaminating genes from all sample cells in the initial cell expression profile to obtain the cell purified expression profile.

[0092] In one embodiment, the difference correction module 330 is further configured to determine the total number of transcripts in all cells of the sample and the number of cells in the sample based on the expression profile after cell purification; and to calculate the average sequencing depth of each sample cell based on the total number of transcripts in all cells of the sample and the number of cells in the sample.

[0093] In one embodiment, the differential correction module 330 is further used to compare the average sequencing depth of each sample cell, select the minimum sequencing depth from each average sequencing depth as a benchmark, and calculate the maximum correction depth; if the average sequencing depth of the sample cells is greater than the maximum correction depth, the number of transcripts to be reduced in the sample cells is calculated and the gene expression level is reduced to complete the sequencing depth correction and obtain the depth-balanced cell expression profile after correction.

[0094] In one embodiment, the cell expression profile correction and integration module 350 is further configured to classify cells of the same origin and cell type using the k-means algorithm, and randomly extract seed cells from the classified cells; determine similar cells corresponding to each seed cell, and merge the expression profiles of similar cells by averaging to generate anchor cells.

[0095] In one embodiment, the cell expression profile correction integration module 350 is further configured to calculate the centroid of the expression profile corresponding to the cell type based on anchor cells of the same cell type from all sources; calculate the expression profile difference between each source anchor cell and the centroid, and calculate the average value of each expression profile difference, using the average value as the source correction factor; subtract the source correction factor from the expression profile of the cell type to complete the correction.

[0096] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for integrated analysis of multi-source single-cell transcriptome data. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device casing, or an external keyboard, touchpad, or mouse.

[0097] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0098] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement steps of an integrated analysis method for multi-source single-cell transcriptome data.

[0099] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program being executed by a processor to implement steps of an integrated analysis method for multi-source single-cell transcriptome data.

[0100] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0101] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0102] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for integrated analysis of multi-source single-cell transcriptome data, characterized in that, The method includes: The gene names of all genes from single-cell transcriptome data from different sources were counted, and each gene name was converted into a uniform version. Genes expressed only in a single source were recorded as source-specific genes and filtered to obtain the initial cellular expression profile of the gene set. A specific background gene pool was constructed, and contaminating genes in each sample cell of the initial cell expression profile were identified by expression ratio threshold and average expression level threshold. Background RNA contamination was eliminated based on the identification results to obtain the cell purification expression profile. Based on the cell purification expression profile, the average sequencing depth of each sample cell is calculated, and the sequencing depth of each sample cell is corrected to obtain the depth-balanced cell expression profile after correction. Based on single-cell transcriptome data from different sources, cell types were identified by cell clustering and characteristic gene expression analysis to obtain various cell types. Anchor cells are constructed for each cell type, including: classifying cells of the same source and cell type using the k-means algorithm, and randomly selecting seed cells from the classified cells; identifying similar cells corresponding to each seed cell, and merging the expression profiles of similar cells by averaging to generate anchor cells; calculating the centroid and source correction factor of the deep balanced cell expression profile based on the anchor cells, and correcting the deep balanced cell expression profile of the same cell type to obtain the integrated cell expression profile.

2. The method for integrating and analyzing multi-source single-cell transcriptome data according to claim 1, characterized in that, Convert all gene names to a uniform version, and classify genes expressed only in a single source as source-specific genes and filter them, including: Based on the stable IDs and gene names of the NCBI and Ensembl reference genomes, all gene names from single-cell transcriptome data from different sources are converted into a unified version; The gene expression levels of all cells within a single source are merged, the expressed genes of each source are counted, and genes expressed only in a single source are recorded as source-specific genes and deleted.

3. The method for integrating and analyzing multi-source single-cell transcriptome data according to claim 1, characterized in that, A specific background gene pool was constructed, and contamination genes in each sample cell of the initial cell expression profile were identified by using expression ratio thresholds and average expression level thresholds, including: A specific background gene pool is constructed, and unsupervised clustering is performed on individual sample cells. The expression ratio and average expression level of each gene in the specific background gene pool in each cell population are calculated. Genes whose expression ratio exceeds the expression ratio threshold and whose average expression level in each group exceeds the average expression level threshold are identified as polluting genes.

4. The method for integrating and analyzing multi-source single-cell transcriptome data according to claim 3, characterized in that, Background RNA contamination was eliminated based on the identification results, and the expression profile after cell purification was obtained, including: The minimum expression level of the polluting gene in each cell population was taken as the polluting expression level, and the polluting expression level was removed. All sample cells in the initial cell expression profile were subjected to contamination gene identification and removal to obtain the purified cell expression profile.

5. The method for integrating and analyzing multi-source single-cell transcriptome data according to claim 1, characterized in that, Based on the cell purification expression profile, the average sequencing depth of each sample cell was calculated, including: Based on the cell purification expression profile, the total number of transcripts in all cells of the sample and the number of cells in the sample were determined. The average sequencing depth of each sample cell was calculated based on the total number of transcripts in all cells of the sample and the number of cells in the sample.

6. The method for integrating and analyzing multi-source single-cell transcriptome data according to claim 5, characterized in that, Sequencing depth correction was performed on each cell sample to obtain a depth-balanced cell expression profile after correction, including: Compare the average sequencing depth of each sample cell, select the minimum sequencing depth from the various average sequencing depths as the benchmark, and calculate the maximum correction depth; If the average sequencing depth of the sample cells is greater than the maximum correction depth, the number of transcripts to be reduced in the sample cells is calculated and the gene expression level is reduced to complete the sequencing depth correction and obtain the depth-balanced cell expression profile after correction.

7. The method for integrating and analyzing multi-source single-cell transcriptome data according to claim 1, characterized in that, Based on the anchor cells, the centroid and source correction factor of the depth-balanced cell expression profile are calculated, and the depth-balanced cell expression profile of the same cell type is corrected, including: Calculate the centroid of the expression profile corresponding to the cell type based on anchor cells of the same cell type from all sources; The expression profile difference between each source anchor cell and the centroid is calculated, and the average value of each expression profile difference is calculated. The average value is used as the source correction factor. The correction is completed by subtracting the source correction factor from the expression profile of the cell type.

8. An integrated analysis system for multi-source single-cell transcriptome data, characterized in that, The system includes: A unified naming and gene filtering module is used to count all gene names in single-cell transcriptome data from different sources, convert each gene name into a unified version, and record genes expressed only in a single source as source-specific genes and filter them to obtain the initial cell expression profile of the gene set. The contamination elimination module is used to construct a specific background gene pool, identify contaminating genes in each sample cell of the initial cell expression profile by means of expression ratio threshold and average expression level threshold, and eliminate background RNA contamination based on the identification results to obtain the cell purified expression profile. The difference correction module is used to calculate the average sequencing depth of each sample cell based on the cell purification expression profile, and to correct the sequencing depth of each sample cell to obtain a depth-balanced cell expression profile after correction. The cell type identification module is used to identify cell types based on single-cell transcriptome data from different sources by analyzing cell clustering and characteristic gene expression levels, thereby obtaining various cell types. The cell expression profile correction and integration module is used to construct anchor cells for each cell type. This includes: classifying cells of the same source and cell type using the k-means algorithm, and randomly selecting seed cells from the classified cells; identifying similar cells corresponding to each seed cell, and merging the expression profiles of similar cells by averaging to generate anchor cells; calculating the centroid and source correction factor of the deep-balanced cell expression profile based on the anchor cells, and correcting the deep-balanced cell expression profile of the same cell type to obtain the integrated cell expression profile.

9. The integrated analysis system for multi-source single-cell transcriptome data according to claim 8, characterized in that, The unified naming and gene filtering module is also used to convert all gene names of single-cell transcriptome data from different sources into a unified version based on the stable IDs and gene names of the NCBI and Ensembl reference genomes; merge the gene expression levels of all cells within a single source, count the expressed genes from each source, and deleting genes that are expressed only in a single source as source-specific genes.

Citation Information

Patent Citations

  • Method for annotating cell identities based on single cell transcriptome clustering results

    CN110060729A

  • Analysis method suitable for 10x single cell transcriptome sequencing data

    CN112599199A