A sequencing analysis method of single-cell transcriptome based on ont sequencing

By combining ONT sequencing technology with various software tools, the limitations of single-cell RNA sequencing technology in capturing and analyzing full-length RNA molecules have been overcome, enabling comprehensive analysis of transcripts, improving data quality, reducing costs, and expanding the understanding of single-cell gene expression.

CN119091964BActive Publication Date: 2025-11-21WUHAN BEINA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411262686.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-11-21
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

Existing single-cell RNA sequencing technologies have limitations in capturing and analyzing full-length RNA molecules, which weakens the comprehensive understanding of post-transcriptional modifications, alternative splicing events, and gene expression regulation mechanisms. Furthermore, ONT single-cell full-length sequencing technology faces challenges in sample preparation, data processing, and cost-effectiveness.

Method used

Using 10×Genomics single-cell capture and labeling technology and ONT sequencing technology, combined with software such as ONT-SC-pipeline, SeuratV5, Harmony, CellPhoneDB V5, Monocle3, isoswitch, and JAFFA, we performed data filtering, alignment, quantification, and advanced analysis. We identified dual-cell structures, removed inactive cells, and performed cell annotation, communication analysis, and alternative splicing analysis to achieve comprehensive analysis of transcripts.

Benefits of technology

This enables more accurate and comprehensive analysis of transcripts, improves sequencing data quality, reduces operating costs, broadens our understanding of the complexity of single-cell gene expression, and provides more data and applications for biological research and clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091964B_ABST
    Figure CN119091964B_ABST
Patent Text Reader

Abstract

The application discloses a sequencing analysis method for single-cell transcriptome based on ONT sequencing and relates to the technical field of gene sequencing analysis.The method comprises the following steps: preparing and constructing a cDNA library of a sample to be measured, sequencing, obtaining three generations of original data, performing quality control on the three generations of original data, obtaining three generations of effective data, performing functional annotation on a reference genome, using software ONT-SC-pipeline to perform data splitting, alignment and quantification, obtaining quantitative data, sequentially performing basic analysis, Marker gene enrichment analysis and advanced analysis on the quantitative data, wherein the basic analysis comprises top gene distribution analysis, hyper-variable gene analysis, clustering and grouping analysis, correlation analysis of each cell group, dimension reduction analysis and Marker gene identification; and the advanced analysis comprises cell annotation analysis, cell communication analysis, cell trajectory analysis, variable splicing analysis and fusion gene analysis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of gene sequencing analysis, in particular to a sequencing analysis method of single-cell transcriptome. BACKGROUND

[0002] In biomedical research and clinical practice, single-cell analysis technology is playing an increasingly crucial role in revealing cell heterogeneity, analyzing complex biological processes and disease pathogenesis. Traditional population-level analysis methods mask the subtle differences between cells, which are essential for understanding disease development, cell decision-making processes, and responses to environmental changes. Single-cell RNA sequencing (scRNA-seq) technology has greatly facilitated the detection and quantification of thousands of RNA molecules from a single cell, providing a unique perspective for in-depth exploration of the internal mechanisms of biological systems.

[0003] Despite the significant progress made in scRNA-seq technology, limitations in capturing and analyzing full-length RNA molecules have weakened the complete information obtained from single cells, limiting the comprehensive understanding of post-transcriptional modifications, alternative splicing events, and gene expression regulation mechanisms.

[0004] In this context, the Oxford Nanopore Technologies (ONT) single-cell full-length sequencing technology emerged, which provides an unprecedented complete transcript perspective by directly sequencing full-length RNA molecules in single cells. This technology not only reveals 5' and 3' ends, coding regions, introns, and alternative splicing events, but also greatly broadens the understanding of the complexity of single-cell gene expression, providing new paths for biological research and clinical diagnosis.

[0005] However, the ONT single-cell full-length sequencing technology faces challenges in sample preparation, data processing, and cost-effectiveness. These challenges involve maintaining the integrity of RNA molecules, improving the quality of sequencing data, and reducing operational costs. Therefore, developing new optimization methods to address these challenges is of great significance to expand the application of ONT single-cell full-length sequencing technology in scientific research and diagnosis. SUMMARY

[0006] The present application provides a sequencing analysis method of single-cell transcriptome based on ONT sequencing. This method utilizes 10x Genomics single-cell capture and labeling technology and ONT sequencing technology to more accurately and comprehensively analyze transcripts and genes in cells, and to deeply mine transcription information in life processes, thereby providing more valuable data and applications in the fields of life science research, medical diagnosis, and drug development. The sequencing analysis method of single-cell transcriptome provided by the present application is specifically implemented through the following technologies.

[0007] A sequencing analysis method of single-cell transcriptome based on ONT sequencing, comprising the following steps:

[0008] A cDNA library of a to-be-tested sample is prepared and constructed, sequencing is performed, and three generations of raw data of the to-be-tested sample are obtained; quality control is performed on the three generations of raw data, and three generations of effective data of the to-be-tested sample are obtained.

[0009] The genes and transcripts of a reference genome are functionally annotated.

[0010] In combination with the three generations of effective data and the gene data of the reference genome, data filtering, data splitting, alignment and quantification are performed using software ONT-SC-pipeline, and quantitative data are obtained.

[0011] Double cells are recognized and removed from the quantitative data, and the data are further filtered.

[0012] The filtered data are subjected to, in sequence, basic analysis, Marker gene enrichment analysis and advanced analysis; the basic analysis includes top gene distribution analysis, high-variable gene analysis, clustering analysis, correlation analysis of each cell cluster, dimension reduction analysis and Marker gene identification; and the advanced analysis includes cell annotation analysis, cell communication analysis, cell trajectory analysis, alternative splicing analysis and fusion gene analysis.

[0013] Further, the quality control processing method is to screen according to sequencing quality values, and to remove sequences with a length of ≤30 bp and a quality value of ≤5.

[0014] Further, the method for performing data splitting, alignment and quantification using software ONT-SC-pipeline is to perform data splitting using vsearch, to perform data alignment using minimap2, and to extract information of the splitting result and the alignment result to obtain the quantitative data.

[0015] Further, the scDblFinder software is used to recognize and remove double cells and low-activity cells, and to remove cells with a high proportion of ribosomal RNA and a high proportion of mitochondrial genes.

[0016] Further, the SeuratV5 software is used to perform top gene distribution analysis; the function FindvariableFeatures is used to perform high-variable gene analysis; and the SeuratV5 software is used to perform clustering analysis, dimension reduction analysis and Marker gene identification.

[0017] Further, the dimension reduction analysis is performed using at least one of the PCA, T-SNE and UMAP analysis methods in the SeuratV5 software.

[0018] Specifically, if the sample to be tested is a multi-sample, the harmony software is used instead of the original PCA analysis method to correct and integrate the filtered multi-sample single cell data.

[0019] Further, the hypergeometric distribution algorithm is used for Marker gene enrichment analysis.

[0020] Further, the known Marker gene and scina software are combined for cell annotation analysis.

[0021] Further, CellPhoneDB V5 is used for cell communication analysis.

[0022] Further, monocle3 software is used for cell trajectory analysis.

[0023] Further, isoswitch software is used for variable splicing analysis.

[0024] Further, JAFFA software is used for fusion gene analysis.

[0025] Further, monocle3 software is used for cell trajectory analysis.

[0026] Further, isoswitch software is used for variable splicing analysis.

[0027] Further, JAFFA software is used for fusion gene analysis.

[0028] Further, based on the results of CellPhoneDB V5, the ggplot2 and igraph methods are used to visualize the communication results of single pathways between various cell subgroups; the contribution of each subunit to the overall ligand and receptor and the subunit of each ligand and receptor pair are visualized.

[0029] Compared with the prior art, the present application has the advantages that:

[0030] 1. Compared with the second-generation sequencing method, the ONT sequencing does not need to break the fragments, can measure the complete transcript, and can quantify the transcript.

[0031] 2. Compared with the existing Pacbio third-generation sequencing, it has a cost advantage, and has a very high cost performance while considering sequencing various results. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 The flowchart of the single cell transcriptome sequencing analysis method based on ONT sequencing provided by the present application.

[0033] Figure 2Flow chart for data filtering, data splitting, alignment and quantification.

[0034] Figure 3 Correlation plot of UMI number before filtering in the specific embodiment: X axis is UMI number, Y axis is gene number, and the number is Pearson correlation coefficient. The left vertical coordinate is the proportion of mitochondria, and the right vertical coordinate is the gene number.

[0035] Figure 4 Statistical chart before and after filtering in the specific embodiment.

[0036] Figure 5 Box plot of TOP genes.

[0037] Figure 6 Highly variable gene dot plot.

[0038] Figure 7 UMAP analysis result plot.

[0039] Figure 8 Plot of the proportion of each Cluster.

[0040] Figure 9 Correlation plot of each category.

[0041] Figure 10 Marker gene heat map.

[0042] Figure 11 Marker gene expression UMAP plot.

[0043] Figure 12 Marker gene GO enrichment bubble plot.

[0044] Figure 13 Cell annotation statistical chart.

[0045] Figure 14 Cell annotation UMAP plot.

[0046] Figure 15 Statistical chart of cell communication number.

[0047] Figure 16 Trajectory plot after selecting root.

[0048] Figure 17 Top gene trend plot.

[0049] Figure 18 Variable splicing example plot.

[0050] Figure 19 Fusion gene example plot. DETAILED DESCRIPTION

[0051] The technical solutions of the present application will be described clearly and completely below. Obviously, the described embodiments are only some of the embodiments of the present application, but not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0052] The present application provides a sequencing analysis method of single-cell transcriptome based on ONT sequencing, and the overall sequencing analysis process is as shown in the figure. Figure 1 The overall steps of the method are as follows:

[0053] Prepare and construct the cDNA library of the sample to be tested, sequence, and obtain the three-generation raw data of the sample to be tested; quality control the three-generation raw data to obtain the three-generation effective data of the sample to be tested;

[0054] Functionally annotate the genes and transcripts of the reference genome;

[0055] Combine the three-generation effective data and the gene data of the reference genome, use the software ONT-SC-pipeline to perform data filtering, data splitting, alignment and quantification, and obtain the quantification data;

[0056] Double-cell recognition and removal are performed on the quantification data, and the data is further filtered;

[0057] The filtered data is sequentially subjected to basic analysis, Marker gene enrichment analysis and advanced analysis; the basic analysis includes top gene distribution analysis, high-variable gene analysis, clustering analysis, correlation analysis of each cell cluster, dimensionality reduction analysis and Marker gene identification; the advanced analysis includes cell annotation analysis, cell communication analysis, cell trajectory analysis, alternative splicing analysis and fusion gene analysis.

[0058] Specifically, the steps of the method are as follows.

[0059] I. Test preparation

[0060] Experimental materials: mouse liver tissue.

[0061] Related reagents, kits and consumables: all purchased from current market products.

[0062] II. Preparation of sample single-cell suspension

[0063] 1. Dissect the mouse, cut part of the liver tissue, wash the blood in the tissue with PBS or physiological saline, and put the washed tissue into a 1 mL preservation tube filled with tissue preservation solution, and transport to the laboratory at 4℃;

[0064] 2. Add 1×PBS solution to a 6-well plate. Remove the tissue from the tissue preservation solution and wash it in PBS to remove residual blood and mucus. Take an appropriate amount of the washed sample and add it to the dissociation solution. Use surgical scissors to cut the tissue into 1mm pieces. 3 Left and right tissue blocks;

[0065] 3. Place the dissociation solution containing the tissue in a constant temperature metal bath and incubate at 37°C for 15-40 minutes until the tissue is completely dissociated. Mix the dissociation solution with a wide-mouth pipette tip every 10 minutes. Large pieces of tissue can be cut into smaller pieces with scissors.

[0066] 4. After complete dissociation, gently mix the dissociation solution and filter it using a 40-70μm sterile cell sieve;

[0067] 5. Rinse the cell sieve with 5 mL of 1640 medium + 2% FBS;

[0068] 6. Mix the filtered cell suspension well and centrifuge at 300g for 5 minutes at 4°C; discard the supernatant.

[0069] 7. Take 1 mL of 1640 medium + 2% FBS to resuspend the cell pellet; after mixing, take 10 μL of the suspension for AO / PI staining and cell counting monitoring;

[0070] 8. Store the cells on ice or at 4°C until the next step of the experiment.

[0071] III. Obtaining and Amplifying Cellular cDNA

[0072] 1. Prepare a 10×Genomics capture and labeling cell reaction master mix;

[0073] 2. Place the PCR tubes on ice and add Master Mix and other necessary components to each reaction;

[0074] 3. Process the chip according to the official 10×Genomics instruction manual;

[0075] 4. Load the single-cell suspension onto the 10×Genomics chip;

[0076] 5. Run 10×Genomics Chromium Controller to prepare Gel Beads-in-emulsion (GEMs);

[0077] 6. Transfer the prepared GEMs to a PCR tube and place it in a PCR instrument for reverse transcription. Follow the reaction procedure shown in Table 1 below, and set the hot cap temperature to 53℃.

[0078] Table 1

[0079] Procedure Temperature Time 1 53℃ 45 min 2 85℃ 5 min 3 4℃ hold

[0080] 7. After reverse transcription is completed, Dynabeads Cleanup Mix is added to purify the reverse transcription product, and the purified first strand of cDNA is obtained;

[0081] 8. The first strand of cDNA is amplified, and the amplification reaction system shown in Table 2 below is prepared, and the amplification is performed according to the reaction procedure shown in Table 3 below.

[0082] Table 2

[0083] Components Volume (ul) cDNA 35 Amp Mix 50 cDNA Primers 15 H2O 0

[0084] Table 3

[0085]

[0086] 9. The SPRIselect reagent kit is used to purify the amplification product, and the double-stranded cDNA amplification product is obtained.

[0087] Four, construction of cDNA library and sequencing based on ONT sequencing platform

[0088] 1. 500 ng of the cDNA amplified is subjected to end repair, and the amplification reaction system is shown in Table 4 below. Incubation conditions: 20°C for 30 min, 65°C for 20 min, and 4°C storage.

[0089] Table 4

[0090] Components Volume DNA 48 ul NEBNext FFPE buffer (NEB) 3.5 ul NEBNext End-prep buffer (NEB) 2 ul FFPE DNA repair Mix (NEB) 2 ul End-prep enzyme Mix (NEB) 3.5 ul Total 60 ul

[0091] 2. The repaired product is purified using 1 volume of magnetic beads, 61 μl of Elution buffer is used to elute the DNA, and 1 μl of the purified product is used for Qubit quantification;

[0092] 3. The sequencing adapter ligation system is shown in Table 5 below; incubation conditions: 25°C, 60 min.

[0093] Table 5

[0094] Components Volume DNA after repair 60 ul LNB (SQK-LSK114) 25 ul Ligation Adapter Mix (SQK-LSK114) 5 ul Quick T4 DNA ligase (NEB) 10 ul Total 100 ul

[0095] 4. The DNA is purified using 0.4 volume of magnetic beads, 25 ul of Elution buffer (SQK-LSK114) is used to elute the DNA, 1 ul of the library is used for Qubit quantification;

[0096] 5. The on-machine library system is configured as shown in Table 6 below, and the cDNA library is constructed.

[0097] Table 6

[0098] Components Volume Library 24 ul Library Buffer (SQK-LSK114) 51 ul Squencing Buffer (SQK-LSK114) 75 ul Total 150 ul

[0099] 6. Sequencing by machine, loading the library into R10.4.1 sequencing chip, sequencing for 72h by PromethION sequencer (Oxford Nanopore Technologies, Oxford, UK), and obtaining the three-generation raw data.

[0100] V. Quality control processing

[0101] In the above step four, the three-generation raw data of the off-machine data of Nanopore sequencing is in the pod5 format containing all raw sequencing signals. After base calling by Dorado software (version: 0.7.2), the data in the pod5 format is converted into the fastq format, and the three-generation effective data is obtained.

[0102] VI. Data splitting, alignment and quantification

[0103] The present embodiment adopts the software ONT-SC-pipeline to process the data, including data splitting, alignment and quantification, and the flow chart is shown in Figure 2 Specifically:

[0104] The vsearch function is used to extract the adaptor and obtain the raw data, filter the low-quality reads (i.e. remove the reads not containing the adaptor), complete the data splitting, and obtain the clean data (splitting result).

[0105] In combination with the reference genome, the minimap2 is used to align the clean data, the sringtie is used to identify the new transcripts, align the transcripts, and extract the reads ID, transcript and gene relationship. At the same time, the vsearch is used to identify the Barcode and UMI of the clean data, cluster the Barcode and UMI, re-allocate the Barcode and UMI, and extract the reads ID, Barcode, UMI relationship.

[0106] The relationships of the reads ID, gene and transcript, Barcode and UMI are combined and calculated, and the quantification data, i.e. the expression amount of the gene and transcript, is obtained.

[0107] The software ONT-SC-pipeline specifically adopts the nextflow workflow management system to integrate methods such as vsearch, seqkit, and minimap, so that they work together. Compared with traditional shell or python scripts, the embodiment solves the problems of multi-threading implementation difficulty and slow running speed, realizes the rapid realization from fastq raw data to gene expression, and has good scalability and can be well linked with other processes.

[0108] The basic information of the filtered full-length sequence data after quality control is shown in Table 7. The alignment results and quantitative results are shown in Tables 8 and 9, respectively.

[0109] Table 7: Full-length sequence statistics table after quality control (excerpts)

[0110] sample n_reads rl_mean n_fl fl_rate (%) A1 508909 982 465435 91.45

[0111] In Table 7, sample: sample name; n_reads: total read number; rl_mean: average length; n_fl: full-length read number; fl_rate: full-length read proportion.

[0112] Table 8: Alignment results table (excerpts)

[0113] Sample Total reads mapped reads map_rate unmapped reads unmapped_rate A2 114,234 114,234 100.00% 0 0.00%

[0114] In Table 8, sample: sample name; Total reads: all reads number; mapped reads: aligned reads number; map_rate: alignment rate; unmapped reads: unaligned reads number; unmapped_rate: unaligned rate.

[0115] Table 9: Quantitative results table (excerpts)

[0116]

[0117] In Table 9, column 1 is the gene name, and the other columns are cell numbers.

[0118] Some cells in the captured cells have low activity or are even dead cells. Generally, cell filtering will be performed based on the number of genes expressed by the cells. The number of genes expressed by different sample types is inconsistent, and different threshold values will be used for specific projects. Common threshold values are 200 or 500; the mitochondrial gene filtering threshold is generally set to 0.1 and 0.25. As shown in Figure 3 the X-axis is the UMI number, the Y-axis is the gene number, and the number is the Pearson correlation coefficient. The filtered cell statistics are shown in Table 10.

[0119] Table 10 Cell statistics table after filtering (excerpts)

[0120] cell_ID orig.ident nCount_RNA nFeature_RNA scDblFinder.class A1_AAACCCAAGACTTCAC-1 A1 1274 881 singlet A1_AAACCCAAGATACAGT-1 A1 4771 2373 singlet A1_AAACCCAAGGGCAACT-1 A1 2418 1571 singlet A1_AAACCCAAGTCGCCCA-1 A1 6593 3228 singlet A1_AAACCCAGTCATGCAT-1 A1 1726 980 singlet

[0121] In Table 10, ell_ID: cell ID; orig.ident: sample name; nCount_RNA: total number of RNA; nFeature_RNA: number of RNA species; scDblFinder.class: number of microdroplet cells; percent.mito: mitochondrial proportion; percent.ribo: rRNA proportion; dropouts.

[0122] As shown in the left graph: the number of genes detected by single cells. In the middle graph: the number of UMIs detected by single cells. In the right graph: the percentage of mitochondrial gene expression. Figure 4

[0123] Seven, basic analysis

[0124] 1. Top gene distribution

[0125] The top N genes with high gene expression (Count) in cells are extracted for plotting. As shown in Figure 5 , the Y axis is the gene name, and the X axis is the proportion of gene count. The larger the proportion, the higher the gene expression.

[0126] 2. High variable genes

[0127] The expression matrix of the remaining cells after quality control is normalized, and then the function FindvariableFeatures is used to identify genes with high variable expression.

[0128] As shown in Figure 6 , red is the most obvious change in some genes, and the gene name is labeled on the right.

[0129] 3. Clustering and cluster analysis

[0130] In the face of numerous cells, the present embodiment clusters cells according to similarity, and then analyzes each Cluster after clustering. As shown in Figure 7 , the horizontal coordinate is the category (Cluster) obtained by clustering, and the vertical coordinate is the proportion of cells in each sample. The clustering information statistics are shown in Table 11.

[0131] Table 11 Clustering information statistics table (excerpts)

[0132] Cluster number stat (%) 0 765 17.14 1 693 15.53 2 616 13.80 3 445 9.97 4 390 8.74

[0133] 4. Correlation analysis of each cell cluster ​

[0134] The top 1000 cells with the largest standard deviation were taken for correlation analysis of each Cluster. The results are shown in FIG. 5, where different colors represent the correlation size. Figure 8

[0135] 5. Dimensionality reduction analysis

[0136] One of the significant features of high-throughput single-cell sequencing data is that the data volume is large, and a large number of cells can be reflected at a time. Therefore, it is a very important work to show the characteristics of cell data through dimensionality reduction and visualization.

[0137] In this embodiment, the commonly used graphs for single-cell dimensionality reduction are t-SNE (t-distributed Stochastic Neighbor Embedding) and UMAP (Uniform Manifold Approximation and Projection).

[0138] (1) PCA

[0139] In transcriptome research, the relationship between samples is determined by a large number of variables (gene expression). Principal component analysis is a method that recombines a large number of indicators (gene expression) with certain correlations into a new set of independent comprehensive indicators, thereby reducing the complexity of the problem, to study the relationship between gene expression and samples.

[0140] In the two-dimensional PCA analysis result, a scatter plot is shown with principal component 1 (PC1) as the X-axis and principal component 2 (PC2) as the Y-axis, and each point represents a cell. In the figure, the farther apart two cells are, the greater the difference in gene expression patterns between the two cells. Conversely, the closer the corresponding cells are, the more similar the overall expression patterns are. Therefore, PCA analysis is often used to evaluate the repeatability of cells.

[0141] If the samples are the results of different batches of processing, even the same sample will have a large difference, at this time we will use the software package harmony to correct the batch, at this time, the display figure is the corrected result.

[0142] (2) T-SNE

[0143] t-SNE (t-distributed Stochastic Neighbor Embedding) is a nonlinear dimensionality reduction technique that calculates the distance between all cells and maps high-dimensional data to low-dimensional space through probability distribution.

[0144] (3) UMAP

[0145] ​UMAP analysis, short for Uniform Manifold Approximation and Projection, is a novel dimensionality reduction algorithm. When calculating spatial distributions, it measures the distances between K neighboring points, preserving local information while maintaining global information accuracy. Therefore, it has seen increasing application in recent years.

[0146] like Figure 9 As shown, a dot represents a cell, and different colors represent different clusters.

[0147] 6. Marker gene analysis

[0148] (1) Marker gene identification analysis

[0149] The identification of marker genes is still performed using the R package Seurat, employing the function FindAllMarkers. This implementation identifies marker genes for each cell cluster. Identification involves performing marker analysis on the target cell population and all remaining cells.

[0150] The FindAllMarkers statistical method used for marker analysis (row count) is Wilcox. Marker gene criteria are: logfc > 0.5, pvalue ≤ 0.05. Marker genes are generally selected from those with a log2fc value within the cell cluster. Downstream analyses such as visualization and cell annotation can be performed based on these marker genes.

[0151] like Figure 10 As shown: the horizontal axis represents the gene with the highest avg_log2FC in each cluster, the vertical axis represents different clusters, and the color of the point represents the expression level after normalization.

[0152] like Figure 11 The color of the dots indicates the level of expression, and the base map is a UMAP map.

[0153] (2) Marker gene enrichment analysis

[0154] Enrichment analysis refers to the use of statistical methods to study whether differentially expressed genes are concentrated in certain specific functional categories. Commonly used functional categories include Gene Ontology and KEGG Pathway. In this implementation, enrichment analysis uses the hypergeometric distribution algorithm to identify GOTerms or KEGG Pathways where differentially expressed genes are significantly enriched relative to all annotated genes.

[0155] The differential gene enrichment analysis software was ClusterProfiler (version: 3.14.3). p.adjust(qvalue) is the pvalue after multiple hypothesis testing correction. The value range of p.adjust(qvalue) is [0,1]. The closer it is to zero, the more significant the enrichment.

[0156] like Figure 12 As shown, this is a bubble chart of GO enrichment. The horizontal axis represents Rich Factor (the proportion of Marker genes in this GO to all Marker genes), and the vertical axis represents Gene Ontology. The color of the dots represents p.adjust(qvalue).

[0157] Table 12 shows an excerpt of GO enrichment examples.

[0158] Table 12 GO enrichment example table (excerpt)

[0159]

[0160] ID: Gene Ontology ID; Description: Description of the Gene Ontology function; GeneRatio: Ratio of Marker genes annotated to this GO Term to all Marker genes with GO annotations; BgRatio: Ratio of all genes annotated to this GO Term to all genes with GO function annotations; pvalue: Statistical significance level index.

[0161] Generally, pvalue < 0.05 indicates significant enrichment; p.adjust: corrected pvalue; qvalue: corrected pvalue, which is different from the p.adjust correction method.

[0162] In the GO database, each term has a clear inclusion and substitution relationship. Based on these relationships, we can draw a directed acyclic graph, which allows us to see the relationships between terms more clearly.

[0163] However, due to the large number of terms, drawing a complete directed acyclic graph would be very complex. Therefore, this implementation only shows a portion of the terms. For example, an ellipse or square represents a GO term, and arrows indicate the relationships between different GO terms.

[0164] 7. Cell Annotation

[0165] Single-cell transcriptome sequencing (scRNA-seq) data analysis can identify different cell types by differences in gene expression levels between different cells, study cell heterogeneity in complex tissues, and ultimately help us understand the pathogenesis of diseases, cell lineage and differentiation trajectory, or intercellular communication mechanisms. In the actual processing of scRNA-seq data, the annotation of cell types is a key step for subsequent analysis and research.

[0166] In this embodiment, the annotation of cell types is divided into cell-based annotation and cluster-based annotation according to the annotation object, and can be divided into supervised annotation and unsupervised annotation according to whether label data is required. The cell annotation statistics are shown in Table 13.

[0167] Table 13 Cell annotation statistics table (extract)

[0168] type count lateral_root 3330 xylem 1782 branch 1142 lateral_root_cap 976 phloem 921

[0169] As shown in Figure 13 , the stacked column chart of each subpopulation, the horizontal coordinate is different samples, the vertical coordinate is the number of cells, and different colors represent different cell types.

[0170] As shown in Figure 14 , the figure is a UMAP graph, and different colors represent different cell types.

[0171] 8. Cell communication analysis

[0172] Cell communication analysis, i.e. receptor ligand analysis. Biological cell ligand-receptor mediated cell interaction regulates various bioinformatics processes such as development, differentiation and inflammation. The mutual communication between cells forms a cell regulation network, which affects the progress of various physiological processes. The tissue ecosystem (microenvironment), especially the tumor microenvironment, is composed of various types of cells, including immune cells, stromal cells, hematopoietic cells, etc. These different types of cells can play an important role in tumor occurrence and development, drug resistance, immune infiltration and inflammation through the interaction between ligands and receptors (e.g. with immune checkpoint inhibitors).

[0173] In this embodiment, software CellPhoneDB V5 is used for communication analysis. This software uses network analysis and pairing patterns to identify the main input and output signals between cells and determine their communication relationship, and provides various visualization methods for customer use.

[0174] In this embodiment, the software CellPhoneDB V5 takes cellular gene expression data as input and combines it with the phase communication of receptors and their cofactors to simulate intercellular communication. To establish intercellular communication, "cell_type" is used as the tag for communication analysis.

[0175] (1) Statistical analysis of cell communications using the CellPhoneDB V5 R package for communication analysis

[0176] The statistical results (visualized) for all cell subpopulations are: overall statistics on the number of receptor-ligand pairs and communication strength between different cell subpopulations. The size of the various colored circles on the periphery represents the number of cells; the larger the circle, the more cells. Cells pointing to arrows express ligands, and cells pointed to by arrows express receptors. The more ligand-receptor pairs, the thicker the line.

[0177] (2) Statistical analysis of cell communication intensity

[0178] like Figure 15 As shown in the visualization statistical analysis results, the size of the various colored circles on the periphery represents the number of cells; the larger the circle, the more cells there are. Cells pointing with arrows express ligands, and cells the arrows point to express receptors. The greater the probability of ligand-receptor pair communication, the thicker the line.

[0179] (3) Visualization of single-path communication

[0180] Multiple ligand-receptor pairs may reside within the same pathway. This analysis focuses on the communication between ligand-receptor pairs within a pathway across different cell populations, organized by pathway. Visualization of single-pathway communication between cell subpopulations is presented using heatmaps, chord diagrams, and network diagrams—these are merely different representations with the same meaning. In the visualization, the vertical axis represents the ligand, the horizontal axis represents the receptor, and the color represents the communication probability. The ends of the line represent the ligand-receptor pair, and the line width represents the communication probability.

[0181] (4) Visualization of ligand receptors

[0182] The pathway contains numerous ligand-receptor pairs, which may also consist of multiple subunits, each with varying degrees of influence. Therefore, the contribution of each subunit to the overall ligand-receptor pair and the subunits of individual ligand-receptor pairs within the pathway can be visualized. The x-axis represents relative contribution, and the y-axis represents different subunits. This can demonstrate the impact of each subunit on different cell types. Furthermore, it can display bubble diagrams of multiple ligand-receptor mediated cell communication, with the x-axis representing communication between cell subpopulations and the y-axis representing ligand-receptor pairs. The dot size represents the p-value, and the color indicates the communication probability.

[0183] (5) Related gene expression point maps and violin diagrams

[0184] The expression levels of receptor-related genes in the pathway can also be visualized across different cell types.

[0185] 9. Cell trajectory analysis (pseudo-time series analysis)

[0186] Pseudotime, or pseudo-time series for short, is a method for measuring the progress of cells in a biological process such as cell differentiation. It is an abstraction of a biological process, describing the distance of the shortest path between a cell's current state and its initial state in time. This implementation uses the software Monocle3 for cell trajectory analysis. Figure 16 As shown, this is the trajectory diagram after selecting the root node.

[0187] (1) Pseudo-temporal trajectory differential genes

[0188] Morans' index (morans_I) can represent spatial correlation, with values ​​ranging from -1 to 1, where 0 indicates the gene is absent. In single cells, it can represent spatial co-expression effects, where 1 indicates that the gene's expression levels are highly similar in spatially close cells.

[0189] like Figure 17 As shown, the time series diagrams of some genes in different clusters are presented, with the horizontal axis representing different time series and the vertical axis representing expression levels.

[0190] (2) Co-expression module analysis

[0191] Once you have a set of genes that vary in some way within a cluster, you can group them into modules, and the software monocle3 provides just that.

[0192] 10. Variable Shear Analysis (Advanced Analysis)

[0193] Thanks to the ability to measure complete transcripts of the full-length ONT, we were able to analyze alternative splicing of genes using different transcripts of the same gene. We used the software isoswitch for alternative splicing analysis and used the plot_assay_stats() function to create a summary plot that includes the number of genes, the number of transcripts, the distribution of isoforms, and the number of genes in each cell type, providing a concise description of the isoform distribution in the dataset.

[0194] like Figure 18 The diagram shows an example of alternative splicing. Top left: Bar chart of gene (single transcript, multiple transcript) and transcript number statistics; Middle left: Bar chart of transcript number statistics for gene; Bottom left: Scatter plot of gene transcript number vs. major transcript number; Right: Box plot of gene expression in different clusters.

[0195] 11. Fusion gene analysis

[0196] Full-length transcript sequencing technology provides a powerful tool for revealing the structural information at the transcript level, and also opens the door for detecting gene fusion events at the single-cell level. Gene fusion, i.e., the recombination of part or all of the sequences of two or more genes to form a new hybrid gene, can be caused by various mechanisms such as chromosomal translocation, deletion or inversion, and plays a crucial role in the occurrence and development of diseases.

[0197] In this embodiment, software JAFFA is used for fusion gene analysis. By aligning the clean reads obtained by sequencing to the reference genome and combining the gene annotation information, we can extract sequence fragments spanning two genes and further identify potential fusion genes. In order to ensure the reliability of the results, we set a coverage depth threshold (default: cutoff≥5) to filter low-quality data, and finally obtain the following fusion gene information, as shown in Table 14 and Table 15. Figure 19

[0198] Table 14 (extract)

[0199]

[0200] The above specific embodiments describe the implementation of the present application in detail, but the present application is not limited to the specific details in the above embodiments. Within the scope of the claims and technical concepts of the present application, various simple modifications and changes can be made to the technical solutions of the present application, and these simple modifications all belong to the protection scope of the present application.​

Claims

1. A sequencing analysis method for single-cell transcriptomes based on ONT sequencing, characterized in that, Includes the following steps: Prepare and construct a cDNA library of the sample to be tested, sequence it, and obtain the third-generation raw data of the sample to be tested; The three generations of raw data are subjected to quality control to obtain the three generations of valid data for the sample to be tested; Functional annotation of genes and transcripts in the reference genome; By combining the three generations of valid data and the gene data of the reference genome, the ONT-SC-pipeline software was used to split, compare, and quantify the data to obtain quantitative data. The quantitative data is subjected to dual-cell identification and removal, and further filtered. The filtered data were subjected to basic analysis, marker gene enrichment analysis, and advanced analysis in sequence. The basic analysis included top gene distribution analysis, hypervariable gene analysis, clustering analysis, correlation analysis of various cell groups, dimensionality reduction analysis, and marker gene identification. The advanced analysis included cell annotation analysis, cell communication analysis, cell trajectory analysis, alternative splicing analysis, and fusion gene analysis. In the ONT-SC-pipeline software, the vsearch function is used to extract the adapter and obtain the raw data, filter low-quality reads, and complete data splitting to obtain clean data. Combined with the reference genome, minimap2 is used to align the clean data, sringtie identifies new transcripts, compares transcripts, and extracts read IDs, transcripts, and gene relationships. vsearch is used to identify the barcodes and UMIs of the clean data, clusters the barcodes and UMIs, reassigns the barcodes and UMIs, and extracts the read ID, barcode, and UMI relationships. The relationship between reads ID, gene and transcript, barcode, and UMI is combined and calculated to obtain quantitative data, namely the expression levels of genes and transcripts.

2. The sequencing analysis method for single-cell transcriptome based on ONT sequencing according to claim 1, characterized in that, The quality control process is as follows: sequences with a length ≤30 bp and a quality value ≤5 are filtered out based on sequencing quality values.

3. The sequencing analysis method for single-cell transcriptome based on ONT sequencing according to claim 1, characterized in that, The method for data splitting, comparison and quantification using the software ONT-SC-pipeline is as follows: vsearch is used to split the data, and minimap2 is used to compare the data. The information from the splitting results and the comparison results is extracted and merged to obtain the quantitative data.

4. The sequencing analysis method for single-cell transcriptome based on ONT sequencing according to claim 1, characterized in that, The scDblFinder software was used to identify and remove double-celled and low-activity cells, and to eliminate cells with a high proportion of ribosomal RNA and a high proportion of mitochondrial genes.

5. The sequencing analysis method for single-cell transcriptome based on ONT sequencing according to claim 1, characterized in that, Top gene distribution analysis was performed using SeuratV5 software; highly variable gene analysis was performed using the FindvariableFeatures function; clustering analysis, dimensionality reduction analysis, and marker gene identification were also performed using SeuratV5 software.

6. The sequencing analysis method for single-cell transcriptome based on ONT sequencing according to claim 5, characterized in that, The dimensionality reduction analysis is specifically performed using at least one of the PCA, T-SNE, and UMAP analysis methods in Seurat V5 software.

7. The sequencing analysis method for single-cell transcriptome based on ONT sequencing according to claim 6, characterized in that, If there are multiple samples to be tested, the Harmony software is used instead of the PCA analysis method to correct and integrate the filtered single-cell data from the multiple samples.

8. The sequencing analysis method for single-cell transcriptome based on ONT sequencing according to claim 1, characterized in that, The hypergeometric distribution algorithm was used for enrichment analysis of marker genes.

9. The sequencing analysis method for single-cell transcriptome based on ONT sequencing according to claim 1, characterized in that, Cell annotation analysis was performed using known marker genes and SCINA software; cell communication analysis was performed using CellPhoneDB V5; cell trajectory analysis was performed using Monocle3 software; alternative splicing analysis was performed using Isoswitch software; and fusion gene analysis was performed using JAFFA software.

10. The sequencing analysis method for single-cell transcriptome based on ONT sequencing according to claim 9, characterized in that, Based on CellPhoneDB V5 results, the communication results of single pathways among various cell subpopulations were visualized using ggplot2 and igraph methods; the contribution of each subunit to the overall ligand and receptor and the subunits of individual ligand and receptor pairs in the pathway were visualized.

Citation Information

Patent Citations

  • Method for detecting pathogens based on bioinformatics of Nanopore metagenome RNA-seq

    CN113265452A

  • Analysis method and system of parametric transcriptome sequencing data adaptive to ONT sequencing

    CN116013415A