A method, device and electronic device for obtaining tissue space transcript information

By combining long read and long sequencing and spatial shearing transcriptomics, nanopore sequencing and spatial barcode/UMI technology, the problem of difficulty in comprehensively exploring the organization of spatial transcript information in the existing technology is solved, and a higher quality and complete transcript data acquisition is achieved.

CN117012278BActive Publication Date: 2025-05-16CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310994114.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-08
Publication Date
2025-05-16
Estimated Expiration
2043-08-08

AI Technical Summary

Technical Problem

The prior art is difficult to fully explore the spatial sequence heterogeneity and RNA editing of full-length cleavage books in tissues in a spatial context, especially in human tissues and tumor samples, which lack spatial analysis capabilities at isomer levels.

Method used

Long read and long sequencing technology combined with spatial shear transcriptomics, and full-length transcripts were obtained through nanopore sequencing, and spatial barcode and UMI allocation technology were used to enhance spatial information at this level.

Benefits of technology

It realizes a more comprehensive and accurate acquisition of tissue spatial transcript information, simplifies the identification and quantification of co-education books, reduces the loss of full-length transcript information, and improves data quality and sequencing saturation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117012278B_ABST
    Figure CN117012278B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and electronic device for obtaining tissue spatial transcript information, and relates to the field of bioinformatics. The method provided by the present invention adds a long-read sequencing step compared to the existing spatial transcriptome technology, thereby increasing the spatial information at the shear level. The identification and quantitative analysis of the entire shear are greatly simplified, which is conducive to quickly obtaining tissue spatial transcript information. In addition, the quality of the transcript data obtained by long-read sequencing using the method provided by the present invention is improved, and a more complete transcript can be obtained, which greatly reduces the loss of full-length transcript information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and in particular to a method, device and electronic device for obtaining tissue spatial transcript information. Background Art

[0002] As an emerging field of drug discovery, targeting alternative splicing is becoming one of the main strategies. Several characteristics of cancer development and metastasis are associated with abnormal splicing patterns, such as unrestricted proliferation, escape from growth inhibitory signals, dysregulated metabolic processes and cellular phenotypes [1]. With the advent of single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics, the relationship between expression heterogeneity in different cell types, states and locations and tumorigenesis can be easily resolved [2]. However, short-read sequencing limits the comprehensive analysis of alternative splicing; although exon junctions can be observed, quantitative analysis of the entire splice version is extremely challenging [3–5].

[0003] 10x Genomics' Visium platform is well suited for spatial transcriptome analysis because in this approach, cDNA from tissue sections is synthesized in situ on spatially barcoded slides. The resulting sequencing reads therefore contain these spatial barcodes and can be assigned to coordinates on the slide. Evidence is currently available to characterize spliceotypes and to identify differences in spliceotype usage by combining scRNA-seq and long-read sequencing technologies (PacBio or Oxford Nanopore sequencing).

[0004] A recent study proposed a hybrid approach in which single-cell transcript splices were sequenced by PacBio to derive their expression levels. The spatial coordinates of the resulting splices were then inferred by shallow nanopore sequencing of spatial transcriptomics (Visium) libraries. However, low coverage and the lack of unique molecular identifiers (UMIs) prevented extensive exploration of spatial sequence heterogeneity of full-length splices in tissues and RNA editing. Another study recently performed full-length spatial transcriptomics of mouse heart using the Visium system. However, insufficient coverage of the 5′ end of the transcripts led the authors to assemble full-length cDNAs from partial transcript sequences without using UMIs for expression quantification [6–8]. However, spatial analysis of isoform levels in human tissues, even tumor samples, has not been reported or documented. Esophageal squamous cell carcinoma (ESCC) is characterized by oncogenic splicing events [9–11]. Esophageal cancer (EC) ranks seventh among the most common cancers worldwide and sixth among the leading causes of cancer death in 2020

[12] . EC can be divided into two subtypes: esophageal adenocarcinoma and ESCC. ESCC accounts for 90% of esophageal cancer cases. The 5-year survival rate of ESCC is approximately 10%-15%. Although previous studies have comprehensively characterized the landscape of AS in ESCC, the spatial information of isoforms is still unclear

[12] . So far, only transcript quantitative analysis has been reported in single-cell transcriptomes, and ScNaUmi-seq technology has been mainly used to analyze long-read single-cell RNA sequencing data.

[0005] ScNaumi-seq technical flow chart reference Figure 1 As shown in the figure, its principle and basic steps are as follows: 1) After single cell dissociation, the 10x3' RNA kit is used to complete cell sorting and reverse transcription to obtain full-length cDNA; 2) Part of the full-length cDNA is used for conventional second-generation sequencing; 3) Part of the full-length cDNA is amplified by PCR to construct a nanopore library and sequenced; 4) Second-generation sequencing is used to correct long read data (cellbarcode splitting and UMI correction); 5) Transcripts are identified through Gencode; 6) Single-cell gene expression is quantified and major cell types are identified through the second-generation single-cell transcriptome basic bioinformatics process

[13] .

[0006] Recent technological advances have made it possible to perform high-throughput quantification of gene expression in a spatial context. These methods can be roughly divided into two types, one that detects the presence of predefined target genes and the other that makes observations by sampling the entire transcriptome. The latter is required for a priori exploratory analysis and new hypothesis generation. This approach is usually based on the in situ capture of polyadenylated RNA on spatially barcoded reverse transcription primers, which allows the capture of all mRNAs in the transcriptome. The captured transcripts are then sequenced ex situ and their spatial barcodes are used to infer their spatial origin. Although several new large-scale transcriptome analysis methods based on positional capture have emerged recently, they only evaluate transcripts as 3' cDNA tags, not as complete transcripts. The fundamental reason behind this is that all of these methods are based on short-read library construction and sequencing, which means the loss of full-length transcript information. Since a large amount of transcriptome diversity originates from post-transcriptional modifications, only the characterization of the full-length sequence of transcripts can truly comprehensively describe the transcriptome.

[0007] In view of this, the present invention is proposed. Summary of the invention

[0008] The object of the present invention is to provide a method, device and electronic device for obtaining tissue spatial transcript information to solve the above technical problems.

[0009] The present invention is achieved in that:

[0010] In a first aspect, the present invention provides a method for obtaining tissue spatial transcript information, comprising the following steps:

[0011] Long-read sequencing: First, purify the full-length cDNA sequence of the tissue to be tested in the spatial transcriptome library, then amplify the purified full-length cDNA sequence using primers to generate materials for preparing nanopore sequencing libraries, and use a nanopore sequencing library kit to prepare the nanopore sequencing library; and perform PCR amplification on the nanopore sequencing library;

[0012] Oxford Nanopore data processing: PCR amplified cDNA was ligated to the adapters in the Oxford Nanopore kit to generate chimeric cDNA; the chimeric cDNA was then scanned for the internal template switching oligonucleotide (TSO), the 3' adapter sequence connected by poly(T), and then the poly(A / T) tail and 3' adapter sequence of all reads; the scanned reads were aligned to the reference genome of the sample to be tested using bioinformatics software, and spatial barcodes and UMIs were assigned to nanopore reads using the strategy and software described for single-cell libraries; the SAM records for each spatial point and gene were then grouped by UMI; the consensus sequence was calculated using SPOA based on the number of available reads for the UMI; the sequence between the end of the TSO and the base before the polyA sequence was used; the Consensus cDNA sequence was aligned to the sequence of the reference genome of the sample to be tested using bioinformatics software in a spliced ​​alignment manner; and the SAM records matching known genes were analyzed to match Gencode vM24 transcript isoforms.

[0013] The inventors applied spatial splice transcriptomics, which is an unbiased method based on spatial in situ capture, combined with long read sequencing to detect and quantify the spatial expression of splice variations. Using nanopore sequencing, the read length is equal to the fragment length, which means that the entire transcript can be sequenced in a single read. Full-length transcripts with a length > 20kb have been sequenced by a single read. This technology adds a long read sequencing step compared to existing spatial transcriptome technologies and thereby increases the spatial information at the splice level. The identification and quantification of the entire isoform are greatly simplified, which is conducive to the rapid acquisition of tissue spatial transcript information. In addition, the transcript data quality of the long read sequencing obtained by the method provided by the present invention is good, the average read length is close to 1000bp, and the longest single length exceeds 330kbp. Therefore, the method provided by the present invention can obtain a more complete transcript and greatly reduce the loss of full-length transcript information.

[0014] Nanopore sequencing as described above is a more attractive option that can generate sufficient numbers of reads to achieve sequencing saturation required for comprehensive transcriptome and sequence heterogeneity exploration.

[0015] In a preferred embodiment of the present invention, the purification method comprises: firstly amplifying the full-length cDNA sequence of the tissue to be tested in the spatial transcription library using a biotin-labeled primer, and then oscillating and washing the biotinylated cDNA amplification product to obtain a purified full-length cDNA sequence, and the product is used for subsequent amplification to generate material for preparing a nanopore sequencing library.

[0016] Nanopore sequencing of full-length cDNA libraries prepared from the 10x Genomics workflow generates 20-50% of reads without 3' sequences, thus lacking spatial barcodes and UMIs. To deplete these fragments, the inventors employed biotin labeling to screen cDNAs containing biotinylated 3' primers, thereby obtaining reads with spatial barcodes and UMIs.

[0017] In an optional embodiment, the number of amplification cycles for amplifying the full-length cDNA sequence of the tissue to be tested in the spatial transcriptome library using biotin-labeled primers is 3-6 cycles, and the amplification primer sequences are shown in SEQ ID NO.1-2.

[0018] In a preferred embodiment of the present invention, the number of cycles for amplifying the purified full-length cDNA sequence is 8-10 cycles, such as 8, 9 or 10 cycles.

[0019] In an alternative embodiment, the Oxford Nanopore kit is selected from the LSK-109 or LSK-110 kit.

[0020] In a preferred embodiment of the present invention, after assigning UMIs to nanopore reads, before UMI grouping, the process also includes removing low-quality mapped reads (mapqv=0) and potential chimeric reads (terminal soft / hard clipping>150 nt).

[0021] In a preferred embodiment of the present invention, the ComputeConsensus sicelore-2.0 method and steps are used to calculate the consensus sequence (UMI) of each molecule based on the number of available reads of the UMI; for molecules with more than two reads (RN>2), SPOA is used to calculate the consensus sequence, and the sequence between the bases before the TSO end (SAM Tag: TE) and the polyA sequence (SAM Tag: PE) will be used by the SPOA method to calculate the consensus sequence. The quality value of the consensus nucleotide is -10 *log10 (n Reads do not meet the consensus nucleotide / n Reads total).

[0022] In a preferred embodiment of the present invention, the tissue to be tested is selected from solid tumor tissue or lymphatic metastasis tissue; the bioinformatics software is selected from minimap2 v2.17, and in other embodiments, it can also be adjusted to other bioinformatics analysis software as needed.

[0023] To assign UMIs to Gencode transcripts, an exact match between the UMI and the Gencode transcript exon-exon junction layout is required, authorizing exon boundaries to have two base edges of sequence added or missing to allow for indexing at exon junctions and imprecise mapping by minimap2.

[0024] In a second aspect, the present invention also provides a data processing method, which comprises: using an R software package to process the original gene expression matrix generated by SpaceRanger to create a Seurat object, adding cell-related information, and performing data processing; the cell-related information comprises: data containing gene-level nanopore long reads; containing tissue spatial transcript information obtained by the above-mentioned method for obtaining tissue spatial transcript information;

[0025] In an optional embodiment, the cell-related information also includes gene-level information from the short read output.

[0026] The R software package is selected from at least one of Bioconductor (version 4.0.2) and Seurat (version 24) software packages (version 3.9.9).

[0027] In a third aspect, the present invention further provides a device for obtaining tissue spatial transcript information, comprising:

[0028] Long-read sequencing module: used to prepare nanopore sequencing libraries and amplify nanopore sequencing libraries;

[0029] Oxford Nanopore Data Processing Module: used to read and calibrate the amplified products of the Nanopore sequencing library, then align the sequences with the reference genome, and analyze the SAM records that match the known genes to match the GencodevM24 transcript isoforms;

[0030] Data reading includes reading all reads that contain both spatial barcodes and UMIs.

[0031] Correction refers to the UMI grouping of the SAM records for each spatial point and gene after removing low-quality mapping reads (mapqv = 0) and potential chimeric reads (terminal soft / hard clipping > 150nt). That is, using short-read sequencing data to match the UMI and barcode in the long-read sequencing data makes the reads containing UMI and barcode in the long-read sequencing more accurate.

[0032] In other embodiments, the apparatus further comprises a data processing module for processing the original gene expression matrix generated by SpaceRanger through an R software package to create a Seurat object, and adding cell-related information for data processing.

[0033] The cell-related information includes: data containing long nanopore reads at the gene level; and tissue space transcript information obtained by the above-mentioned method for obtaining tissue space transcript information.

[0034] In a fourth aspect, the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned method for obtaining tissue spatial transcript information when executing the program.

[0035] In addition, the electronic device further includes a bus and a communication interface, and the memory, processor and communication interface are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these elements can be electrically connected to each other via one or more buses or signal lines. The processor can process information and / or data related to target identification to perform one or more functions described in the present application.

[0036] The memory can be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0037] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0038] In a fifth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned method for obtaining tissue spatial transcript information when the computer program is executed by a processor.

[0039] The present invention has the following beneficial effects:

[0040] The present invention applies spatial shear transcriptomics, which is an unbiased method based on spatial in situ capture, combined with long-read sequencing to detect and quantify the spatial expression of shear variants. Using nanopore sequencing, the read length is equal to the fragment length, which means that the entire transcript can be sequenced in a single read. The method provided by the present invention adds a long-read sequencing step compared to existing spatial transcriptomics technology and thereby increases the spatial information at the shear level. The identification and quantification of the entire isoform is greatly simplified, which is conducive to the rapid acquisition of tissue spatial transcript information.

[0041] In addition, the quality of transcript data obtained by long-read sequencing using the method provided by the present invention is improved, more complete transcripts can be obtained, and the loss of full-length transcript information is greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0043] Figure 1 A technical flowchart of ScNaUmi-seq that combines ONT and single-cell transcriptomics technology;

[0044] Figure 2 A flow chart and a data processing flow chart for obtaining tissue spatial transcript information provided by the present invention;

[0045] Figure 3 Schematic diagram of the device for obtaining tissue spatial transcript information provided by the present invention.

[0046] Figure numbers: 10-long read sequencing module; 20-Oxford Nanopore data processing module; 30-data processing module. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the technical scheme in the embodiments of the present invention will be described clearly and completely below. If the specific conditions are not specified in the embodiments, they are carried out according to conventional conditions or conditions recommended by the manufacturer. If the manufacturer of the reagents or instruments used is not specified, they are all conventional products that can be purchased commercially.

[0048] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.

[0049] Example 1

[0050] This embodiment provides a method for obtaining tissue spatial transcript information and data processing, and its flowchart is as follows: Figure 2 As shown, it includes the following steps:

[0051] 1. 10X spatial transcriptome library

[0052] After cryoembedding of solid tumor tissue or lymph node metastasis tissue, the permeabilization conditions of esophageal tissue were optimized using Visium Spatial Tissue Optimization Slides and Kits (10x Genomics, Pleasanton, CA, USA). Spatial barcoded full-length cDNA was then generated using Visium Spatial Gene Expression Slides and Kits (10x Genomics).

[0053] In this embodiment (sample ID), samples A and B are solid tumor tissues, and C is lymph node tissue.

[0054] The cDNA amplification procedure is shown in Table 1 below, and the number of cDNA amplification cycles of 10X Genomics spatial transcriptome is shown in Table 2. In this example, the cDNA library of each sample was divided into two parts, one part will be used for long-read sequencing (Oxford Nanoporesequencing), and the other part will be used for short-read sequencing (NextSeq500 (Illumina)). The raw sequencing data was processed using the pre-launch (version 4509.7.5) of the Space Ranger pipeline (10x Genomics) and mapped to the mm10 genome assembly.

[0055] Table 1: 10X Genomics spatial transcriptome cDNA amplification time and temperature

[0056]

[0057] Table 2: 10X Genomics spatial transcriptome cDNA amplification cycle

[0058]

[0059] 2. Long read sequencing (Oxford Nanopore sequencing)

[0060] Nanopore sequencing of full-length cDNA libraries prepared from the 10x Genomics workflow generates 20-50% of reads without 3' sequence and therefore lacking spatial barcodes and UMIs. To deplete these fragments, we initially selected cDNAs with biotinylated 3' primers. 10 ng of 10x Genomics Visium PCR product was amplified for 5 cycles with 5'-AAGCAGTGGTATCAACGCAGAGTACAT-3' (SEQ ID NO.1) and 5'-biotin-aaaaactacacgctcttccgatct -3' (SEQ ID NO.2).

[0061] Excess biotinylated primers were removed by 0.55x SPRIselect (Beckman Coulter) purification, and biotinylated cDNA (biotinylated cDNA mixed with 40 μl EB, Qiagen) was bound to 15 μl 1× SSPE-washed Dynabeads™ M-270 Streptavidin beads (Thermo), shaken in 10 μl 5× SSPE at room temperature for 15 min, washed twice with 100 μl 1× SSPE, and once with 100 μl EB.

[0062] The beads were suspended in 100 μl of 1× PCR mix and amplified for 8 cycles with primers NNNAAGCAGTGGTATCAACGCAGAGTACAT and nnnctacacgacctcttccgatct to generate sufficient material (1-2 μg) for preparation of nanopore sequencing libraries.

[0063] Nanopore sequencing libraries were prepared using the Oxford Nanopore LSK-109 kit (1 μg cDNA) following the manufacturer’s instructions. PromethION flow cells were loaded with 200 ng of library each.

[0064] The prepared nanopore library was subjected to PCR amplification using Kapa HiFi Hotstart polymerase (Roche Sequencing Solutions): initial denaturation at 95°C for 3 min; cycles: 98°C for 30 s, 64°C for 30 s, 72°C for 5 min; final extension: 72°C for 10 min, with a primer concentration of 1 μM.

[0065] 3. Oxford Nanopore Data Processing

[0066] Nanopore read processing was performed according to the scNaUmi-seq protocol with modifications. Amplified cDNAs were ligated with Oxford Nanopore adapters (LSK-109 or LSK-110 kits) to generate 3–20% chimeric cDNAs, where two or more cDNA molecules are ligated together, to a much lesser extent. cDNAs from the 10xGenomics Visium system were ligated with a template-switching oligonucleotide (TSO, AAGCAGTGGTATCAACGCAGAGTACAT) at the 5’ and a poly(A) followed by an adapter sequence (CTACACGACGCTCTTCCGATCT) at the 3’. Junctions of individual cDNAs in chimeric reads were characterized by the presence of those 5’ and / or 3’ end sequences that were ligated together in the read sequence.

[0067] To split reads from chimeric cDNAs, this example first scanned the internal (>200 nucleotides from the end) TSO and the 3' adapter sequence connected by poly(T) (poly(T)-adapter). When two adjacent poly(T)-adapters, two TSOs, or one TSO adjacent to a poly(T)-adapter were found, thus finding the junction of two cDNA molecules, the read was split into two independent reads.

[0068] Next, all reads were scanned for poly(A / T) tails and 3' adapter sequences to determine the orientation and strand specificity of the reads.

[0069] The scanned reads were then aligned to Mus musculus mm10 using Minimap2 v2.17 in spliced ​​alignment mode. Spatial barcodes and UMIs were then assigned to the Nanopore reads using the strategy and software previously described for single-cell libraries.

[0070] After removing low-quality mapped reads (mapqv = 0) and potential chimeric reads (end soft / hard clipping > 150 nt), SAM records for each spatial point and gene were grouped by UMI.

[0071] The consensus sequence (UMI) of each molecule was calculated based on the number of available reads for the UMI using the ComputeConsensus sicelore-2.0 steps and methods. For molecules with more than two reads (RN>2), the consensus sequence was calculated using SPOA, using the sequence between the bases before the end of the TSO (SAM Tag: TE) and the polyA sequence (SAM Tag: PE). The quality value of the consensus nucleotide was -10 *log10 (n reads that did not meet the consensus nucleotide / n total number of reads).

[0072] Consensus cDNA sequences were aligned to the Mus musculus mm10 sequence in a spliced ​​alignment using minimap2 v2.17. SAM records matching known genes were analyzed for matches to Gencode vM24 transcript isoforms (same exon composition).

[0073] To assign a UMI to a Gencode transcript, an exact match between the UMI and the exon-exon junction layout in the Gencode transcript database is required, with the exception of an extra or missing two bases at the edges of the exon sequence so that exon junctions can be imprecisely mapped with minimap2.

[0074] 4. Data Processing

[0075] The original gene expression matrix generated by SpaceRanger was processed using R / Bioconductor (version 4.0.2) and Seurat (version 24) software packages (version 3.9.9). After creating the Seurat object, this embodiment adds three cell-related information: (i) "GENE" contains gene-level information from short read output, (ii) "ISOG" contains gene-level nanopore long read data, and (iii) "ISO" contains isoform-level transcription information, in which only the spatial transcription of the molecular follow-up data of all exons is retained. Subsequent data processing will be the same as the steps and methods of spatial transcriptome analysis at the gene level.

[0076] The spatial information after spatial transcribed cDNA long read sequencing is shown in Table 3, and the data quality after spatial transcribed cDNA long read sequencing is shown in Table 4. As can be seen from the table, this embodiment successfully detected the number of reads from 600000000 to 95000000. Although the reads with both UMI and coding markers are less than 50%, the final average read length reaches 969.2 bp and the longest reaches 338 kbp.

[0077] Table 3: Spatial information after long-read sequencing of spatially transcribed cDNA

[0078]

[0079] Table 4: Data quality after long-read sequencing of spatially transcribed cDNA

[0080]

[0081] Example 2

[0082] This embodiment provides a device for obtaining tissue spatial transcript information, referring to Figure 3 As shown, it includes:

[0083] Long-read sequencing module 10: used to prepare nanopore sequencing libraries and amplify the nanopore sequencing libraries.

[0084] Oxford Nanopore Data Processing Module 20: used to read and calibrate the amplified products of the Nanopore sequencing library, and then perform sequence alignment with the reference genome, and analyze the SAM records matching the known genes to match the GencodevM24 transcript isoforms;

[0085] Data reading includes reading all reads that contain both spatial barcodes and UMIs.

[0086] Data processing module 30: used to process the original gene expression matrix generated by Space Ranger through the R software package to create a Seurat object, and add cell-related information for data processing.

[0087] Specifically, the functions of the long-read sequencing module 10 include: first purifying the full-length cDNA sequence of the tissue to be tested in the spatial transcription library, then using primers to amplify the purified full-length cDNA sequence to generate materials for preparing a nanopore sequencing library, using a nanopore sequencing library kit to prepare a nanopore sequencing library; and performing PCR amplification on the nanopore sequencing library.

[0088] The functions of the Oxford Nanopore data processing module 20 include: ligating PCR amplified cDNA with adapters from the Oxford Nanopore kit to generate chimeric cDNA; then scanning the chimeric cDNA for the internal template switching oligonucleotide (TSO), the 3' adapter sequence connected by poly(T), and then scanning all reads for the poly(A / T) tail and 3' adapter sequence; aligning the scanned reads to the reference genome of the sample to be tested using minimap2 v2.17, and then assigning spatial barcodes and UMIs to Nanopore reads using the strategy and software described for single-cell libraries; then grouping the SAM records for each spatial point and gene by UMI; calculating the consensus sequence using SPOA based on the number of available reads for the UMI; using the sequence between the end of the TSO and the base before the polyA sequence; aligning the Consensus cDNA sequence to the sequence of the reference genome of the sample to be tested using minimap2 v2.17 in a spliced ​​alignment; and analyzing SAM records that match known genes for matching Gencode vM24 transcript isoforms.

[0089] The functions of the data processing module 30 include: using R / Bioconductor (version 4.0.2) and Seurat (version 24) software packages (version 3.9.9) to process the original gene expression matrix generated by Space Ranger. After creating the Seurat object, this embodiment adds three cell-related information: (i) "GENE" contains gene-level information from short read output, (ii) "ISOG" contains gene-level nanopore long read data, and (iii) "ISO" contains isomer-level transcription information, in which only the spatial transcription of the molecular follow-up data of all exons is observed is retained. The subsequent data processing will be the same as the steps and methods of spatial transcriptome analysis at the gene level.

[0090] References:

[0091] 1.Zhao, S. Alternative splicing, RNA-seq and drug discovery. DrugDiscov Today 24, (2019).

[0092] 2.Marzese, D. M., Manughian-Peter, A. O., Orozco, J. I. J.&Hoon, D.S. B. Alternative splicing and cancer metastasis: prognostic and therapeuticapplications. Clin Exp Metastasis 35, 393–402 (2018).

[0093] 3.E, B. et al. Spatial maps of prostate cancer transcriptomes revealan unexplored landscape of heterogeneity. Nat Commun 9, (2018).

[0094] 4.A, L.-J., X, T., VN, K.&Z, Y. Spatial transcriptomics inferred frompathology whole-slide images links tumor heterogeneity to survival in breastand lung cancer. Sci Rep 10, (2020).

[0095] 5.R, M. et al. Integrating microarray-based spatial transcriptomicsand single-cell RNA-seq reveals tissue architecture in pancreatic ductaladenocarcinomas. Nat Biotechnol 38, 333–342 (2020).

[0096] 6.Philpott, M. et al. Nanopore sequencing of single-celltranscriptomes with scCOLOR-seq. Nat Biotechnol 39, 1517–1520 (2021).

[0097] 7.Shi, Z.-X. et al. High-throughput and high-accuracy single-cell RNAisoform analysis using PacBio circular consensus sequencing. NatureCommunications 2023 14:1 14, 1–13 (2023).

[0098] 8.Gupta, I. et al. Single-cell isoform RNA sequencing characterizesisoforms in thousands of cerebellar cells. Nat Biotechnol 36, 1197–1202(2018).

[0099] 9.Dlamini, Z. et al. Prognostic Alternative Splicing Signatures inEsophageal Carcinoma. Cancer Manag Res 13, 4509 (2021).

[0100] 10.Ding, J. et al. Alterations of RNA splicing patterns in esophagussquamous cell carcinoma. Cell Biosci 11, (2021).

[0101] 11.Ye, P., Yang, Y., Zhang, L.&Zheng, G. Prognostic Signatures ofAlternative Splicing Events in Esophageal Carcinoma Based on TCGA Splice-SeqData. Front Oncol 11, (2021).

[0102] 12.Sung, H. et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. 71,209–249 (2021).

[0103] 13.Lebrigand, K., Magnone, V., Barbry, P.&Waldmann, R. Highthroughput error corrected Nanopore single cell transcriptome sequencing. NatCommun 11, (2020).

[0104] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for obtaining tissue spatial transcript information, characterized in that: It includes the following steps: Long-read sequencing: first purify the full-length cDNA sequence of the tissue to be tested in the spatial transcriptome library, then amplify the purified full-length cDNA sequence using primers to generate materials for preparing a nanopore sequencing library, and prepare the nanopore sequencing library using a nanopore sequencing library kit; and perform PCR amplification on the nanopore sequencing library; Oxford Nanopore data processing: The PCR amplified cDNA was ligated to the adapter in the Oxford Nanopore kit to generate a chimeric cDNA; the internal template switching oligonucleotide (TSO) of the chimeric cDNA, the 3' adapter sequence connected by poly(T), and then the poly(A / T) tail and 3' adapter sequence of all reads were scanned; the scanned reads were aligned to the reference genome of the sample to be tested using bioinformatics software, and the spatial barcodes and UMIs were assigned to the nanopore reads using the strategy and software described for the single-cell library; the SAM records for each spatial point and gene were then grouped by UMI; the consensus sequence was calculated using SPOA based on the number of available reads for the UMI; the sequence between the end of the TSO and the base before the polyA sequence was used; the Consensus cDNA sequence was aligned to the sequence of the reference genome of the sample to be tested using bioinformatics software in a spliced ​​alignment manner; the SAM records matching known genes were analyzed to match the Gencode vM24 transcript isoforms.

2. The method for obtaining tissue spatial transcript information according to claim 1, characterized in that: The purification method comprises: firstly amplifying the full-length cDNA sequence of the tissue to be tested in the spatial transcriptome library using a primer labeled with biotin, and then oscillating and washing the biotinylated cDNA amplification product to obtain a purified full-length cDNA sequence, and the product is used for subsequent amplification to generate materials for preparing a nanopore sequencing library.

3. The method for obtaining tissue spatial transcript information according to claim 2, characterized in that: The number of amplification cycles for amplifying the full-length cDNA sequence of the tissue to be tested in the spatial transcriptome library using biotin-labeled primers is 3-6 cycles, and the amplification primer sequences are shown in SEQ ID NO.1-2.

4. The method for obtaining tissue spatial transcript information according to claim 2, characterized in that: The number of cycles for amplifying the purified full-length cDNA sequence is 8-10 cycles.

5. The method for obtaining tissue spatial transcript information according to claim 1, characterized in that: The Oxford Nanopore kit is selected from the LSK-109 or LSK-110 kit.

6. The method for obtaining tissue spatial transcript information according to claim 1, characterized in that: After assigning UMIs to nanopore reads, the UMI grouping process also includes removing low-quality mapped reads and potential chimeric reads.

7. The method for obtaining tissue spatial transcript information according to claim 6, characterized in that: The consensus sequence of each molecule was calculated based on the number of available reads of the UMI using the ComputeConsensus sicelore-2.0 method and procedure; for molecules with more than two reads, the sequence between the TSO end and the base before the polyA sequence was used by the SPOA method to calculate the consensus sequence.

8. The method for obtaining tissue spatial transcript information according to claim 7, characterized in that: The tissue to be tested is selected from solid tumor tissue or lymph node metastasis tissue; and the bioinformatics software is selected from minimap2 v2.

17.

9. A data processing method, characterized in that: It includes: The original gene expression matrix generated by Space Ranger is processed using the R software package to create a Seurat object, add cell-related information, and perform data processing; the cell-related information includes: data containing gene-level nanopore long reads; and tissue space transcript information obtained by the method for obtaining tissue space transcript information according to any one of claims 1 to 8.

10. The data processing method according to claim 9, characterized in that: The cell-related information also includes gene-level information from the short-read output.

11. The data processing method according to claim 9, characterized in that: The R software package is selected from at least one of Bioconductor version 4.0.2 and Seurat 24 version software package 3.9.

9.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for obtaining tissue space transcript information as described in any one of claims 1 to 8 are implemented.

13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for obtaining tissue spatial transcript information according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Novel third-generation sequencing method

    CN114540472A