Multicomponent analysis methods, systems, devices, and storage media for direct RNA sequencing

By acquiring multi-dimensional RNA sequencing data in a single experiment and combining it with multi-omics analysis methods, the problem of not being able to fully utilize direct RNA sequencing results was solved, achieving efficient acquisition and accurate analysis of multi-dimensional information.

CN114464254BActive Publication Date: 2026-02-13GUANGZHOU APPARENT BIOTECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111595287.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-24
Publication Date
2026-02-13
Estimated Expiration
2041-12-24

AI Technical Summary

Technical Problem

In existing technologies, direct RNA sequencing results fail to fully extract multi-dimensional information, resulting in incomplete analysis and high costs, and multiple experiments introduce errors and effects.

Method used

This study obtains full-length transcript sequence data, transcript quantification data, methylation modification data, and nascent mRNA data in a single experiment. Multi-omics analysis is then performed using random forest and deep learning models, combined with association analysis, to integrate multi-dimensional information.

Benefits of technology

This approach enables the acquisition of multi-dimensional information from direct RNA sequencing results, reducing experimental errors and batch effects, and improving sequencing accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114464254B_ABST
    Figure CN114464254B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a direct RNA sequencing multi-omics analysis method, system, device and storage medium, relating to the technical field of biological information processing. The method comprises: obtaining sequencing data of direct RNA sequencing; aligning the sequencing data with a reference genome to obtain sequencing alignment data; performing full-length transcript identification according to the sequencing alignment data to obtain full-length transcript sequence data; performing transcript quantification processing on the sequencing data based on the full-length transcript sequence data to obtain transcript quantification data; processing the sequencing data according to a methylation modification prediction model to obtain methylation modification data; processing the full-length transcript sequence data according to a nascent mRNA prediction model to obtain nascent mRNA data; and performing correlation analysis according to the full-length transcript sequence data, the transcript quantification data, the methylation modification data and the nascent mRNA data to obtain direct RNA sequencing multi-dimensional information. The method can achieve the technical effect of sequencing accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biological information processing, in particular, relates to a multi-omics analysis method, system, device and storage medium for direct RNA sequencing. BACKGROUND

[0002] At present, in the third generation sequencing technology, nanopore sequencing technology can directly sequence RNA molecules, also known as direct RNA sequencing technology; the sequencing results obtained by the direct RNA sequencing technology contain multiple dimensional information of the RNA molecules, but there is no method for mining the direct RNA sequencing results at present, so it is of great practical significance to realize the multi-dimensional information acquisition of the direct RNA sequencing results.

[0003] In the prior art, in the RNA molecule data analysis process of the third generation sequencing technology, the following technologies mainly exist: 1. Full-length transcript quantitative analysis combined with second generation sequencing technology, which generally uses the RNA sequence obtained by the second generation sequencing technology combined with the RNA sequence measured by the third generation sequencing technology for quantitative analysis, the disadvantage is that the sample needs to be subjected to second generation sequencing once in the quantitative analysis, which requires a large amount of sample and is expensive; 2. Full-length transcript identification analysis, which uses a qualitative method to identify only the full-length structure sequence of the RNA molecule in the third generation sequencing result, the disadvantage is that other dimensional data of the direct RNA sequencing cannot be effectively utilized; 3. Methylation modification identification of the RNA molecule, which uses a model to predict the methylation site of the RNA molecule, the disadvantage is that the prediction accuracy needs to be improved, and the methylation site cannot be determined in the original RNA molecule. There are few analysis schemes for direct RNA sequencing technology at home and abroad, only one or two dimensional information of the RNA molecule is analyzed, and the third generation sequencing technology is expensive, and the direct RNA sequencing results are not fully utilized. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide a multi-omics analysis method, system, device and storage medium for direct RNA sequencing, which can realize the technical effect of improving sequencing accuracy.

[0005] In a first aspect, the embodiments of the present application provide a multi-omics analysis method for direct RNA sequencing, comprising:

[0006] obtaining sequencing data of direct RNA sequencing;

[0007] aligning the sequencing data with a reference genome to obtain sequencing alignment data;

[0008] performing full-length transcript identification according to the sequencing alignment data to obtain full-length transcript sequence data;

[0009] transcript quantification data is obtained based on the full-length transcript sequence data;

[0010] methylation modification data is obtained by processing the sequencing data according to a methylation modification prediction model;

[0011] new mRNA data is obtained by processing the full-length transcript sequence data according to a new mRNA prediction model;

[0012] correlation analysis is performed according to the full-length transcript sequence data, the transcript quantification data, the methylation modification data and the new mRNA data, and direct RNA sequencing multi-dimensional information is obtained.

[0013] In the implementation process, the multi-omics analysis method of direct RNA sequencing processes the sequencing data and the sequencing alignment data of direct RNA sequencing to obtain full-length transcript sequence data, transcript quantification data, methylation modification data and new mRNA data in sequence, can perform multi-dimensional data analysis on the RNA sequence of direct RNA sequencing, four-dimensional information is obtained from one sequencing data analysis, and the technical effect of improving sequencing accuracy is realized by realizing multi-omics data integration. Therefore, the multi-omics analysis method of direct RNA sequencing can obtain multiple data schemes at the same time through one experiment, reduces experimental errors and batch effects caused by multiple biological experiments, and realizes the technical effect of improving sequencing accuracy.

[0014] Further, the step of identifying full-length transcripts according to the sequencing alignment data to obtain full-length transcript sequence data comprises:

[0015] The sequencing alignment data is corrected to obtain corrected sequencing alignment data;

[0016] The corrected sequencing alignment data is clustered to obtain the full-length transcript sequence data.

[0017] Further, the step of performing transcript quantification on the sequencing data based on the full-length transcript sequence data to obtain transcript quantification data comprises:

[0018] The full-length transcript sequence data is used as a reference sequence of quantitative transcripts, and the sequencing data is quantified based on the reference sequence to obtain the transcript quantification data.

[0019] Further, the step of obtaining methylation modification data by processing the sequencing data according to a methylation modification prediction model comprises:

[0020] Raw reads are obtained from the sequencing data;

[0021] recomputing the electrical signal of the original read length to obtain electrical signal change information caused by methylation modification;

[0022] identifying methylation modification according to the original read length and the electrical signal change information to obtain methylation modification site data;

[0023] filtering the methylation modification site data to obtain the methylation modification data.

[0024] Further, the step of processing the full-length transcript sequence data according to the nascent mRNA prediction model to obtain nascent mRNA data comprises:

[0025] processing the sequence of the direct sequencing read length of the RNA incubated with 5EU and the direct sequencing read length of the RNA not incubated with 5EU to obtain a plurality of feature information;

[0026] establishing a random forest model;

[0027] training the plurality of feature information according to the random forest model to obtain a training model;

[0028] inputting the sequencing data into the training model to obtain the nascent mRNA data.

[0029] Further, the direct RNA sequencing multi-dimensional information comprises first association information, second association information and third association information, and the step of obtaining direct RNA sequencing multi-dimensional information according to the full-length transcript sequence data, the transcript quantification data, the methylation modification data and the nascent mRNA data comprises:

[0030] associating the full-length transcript sequence data with the transcript quantification data to obtain the first association information;

[0031] associating the transcript quantification data with the methylation modification data to obtain the second association information;

[0032] associating the methylation modification data with the nascent mRNA data to obtain the third association information.

[0033] In a second aspect, the embodiments of the present application provide a direct RNA sequencing multi-omics analysis system, comprising:

[0034] an acquisition module configured to acquire sequencing data of direct RNA sequencing;

[0035] a sequencing alignment module configured to align the sequencing data with a reference genome to obtain sequencing alignment data;

[0036] The full-length transcript module is used to identify full-length transcripts based on the sequencing alignment data and obtain full-length transcript sequence data.

[0037] A quantification module is used to perform transcript quantification processing on the sequencing data based on the full-length transcript sequence data to obtain transcript quantification data.

[0038] A methylation modification module is used to process the sequencing data according to a methylation modification prediction model to obtain methylation modification data.

[0039] The newborn mRNA module is used to process the full-length transcript sequence data according to the newborn mRNA prediction model to obtain newborn mRNA data.

[0040] The association analysis module is used to perform association analysis based on the full-length transcript sequence data, the transcript quantification data, the methylation modification data, and the nascent mRNA data to obtain multi-dimensional information from direct RNA sequencing.

[0041] Furthermore, the sequencing alignment module includes:

[0042] A correction unit is used to correct the sequencing alignment data to obtain corrected sequencing alignment data;

[0043] The sequencing alignment unit is used to perform clustering processing on the corrected sequencing alignment data to obtain the full-length transcript sequence data.

[0044] Furthermore, the quantification module is also used to use the full-length transcript sequence data as a reference sequence for quantifying transcripts, and to perform transcript quantification on the sequencing data based on the reference sequence to obtain the transcript quantification data.

[0045] Furthermore, the methylation modification module includes:

[0046] The raw read unit is used to obtain the raw read length based on the sequencing data;

[0047] An electrical signal transformation unit is used to recalculate the electrical signal of the original read length to obtain information on the electrical signal changes caused by methylation modification;

[0048] The modification site unit is used to identify methylation modifications based on the original read length and the electrical signal change information, and to obtain methylation modification site data.

[0049] A methylation modification unit is used to filter the methylation modification site data to obtain the methylation modification data.

[0050] Furthermore, the newly generated mRNA module includes:

[0051] a feature information unit, configured to process sequences of the direct sequencing read length of the RNA subjected to the 5EU incubation and the direct sequencing read length of the RNA not subjected to the 5EU incubation, to obtain a plurality of feature information;

[0052] a model establishing unit, configured to establish a random forest model;

[0053] a training unit, configured to train the plurality of feature information according to the random forest model, to obtain a training model;

[0054] a nascent mRNA unit, configured to input the sequencing data into the training model, to obtain the nascent mRNA data.

[0055] Further, the correlation analysis module comprises:

[0056] a first correlation unit, configured to perform correlation analysis on the full-length transcript sequence data and the transcript quantification data, to obtain first correlation information;

[0057] a second correlation unit, configured to perform correlation analysis on the transcript quantification data and the methylation modification data, to obtain second correlation information;

[0058] a third correlation unit, configured to perform correlation analysis on the methylation modification data and the nascent mRNA data, to obtain third correlation information.

[0059] In a third aspect, an electronic device is provided, which comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the steps of the method according to any one of the first aspect when executing the computer program.

[0060] In a fourth aspect, a storage medium is provided, and the storage medium stores instructions, and the instructions, when running on a computer, cause the computer to execute the method according to any one of the first aspect.

[0061] In a fifth aspect, a computer program product is provided, and the computer program product, when running on a computer, causes the computer to execute the method according to any one of the first aspect.

[0062] Other features and advantages of the present application will be illustrated in the following description, or can be known or determined by no doubt from the description, or can be known by implementing the above-mentioned technology disclosed in the present application.

[0063] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the following preferred embodiments are specifically described with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0065] Figure 1 The flowchart of the direct RNA sequencing multi-omics analysis method provided by the embodiments of the present application is shown in the figure.

[0066] Figure 2 The flowchart of obtaining full-length transcript sequence data provided by the embodiments of the present application is shown in the figure.

[0067] Figure 3 The flowchart of obtaining methylation modification data provided by the embodiments of the present application is shown in the figure.

[0068] Figure 4 The flowchart of obtaining nascent mRNA data provided by the embodiments of the present application is shown in the figure.

[0069] Figure 5 The flowchart of obtaining direct RNA sequencing multi-dimensional information provided by the embodiments of the present application is shown in the figure.

[0070] Figure 6 The gene expression isoform graph provided by the embodiments of the present application is shown in the figure.

[0071] Figure 7 The RNA expression box plot provided by the embodiments of the present application is shown in the figure.

[0072] Figure 8 The differential RNA heat map provided by the embodiments of the present application is shown in the figure.

[0073] Figure 9 The differential RNA volcano plot provided by the embodiments of the present application is shown in the figure.

[0074] Figure 10 The distribution pie chart of m6A Peak on the RNA structure provided by the embodiments of the present application is shown in the figure.

[0075] Figure 11 The distribution metagene plot of m6A Peak on the RNA structure provided by the embodiments of the present application is shown in the figure.

[0076] Figure 12 The differential heat map of m6A site provided by the embodiments of the present application is shown in the figure.

[0077] Figure 13 The nascent mRNA expression box plot provided by the embodiments of the present application is shown in the figure.

[0078] Figure 14 a new mRNA differential heat map provided for the embodiments of the present application;

[0079] Figure 15 a gene level multifunctional analysis chart provided for the embodiments of the present application;

[0080] Figure 16 a differential gene distribution DIU pie chart provided for the embodiments of the present application;

[0081] Figure 17 a DIU analysis volcano chart provided for the embodiments of the present application;

[0082] Figure 18 an m6A and isodorm combined analysis four-quadrant chart provided for the embodiments of the present application;

[0083] Figure 19 a structure schematic diagram of a multi-omics analysis system of direct RNA sequencing provided for the embodiments of the present application;

[0084] Figure 20 a structure schematic diagram of a sequencing alignment module provided for the embodiments of the present application;

[0085] Figure 21 a structure schematic diagram of a methylation modification module provided for the embodiments of the present application;

[0086] Figure 22 a structure schematic diagram of a new mRNA module provided for the embodiments of the present application;

[0087] Figure 23 a structure schematic diagram of a correlation analysis module provided for the embodiments of the present application;

[0088] Figure 24 a structure block diagram of a device provided for the embodiments of the present application. DETAILED DESCRIPTION

[0089] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.

[0090] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. Meanwhile, in the description of the present application, the terms “first”, “second” and the like are only used for distinguishing description, and cannot be understood as indicating or implying relative importance.

[0091] The embodiment of the present application provides a multi-omics analysis method, system, device and storage medium of direct RNA sequencing, which can be applied to RNA sequencing analysis; the multi-omics analysis method of direct RNA sequencing obtains full-length transcript sequence data, transcript quantification data, methylation modification data and nascent mRNA data in sequence by processing sequencing data and sequencing alignment data of direct RNA sequencing, can perform multi-dimensional data analysis on RNA sequence of direct RNA sequencing, four-dimensional information is obtained from one sequencing data analysis, so that the data integration of multi-omics is realized; therefore, the multi-omics analysis method of direct RNA sequencing can obtain multiple data schemes at the same time through one experiment, reduces experimental errors and batch effects caused by multiple biological experiments, and realizes the technical effect of improving sequencing accuracy.

[0092] See Figure 1 , Figure 1 The flowchart of the multi-omics analysis method of direct RNA sequencing provided by the embodiment of the present application is shown, and the multi-omics analysis method of direct RNA sequencing comprises the following steps:

[0093] S100: obtaining sequencing data of direct RNA sequencing.

[0094] Exemplarily, RNA (Ribonucleic Acid) is a genetic information carrier existing in biological cells and part of viruses and viroids. RNA is a long-chain molecule condensed by ribonucleotides through phosphodiester bonds. One ribonucleotide molecule is composed of phosphate, ribose and base. The bases of RNA mainly include four kinds, namely A (adenine), G (guanine), C (cytosine) and U (uracil), wherein U (uracil) replaces T in DNA (DeoxyriboNucleic Acid). The main function of ribonucleic acid in vivo is to guide the synthesis of protein.

[0095] In some embodiments, the embodiment of the present application obtains the result of RNA direct sequencing, i.e. the sequencing data of direct RNA sequencing, by incubating a biological tissue sample or a cell sample with 5EU, extracting RNA and using a three-generation sequencer.

[0096] S200: aligning the sequencing data with a reference genome to obtain sequencing alignment data.

[0097] Exemplarily, a genome is the sum of the genetic material of all chromosomes of a specific organism, which is expressed by the total number of DNA base pairs. For a haploid, the genome represents the total DNA of the organism; for a diploid higher organism, the total DNA of its gametes is a set of genomes. Diploids have two homologous genomes, and eukaryotic cells have several chromosome sets, so they have several genomes. The genome of bacteria is a unique chromosome.

[0098] In some embodiments, the sequencing data of direct RNA sequencing is aligned with the reference genome using alignment tool minimap2 (Li Heng), and the aligned bam file result, i.e., the sequencing alignment data, is output.

[0099] S300: Full-length transcript identification is performed according to the sequencing alignment data, and full-length transcript sequence data is obtained.

[0100] Exemplarily, a transcript is one or more mature mRNAs that can encode proteins formed by transcription of a gene.

[0101] In some embodiments, full-length transcript identification is performed according to the aligned bam file result, and new full-length transcripts are identified and discovered, for example, the aligned bam file result is subjected to full-length transcript identification using FLAIR (van Baren, M.J) software, thereby obtaining full-length transcript sequence data.

[0102] S400: Transcript quantification is performed on the sequencing data based on the full-length transcript sequence data, and transcript quantification data is obtained.

[0103] In some embodiments, during transcript quantification, the full-length transcript sequence data can be used as a reference sequence for quantifying transcripts, and each sample subjected to direct RNA sequencing is subjected to sequencing to obtain sequencing results, and then accurate transcript quantification, i.e., transcript quantification data, is obtained by using salmon (Rob Patro) quantification software.

[0104] S500: Methylation modification data is obtained by processing the sequencing data according to a methylation modification prediction model.

[0105] Exemplarily, methylation refers to the process of catalytically transferring a methyl group from an active methyl compound (such as S-adenosyl methionine) to other compounds. Various methyl compounds can be formed, or methylated products can be formed by chemical modification of certain proteins or nucleic acids. In biological systems, methylation is catalyzed by enzymes, and such methylation is involved in heavy metal modification, regulation of gene expression, regulation of protein function, and RNA processing. In eukaryotes, 5' Cap and 3' ploy A modification play a very important role in transcriptional regulation, and internal modification of mRNA is used to maintain the stability of mRNA. The most common internal modification of mRNA includes N6-adenosine methylation (m6A), N1-adenosine methylation (m1A), cytosine hydroxylation (m5C), etc.

[0106] Exemplarily, when methylation occurs at the sixth nitrogen atom of the adenosine of the RNA molecule, it is referred to as m6A; it should be understood that the m6A methylation modification identification in the embodiments of the present application is taken as a general example.

[0107] In some embodiments, the m6A methylation modification identification is to predict the methylation modification information in the RNA molecule from the sequencing data of direct RNA sequencing by software tombo (Stoiber, M.H) and MINES (Gene W. Yeo).

[0108] S600: processing the full-length transcript sequence data according to the nascent mRNA prediction model to obtain nascent mRNA data.

[0109] In some embodiments, the nascent mRNA identification is to identify the nascent mRNA from the full-length transcript sequence data by a prediction model established by deep learning.

[0110] S700: performing correlation analysis according to the full-length transcript sequence data, the transcript quantification data, the methylation modification data and the nascent mRNA data to obtain direct RNA sequencing multidimensional information.

[0111] In some embodiments, the multi-omics fusion analysis is to perform correlation analysis on the RNA molecule multidimensional information (full-length transcript sequence data, transcript quantification data, methylation modification data and nascent mRNA data) obtained in steps S300, S400, S500 and S600 to obtain direct RNA sequencing multidimensional information.

[0112] In some implementation scenarios, the biological tissue sample or the cell sample is incubated with 5EU, RNA is extracted, and the result of direct RNA sequencing is obtained by a third-generation sequencer. The present application designs a multi-omics analysis method based on the third-generation sequencing technology-direct RNA sequencing for the direct RNA sequencing data containing RNA molecule multidimensional information, which overcomes the defects of the prior art that can only analyze 1 or 2 dimensional information of the direct RNA sequencing result and cannot completely mine the multidimensional information of the direct RNA sequencing result, and at the same time reduces the experimental error and batch effect caused by multiple experiments of the sample. The present application analyzes multiple sets of results through one sequencing, realizes the organic combination of the data, and one-time realizes the multi-omics analysis content. Therefore, the multi-omics analysis method of direct RNA sequencing can simultaneously obtain multiple sets of data schemes through one experiment, reduces the experimental error and batch effect caused by multiple biological experiments, and realizes the technical effect of improving the sequencing accuracy.

[0113] Please refer to Figure 2 , Figure 2 The flowchart for obtaining full-length transcript sequence data provided by the embodiments of the present application includes the following steps:

[0114] S301: correcting the sequencing alignment data to obtain corrected sequencing alignment data;

[0115] S302: clustering the corrected sequencing alignment data to obtain full-length transcript sequence data.

[0116] In some embodiments, full-length transcript identification is performed, and based on the sequencing alignment data, full-length transcript identification and discovery of new full-length transcripts can be performed, for example, using the FLAIR (van Baren, M.J) software to perform full-length transcript identification on the alignment bam file:

[0117] Step 1: correcting sequencing error bases in the sequencing alignment data to obtain corrected sequencing alignment data.

[0118] Step 2: clustering the corrected sequencing alignment data to obtain transcript isoform sequences, i.e., identifying full-length transcript sequence fasta files (full-length transcript sequence data).

[0119] In some embodiments, in the step S400 of performing transcript quantification on the sequencing data based on the full-length transcript sequence data to obtain transcript quantification data, the following steps are included:

[0120] The full-length transcript sequence data is used as a reference sequence for quantifying the sequencing data to obtain transcript quantification data.

[0121] By way of example, for transcript quantification, the full-length transcript sequence data is used as a reference sequence for quantifying the sequencing data to obtain accurate transcript quantification:

[0122] Step 1: using the full-length transcript sequence data as a reference sequence for quantifying the sequencing data;

[0123] Step 2: using the salmon software to quantify the sequencing data to obtain reads count for each full-length transcript, and further obtain accurate transcript quantification, i.e., transcript quantification data.

[0124] By way of example, due to the limitation of current sequencing level, the genome needs to be broken into DNA fragments before library construction and sequencing. Reads refer to the base sequence obtained by a single sequencing of a sequencer, i.e., a series of ATCGGGTA, which is not a component of the genome. Different sequencing instruments have different read lengths. Sequencing the entire genome will produce hundreds of millions of reads.

[0125] Referring to Figure 3 , Figure 3 A flowchart for obtaining methylation modification data provided by the embodiment of the present application includes the following steps:

[0126] S501: Obtain raw reads according to sequencing data;

[0127] S502: Recalculate the electrical signal of the raw reads to obtain the electrical signal change information caused by methylation modification;

[0128] S503: Perform methylation modification identification according to the raw reads and the electrical signal change information to obtain methylation modification site data;

[0129] S504: Filter the methylation modification site data to obtain methylation modification data.

[0130] Exemplarily, taking m6A methylation modification identification as an example, the sequencing data of direct RNA sequencing is used to predict the methylation modification information in the RNA molecule through the software tombo (Stoiber, M.H) and MINES (Gene W. Yeo):

[0131] First step: through the tombo (Stoiber, M.H) software, the raw reads of direct RNA sequencing data are written into the fast5 file of the storage data current signal corresponding to the original data;

[0132] Second step: through the tombo (Stoiber, M.H) software, the electrical signal of the reads is recalculated, the electrical signal change caused by m6A modification is found out, and the change is written into the fast5 file again;

[0133] Third step: through the tombo (Stoiber, M.H) software, the fast5 files processed in the first step and the second step are used to identify m6A modification, and a wig file (methylation modification site data) for identifying m6A modification is outputted;

[0134] Fourth step: the wig file obtained in the above step is filtered through the MINES (Gene W. Yeo) software to obtain the filtered m6A modification site file, that is, the methylation modification information (methylation modification data) in the RNA molecule.

[0135] Referring to Figure 4 , Figure 4 A flowchart for obtaining new mRNA data provided by the embodiment of the present application includes the following steps:

[0136] S601: Process the sequence of the direct sequencing read length of the RNA incubated with 5EU and the direct sequencing read length of the RNA not incubated with 5EU to obtain a plurality of characteristic information;

[0137] S602: Train the plurality of characteristic information according to the random forest model to obtain a training model;

[0138] S603: Input the sequencing data into the training model to obtain new mRNA data.

[0139] Exemplarily, when identifying the new mRNA, the new mRNA can be identified according to the full-length transcript sequence data and the prediction model established by deep learning, and the new mRNA data is obtained:

[0140] First step: First, the pre-model training data processing is performed, the direct sequencing reads of the RNA incubated with 5EU (experimental group) and the direct sequencing reads of the RNA not incubated with 5EU (control group) are processed to obtain 9 characteristic information (the characteristic of each read is: U base mismatch rate, U base mismatch sequencing quality average value, U measured as A base mismatch rate, U measured as A base mismatch sequencing quality average value, U measured as C base mismatch rate, U measured as C base mismatch sequencing quality average value, U measured as G base mismatch rate, U measured as G base mismatch sequencing quality average value, and base deletion rate);

[0141] Second step: Establish a random forest model, train the 9 characteristic information obtained above to predict new mRNA, and obtain a training model;

[0142] Third step: Input the sequencing data of direct RNA sequencing into the training model to predict new mRNA, and obtain the read sequence of new mRNA, that is, new mRNA data.

[0143] See Figure 5 , Figure 5 The flowchart for obtaining the multi-dimensional information of direct RNA sequencing provided by the embodiments of the present application includes the following steps:

[0144] S701: Perform correlation analysis on the full-length transcript sequence data and the transcript quantification data to obtain first correlation information;

[0145] S702: Perform correlation analysis on the transcript quantification data and the methylation modification data to obtain second correlation information;

[0146] S703: Perform correlation analysis on the methylation modification data and the new mRNA data to obtain third correlation information.

[0147] Exemplarily, when performing multi-omics fusion analysis, the multi-dimensional information of the RNA molecules obtained in each step S300-S600 (full-length transcript sequence data, transcript quantification data, methylation modification data, and nascent mRNA data) is subjected to correlation analysis to obtain direct RNA sequencing multi-dimensional information:

[0148] Step 1: Correlation analysis of full-length transcript sequence data and transcript quantification data, isoform matrix is obtained by transcript quantification, and isoform correlation can be used to view the expression of differential isoforms on genes;

[0149] Step 2: Correlation analysis of transcript quantification data and methylation modification data, isoform matrix and methylation modification data can be correlated to find isoforms affected by m6A modification;

[0150] Step 3: Correlation analysis of methylation modification data and nascent mRNA data, correlation of methylation modification data and nascent mRNA data can find the reason for the change of nascent mRNA.

[0151] Exemplarily, isoform refers to different versions of proteins of the same gene.

[0152] In some embodiments, the present application realizes a direct RNA sequencing data 4-dimensional multi-omics analysis method (as shown in Figures 1 to 5 ), which breaks the method of analyzing only 1-2 dimensional data of direct RNA sequencing data. Among them, the transcript quantification establishes the original data filtering standard, the m6A methylation modification identification constructs the deep learning prediction model, the AUC reaches 88%, which is higher than the prediction accuracy of the prior art, the nascent mRNA identification simultaneously constructs the deep learning model, the AUC reaches 90%, which is higher than the prediction standard of the prior art. In addition, the present application realizes a scheme of obtaining multiple sets of data at one time, reduces the experimental error and batch effect caused by multiple biological experiments, and the multi-omics fusion analysis module provides cross fusion of 4-dimensional data, which brings a new direction of biological research.

[0153] In some embodiments, the method embodiment according to Figures 1 to 5 is applied in actual application, and the obtained results are as follows:

[0154] (1) Dimension 1 full-length transcript result:

[0155]

[0156] Table 1-fastq data statistics table

[0157] As shown in Table 1, Sample Name is sample name; ReadNum is sequence number; Base (M) is total base number; N50 is N50 length; MeanLength is average length of reads; MaxLength is the length of the longest reads; MeanQscore is average quality value.

[0158] See Figure 6 , Figure 6 The gene expression isoform graph provided in the embodiment of the present application, Figure 6 The black block in the figure represents the exon region of the transcript; the transcript without ENST number is a newly identified transcript.

[0159] (2) Dimension 2 full-length transcript results:

[0160]

[0161] Table 2 - Comparison rate statistics table

[0162] As shown in Table 2, Sample is sample name; AllRead is total sequencing reads number; MapRead is the number of reads aligned to the reference sequence; map is the alignment rate.

[0163] See Figure 7 , Figure 7 The RNA expression box plot provided in the embodiment of the present application, the horizontal coordinate is the sample name, and the vertical coordinate is the gene count value.

[0164] See Figure 8 , Figure 8 The differential RNA heat map provided in the embodiment of the present application, the horizontal coordinate is the sample name, and the vertical coordinate is the expression amount of the differential transcript. The darker the color, the higher the expression amount.

[0165] See Figure 9 , Figure 9 The differential RNA volcano plot provided in the embodiment of the present application, the horizontal coordinate is the fold difference of transcript expression amount under two conditions, and the vertical coordinate is the significance FDR value. Wherein, FDR < 0.05 is the significant difference result, and the dark point in the figure is the significant difference result.

[0166] (3) RNA methylation modification module results:

[0167] See Figure 10 , Figure 10 The pie chart of m6APeak distribution on the RNA structure provided in the embodiment of the present application; as Figure 10The pie chart shows the methylation modification position statistics of the 3' UTR, 5' UTR, CDS, and ncRNA of the gene.

[0168] Please refer to Figure 11 , Figure 11 The metagene plot of the distribution of m6APeak on the RNA structure provided in the embodiments of the present application, Figure 11 The horizontal coordinate is the region position of the gene, and the vertical coordinate is the proportion of methylation modification.

[0169] Please refer to Figure 12 , Figure 12 The differential heat map of m6A sites provided in the embodiments of the present application, Figure 12 The horizontal coordinate is the sample name, and the vertical coordinate is the expression amount of the differential transcript. The darker the color, the higher the expression amount.

[0170] (4) New mRNA results:

[0171] Please refer to Figure 13 , Figure 13 The box plot of the expression of new mRNA provided in the embodiments of the present application, Figure 13 The horizontal coordinate is the sample name, and the vertical coordinate is the gene count value.

[0172] Please refer to Figure 14 , Figure 14 The differential heat map of new mRNA provided in the embodiments of the present application, Figure 14 The horizontal coordinate is the sample name, and the vertical coordinate is the expression amount of the differential transcript. The darker the color, the higher the expression amount.

[0173] (5) Joint analysis module results based on (1) to (4):

[0174] Please refer to Figure 15 , Figure 15 The multifunctional analysis chart of the gene level provided in the embodiments of the present application, Figure 15 The horizontal coordinate is the gene proportion, and the vertical coordinate is the number of gene region statistical index statistics.

[0175] Please refer to Figure 16 , Figure 16 The DIU pie chart of the distribution of differential genes provided in the embodiments of the present application, wherein DIU refers to significant differential genes, and Major isoform Switching represents structural changes in transcripts.

[0176] Please refer to Figure 17 , Figure 17The DIU analysis volcano plot provided in this application embodiment shows that the horizontal axis represents the fold difference in transcript expression levels under two conditions, and the vertical axis represents the significance FDR value, where FDR < 0.05 indicates a significant difference. Dark dots in the figure represent significant differences.

[0177] Please see Figure 18 , Figure 18 The four-quadrant diagram for the combined analysis of m6A and isodorm provided in the embodiments of this application shows that the horizontal axis represents the fold difference in transcript expression levels under two conditions, and the vertical axis represents the fold difference in methylation modification expression levels under two conditions.

[0178] Please see Figure 19 , Figure 19 This is a schematic diagram of the structure of a multi-omics analysis system for direct RNA sequencing provided in an embodiment of this application. The multi-omics analysis system for direct RNA sequencing includes:

[0179] The acquisition module 100 is used to acquire sequencing data from direct RNA sequencing.

[0180] Sequencing alignment module 200 is used to align sequencing data with a reference genome to obtain sequencing alignment data;

[0181] The full-length transcript module 300 is used to identify full-length transcripts based on sequencing alignment data and obtain full-length transcript sequence data.

[0182] The quantification module 400 is used to perform transcript quantification processing on sequencing data based on full-length transcript sequence data to obtain transcript quantification data.

[0183] The methylation modification module 500 is used to process sequencing data according to the methylation modification prediction model to obtain methylation modification data;

[0184] The newly generated mRNA module 600 is used to process full-length transcript sequence data according to the newly generated mRNA prediction model to obtain newly generated mRNA data.

[0185] The association analysis module 700 is used to perform association analysis based on full-length transcript sequence data, transcript quantification data, methylation modification data, and nascent mRNA data to obtain multi-dimensional information from direct RNA sequencing.

[0186] Please see Figure 20 , Figure 20 This is a schematic diagram of the sequencing alignment module provided in an embodiment of this application.

[0187] For example, the sequencing alignment module 200 includes:

[0188] The correction unit 301 is used to correct the sequencing alignment data to obtain corrected sequencing alignment data;

[0189] The sequencing alignment unit 302 is configured to cluster the corrected sequencing alignment data to obtain full-length transcript sequence data.

[0190] The quantification module 400 is configured to take the full-length transcript sequence data as a reference sequence of a quantified transcript, perform transcript quantification on the sequencing data based on the reference sequence, and obtain transcript quantification data.

[0191] Referring to Figure 21 , Figure 21 A structural diagram of the methylation modification module provided in the embodiments of the present application is shown.

[0192] The methylation modification module 500 includes:

[0193] The raw read length unit 501 is configured to obtain raw read length from the sequencing data.

[0194] The electrical signal change unit 502 is configured to re-calculate the electrical signal of the raw read length to obtain electrical signal change information caused by methylation modification.

[0195] The modification site unit 503 is configured to perform methylation modification identification based on the raw read length and the electrical signal change information to obtain methylation modification site data.

[0196] The methylation modification unit 504 is configured to filter the methylation modification site data to obtain methylation modification data.

[0197] Referring to Figure 22 , Figure 22 A structural diagram of the nascent mRNA module provided in the embodiments of the present application is shown.

[0198] The nascent mRNA module 600 includes:

[0199] The feature information unit 601 is configured to process the sequences of the 5EU incubation-added RNA direct sequencing read length and the 5EU incubation-unadded RNA direct sequencing read length to obtain a plurality of feature information.

[0200] The model establishment unit 602 is configured to establish a random forest model.

[0201] The training unit 603 is configured to train the plurality of feature information based on the random forest model to obtain a training model.

[0202] The nascent mRNA unit 604 is configured to input the sequencing data into the training model to obtain nascent mRNA data.

[0203] Referring to Figure 23 , Figure 23A structural diagram of the correlation analysis module provided in the embodiments of the present application is shown.

[0204] The correlation analysis module 700 includes, for example:

[0205] The first correlation unit 701 is configured to perform correlation analysis on the full-length transcript sequence data and the transcript quantitative data to obtain first correlation information.

[0206] The second correlation unit 702 is configured to perform correlation analysis on the transcript quantitative data and the methylation modification data to obtain second correlation information.

[0207] The third correlation unit 703 is configured to perform correlation analysis on the methylation modification data and the nascent mRNA data to obtain third correlation information.

[0208] It should be understood that, Figures 19 to 23 The direct RNA sequencing multi-omics analysis system shown in the embodiments of the present application corresponds to the method embodiments shown in the embodiments of the present application, and thus will not be described herein again to avoid repetition. Figures 1 to 5 The method embodiments shown in the embodiments of the present application correspond to each other, and thus will not be described herein again to avoid repetition.

[0209] The present application also provides an electronic device, which is shown in Figure 24 , Figure 24 A structural diagram of a device provided in the embodiments of the present application is shown. The electronic device can include a processor 510, a communication interface 520, a memory 530, and at least one communication bus 540. The communication bus 540 is configured to realize direct connection and communication among the components. The communication interface 520 of the electronic device in the embodiments of the present application is configured to perform signaling or data communication with other node devices. The processor 510 can be an integrated circuit chip having a signal processing capability.

[0210] The processor 510 described above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. The processor 510 can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a ready programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The processor 510 can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor 510 can also be any conventional processor.

[0211] The memory 530 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory 530 stores computer readable instructions which, when executed by the processor 510, enable the electronic device to perform the above-mentioned Figures 1 to 5 The method embodiments relate to each step.

[0212] Optionally, the electronic device can further include a storage controller, an input / output unit.

[0213] The memory 530, the storage controller, the processor 510, the peripheral interface, and the input / output unit are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses 540. The processor 510 is configured to execute executable modules stored in the memory 530, such as software function modules or computer programs included in the electronic device.

[0214] The input / output unit is configured to provide a user with a creation task and create an optional time period or a preset execution time for the task to achieve user interaction with a server. The input / output unit can be, but is not limited to, a mouse, a keyboard, etc.

[0215] It can be understood that Figure 24 The structure shown is only schematic, and the electronic device can include more or fewer components than those shown in the figures, or have a different configuration from that shown in the figures. Figure 24 The components shown in the figures can be implemented in hardware, software, or a combination thereof. Figure 24 The components shown in the figures can be implemented in hardware, software, or a combination thereof. Figure 24 The components shown in the figures can be implemented in hardware, software, or a combination thereof.

[0216] The embodiments of the present application also provide a storage medium, which stores instructions. When the instructions are run on a computer, the computer program is executed by a processor to implement the method of the method embodiments. To avoid repetition, this will not be repeated here.

[0217] The present application also provides a computer program product, which, when run on a computer, causes the computer to execute the method of the method embodiments.

[0218] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can also be implemented by other means. The apparatus embodiments described above are only illustrative, for example, the flowcharts and block diagrams in the drawings show the possible implementation architecture, function and operation of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different order from that shown in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0219] In addition, the functional modules in the embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0220] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0221] The above merely provides an example of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, and thus, once an item is defined in one drawing, it need not be further defined and explained in subsequent drawings.

[0222] The above merely provides an example of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, and thus, once an item is defined in one drawing, it need not be further defined and explained in subsequent drawings.

[0223] It should be noted that the relational terms herein such as first and second and the like are used solely to distinguish one from another entity or action without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

Claims

1. A method of multi-omic analysis of direct RNA sequencing, characterized in that, The method comprises the following steps: obtaining sequencing data of direct RNA sequencing; aligning the sequencing data with a reference genome to obtain sequencing alignment data; performing full-length transcript identification according to the sequencing alignment data to obtain full-length transcript sequence data; performing transcript quantification processing on the sequencing data based on the full-length transcript sequence data to obtain transcript quantification data; processing the sequencing data according to a methylation modification prediction model to obtain methylation modification data; processing the full-length transcript sequence data according to a nascent mRNA prediction model to obtain nascent mRNA data; performing correlation analysis according to the full-length transcript sequence data, the transcript quantification data, the methylation modification data and the nascent mRNA data to obtain direct RNA sequencing multi-dimensional information. The step of performing full-length transcript identification according to the sequencing alignment data to obtain full-length transcript sequence data comprises the following steps: correcting the sequencing alignment data to obtain corrected sequencing alignment data; performing clustering processing on the corrected sequencing alignment data to obtain the full-length transcript sequence data.

2. The method of multi-omic analysis of direct RNA sequencing according to claim 1, wherein, The step of performing transcript quantification processing on the sequencing data based on the full-length transcript sequence data to obtain transcript quantification data comprises the following step: taking the full-length transcript sequence data as a reference sequence of a quantitative transcript, and performing transcript quantification on the sequencing data based on the reference sequence to obtain the transcript quantification data.

3. The method of multi-omic analysis of direct RNA sequencing of claim 1, wherein, The step of processing the sequencing data according to a methylation modification prediction model to obtain methylation modification data comprises the following steps: obtaining original read lengths from the sequencing data; recomputing the electrical signals of the original read lengths to obtain electrical signal change information caused by methylation modification; performing methylation modification identification according to the original read lengths and the electrical signal change information to obtain methylation modification site data; filtering the methylation modification site data to obtain the methylation modification data.

4. The method of multi-omic analysis of direct RNA sequencing of claim 1, wherein, The step of processing the full-length transcript sequence data according to a nascent mRNA prediction model to obtain nascent mRNA data comprises the following steps: processing the sequences of RNA direct sequencing read lengths with 5EU incubation and RNA direct sequencing read lengths without 5EU incubation to obtain a plurality of feature information; establishing a random forest model; training the plurality of feature information according to the random forest model to obtain a training model; inputting the sequencing data into the training model to obtain the nascent mRNA data.

5. The method of polyomic analysis of direct RNA sequencing according to claim 1, characterized in that, The direct RNA sequencing multi-dimensional information comprises first correlation information, second correlation information and third correlation information, and the step of performing correlation analysis according to the full-length transcript sequence data, the transcript quantification data, the methylation modification data and the nascent mRNA data to obtain direct RNA sequencing multi-dimensional information comprises the following steps: performing correlation analysis on the full-length transcript sequence data and the transcript quantification data to obtain the first correlation information; performing correlation analysis on the transcript quantification data and the methylation modification data to obtain the second correlation information; and performing correlation analysis on the transcript quantification data, the methylation modification data and the nascent mRNA data to obtain the third correlation information. The methylation modification data and the nascent mRNA data are associated to obtain the third association information.

6. A multi-omic analysis system for direct RNA sequencing, characterized by, The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of:

7. An apparatus, comprising: The method comprises the steps of: The method comprises the steps of:

8. A storage medium, characterized by The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The method comprises the steps of: The

Citation Information

Patent Citations

  • Proton-based transcriptome sequencing data comparison and analysis method and system

    CN104657628A

  • Prediction method for protein post-translational modification methylation loci

    CN105893787A

  • Transcript classification method

    CN111916147A

  • Nanopore sequencing-based m6A analysis method

    CN113130006A