The present invention relates to the technical field of
bioinformatics, and particularly to a method for identifying transposable element-derived RNAs based on
RNA sequencing data. The present invention proposes to take the
exon level of teRNAs as the analysis object. Compared with the teRNA locus, the teRNA
exon not only achieves more accurate locus identification, but also retains the form of the TE in the transcript, including whether it provides splicing sites and
polyadenylation signals, and whether it is chimeric with genes or other TE loci. And compared with the teRNA
transcript level, as a component of the transcript, the teRNA
exon significantly reduces the
sequence assembly difficulty of short-read
sequencing data. The present invention combines Bayesian models, reference-based
assembly algorithms, and local de novo
assembly algorithms, and through the mutual correction of different
assembly algorithms and repeated alignment of sequencing reads before and after assembly, reduces the false positives of assembly, thereby improving the accuracy of teRNA identification.