Lncrna recognition method based on hierarchical alignment gate fusion network

CN122598786APending Publication Date: 2026-08-18JILIN UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611095638.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0009]基于手工特征的机器学习方法虽然具备一定的生物学可解释性,但其表征能力有限,难以充分捕获序列中隐含的长程编码规律

Benefits of technology

[0012] Compared with existing technologies, the beneficial effects of this invention are: by constructing a semantic hierarchical alignment and adaptive fusion mechanism of multi-scale handmade features and deep learning hidden representations, it achieves accurate discrimination of RNA sequence coding potential, thereby improving recognition accuracy and cross-species generalization ability while maintaining interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598786A_ABST
    Figure CN122598786A_ABST
Patent Text Reader

Abstract

This invention relates to the field of genome sequencing technology, specifically disclosing a lncRNA identification method based on a hierarchical alignment-gated fusion network. The method includes extracting multi-scale hierarchical handcrafted features, which are divided into shallow, mid-level, and deep features according to semantic abstraction levels; inputting the RNA sequence into a deep learning coding model, extracting hidden representations output from multiple hidden layers of different depths in the deep learning coding model; gating and fusing handcrafted features and hidden representations belonging to the same semantic level to obtain hierarchical fusion representations for each semantic level; weighted aggregation of the hierarchical fusion representations for each semantic level to obtain a global sequence representation; and identifying lncRNAs in the RNA sequence based on the global sequence representation. This invention achieves accurate discrimination of the coding potential of RNA sequences by constructing a semantic hierarchical alignment and adaptive fusion mechanism between multi-scale handcrafted features and deep learning hidden representations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of genome sequencing technology, specifically a lncRNA recognition method based on a hierarchical alignment-gated fusion network. Background Technology

[0002] Since the initial sequencing of the human genome, researchers have successively completed genome sequencing of various organisms. In-depth research on genome sequences has revealed an important phenomenon: only a very small percentage of the genome, approximately 1.5%-2%, encodes proteins, while over 90% of the genome sequence can be transcribed into RNA. In early biological research, influenced by the central dogma, these non-protein-coding sequences were generally considered "transcriptional noise" or "junk genes" and were long ignored by researchers. However, with the implementation of the ENCODE project and the publication of its analysis results, approximately 80% of the genome was confirmed to participate in important chemical reactions affecting biological activities. This discovery has completely changed our understanding of non-coding sequences, demonstrating that a large number of non-coding RNAs (ncRNAs) play a crucial role in life activities.

[0003] Based on transcript length, non-coding RNAs can be divided into short non-coding RNAs (sncRNAs) and long non-coding RNAs (lncRNAs). lncRNAs are defined as transcripts longer than 200 nucleotides that do not have protein-coding capabilities. Similar to messenger RNA (mRNA), most lncRNAs are transcribed by RNA polymerase II and possess post-transcriptional modification features such as a 5' cap, a 3' polyadenylated tail, and alternative splicing. However, compared to mRNAs, lncRNAs exhibit poor interspecies conservation, generally lower expression levels, and strong tissue and condition specificity, making accurate identification and functional studies challenging. Based on their position in the genome, lncRNAs can be further classified into intronic lncRNAs, intergenic lncRNAs, and antisense lncRNAs, originating from intronic regions of coding genes, regions between two coding genes, and the antisense strand of coding genes, respectively.

[0004] With the rapid development of high-throughput sequencing technology, more and more lncRNAs are being discovered, and their important functions in various biological processes are gradually being revealed. lncRNAs can interact with DNA, RNA, and protein molecules, finely regulating gene expression at the epigenetic, transcriptional, and post-transcriptional levels. Specifically, lncRNAs can alter chromatin structure by recruiting chromatin remodeling complexes, achieving epigenetic silencing; they can interfere with the assembly and function of transcription machinery, affecting gene transcription initiation and elongation; and they can regulate the stability, translation efficiency, and splicing processes of mRNA by binding to it. These regulatory mechanisms enable lncRNAs to participate extensively in various life activities such as chromosomal dosage compensation, genomic imprinting, cell cycle regulation, cell differentiation, individual development, and stress responses.

[0005] In the medical field, aberrant expression of lncRNAs is closely related to the development and progression of various human diseases, including various cancers, neurodegenerative diseases (such as Alzheimer's disease and Huntington's disease), cardiovascular diseases, and immune diseases. Particularly in cancer research, lncRNAs have become a new research hotspot. Studies have shown that aberrantly expressed lncRNAs can affect the proliferation, survival, invasion, metastasis, and drug resistance of cancer cells. For example, the lncRNA HOTAIR can promote the malignant phenotype of lung cancer cells, while the prostate cancer-specific lncRNA PCA3 is highly expressed in the urine of prostate cancer patients and has become a clinically effective biomarker for prostate cancer diagnosis. These findings make lncRNAs promising early diagnostic markers and potential therapeutic targets for various diseases, providing new directions for precision medicine.

[0006] In the field of plant biology, research on lncRNAs started relatively late, and early studies focused primarily on short non-coding RNAs. In recent years, with in-depth research, increasing evidence suggests that lncRNAs play an indispensable role in plant growth, development, and stress responses. Plant lncRNAs participate in regulating flowering time, gene silencing, root organ growth, seedling morphogenesis, and reproductive processes, and exhibit significant expression differences across different tissues, developmental stages, and under various stress conditions. These characteristics provide new insights and directions for plant genomics research and crop genetics and breeding.

[0007] Despite the widespread recognition of the importance of lncRNAs, accurately identifying them from massive transcriptome datasets remains a challenging task. Traditional biological experimental methods are not only time-consuming, labor-intensive, and costly, but also struggle to detect lncRNAs with extremely low expression levels. Furthermore, the massive amounts of data generated by high-throughput sequencing technologies bring problems such as insufficient sequencing depth, sequencing bias, and sequencing errors, further increasing the difficulty of lncRNA identification. Therefore, developing efficient, accurate, and robust lncRNA identification methods has become an urgent need in the field of non-coding RNA research.

[0008] In recent years, the rapid development of machine learning technology has provided a powerful tool for solving this problem. Machine learning algorithms can automatically learn the differential characteristics between lncRNA and mRNA from large amounts of sequence data and build classification models to predict whether a new transcript is a lncRNA.

[0009] While machine learning methods based on handcrafted features possess a degree of biological interpretability, their representational capabilities are limited, making it difficult to fully capture the long-range coding patterns implicit in sequences. Deep learning methods can automatically extract deep semantic representations; however, the abstract statistical features learned from data are insufficient to capture the essential biological laws governing RNA coding potential. They also fail to effectively incorporate structured prior knowledge such as sequence composition and open reading frame structure, resulting in insufficient generalization ability in cross-species scenarios. Summary of the Invention

[0010] The purpose of this invention is to provide a lncRNA recognition method based on a hierarchical alignment gating fusion network to solve the problems mentioned in the background art.

[0011] To achieve the above objectives, the present invention provides the following technical solution: A method for lncRNA recognition based on a hierarchical alignment-gated fusion network, the method comprising: An RNA sequence is obtained, and multi-scale hierarchical handcrafted features are extracted from the RNA sequence. The multi-scale hierarchical handcrafted features are divided into shallow features, medium features, and deep features according to the semantic abstraction level. The RNA sequence is input into a deep learning coding model, and hidden representations output by multiple hidden layers of different depths in the deep learning coding model are extracted. The hidden representations include shallow representations, medium representations, and deep representations corresponding to the shallow features, medium features, and deep features, respectively. The handcrafted features and hidden representations belonging to the same semantic level are gated and fused. The handcrafted features and hidden representations are adaptively fused by gating weights of each dimension to obtain hierarchical fusion representations of each semantic level. We perform weighted aggregation on the hierarchical fusion representations of each semantic level to obtain a global sequence representation; The lncRNA in the RNA sequence is identified based on the global sequence characterization.

[0012] Compared with existing technologies, the beneficial effects of this invention are: by constructing a semantic hierarchical alignment and adaptive fusion mechanism of multi-scale handmade features and deep learning hidden representations, it achieves accurate discrimination of RNA sequence coding potential, thereby improving recognition accuracy and cross-species generalization ability while maintaining interpretability. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention.

[0014] Figure 1 This is an architecture diagram of a lncRNA recognition method based on a hierarchical alignment gating fusion network provided in an embodiment of the present invention.

[0015] Figure 2 A comparison of ROC curves of different methods provided in the embodiments of the present invention on a human benchmark set. Detailed Implementation

[0016] To make the technical problems to be solved, the technical solutions, and the beneficial effects of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.

[0017] In this embodiment of the invention, a method for lncRNA recognition based on a hierarchical alignment-gated fusion network is provided, the method comprising: An RNA sequence is obtained, and multi-scale hierarchical handcrafted features are extracted from the RNA sequence. The multi-scale hierarchical handcrafted features are divided into shallow features, medium features, and deep features according to the semantic abstraction level. The RNA sequence is input into a deep learning coding model, and hidden representations output by multiple hidden layers of different depths in the deep learning coding model are extracted. The hidden representations include shallow representations, medium representations, and deep representations corresponding to the shallow features, medium features, and deep features, respectively. The handcrafted features and hidden representations belonging to the same semantic level are gated and fused. The handcrafted features and hidden representations are adaptively fused by gating weights of each dimension to obtain hierarchical fusion representations of each semantic level. We perform weighted aggregation on the hierarchical fusion representations of each semantic level to obtain a global sequence representation; The lncRNA in the RNA sequence is identified based on the global sequence characterization.

[0018] In this embodiment, the hand-designed biological features possess clear physical meaning and interpretability, and can strongly complement the abstract features automatically learned by the deep learning model. To achieve precise semantic granularity alignment between the hand-designed features and the hidden layer representation of the deep learning encoding model, and to provide structured feature input for the subsequent hierarchical gating fusion mechanism, this paper divides all hand-designed features into three scales—shallow, medium, and deep—based on the information abstraction level and biological semantic depth, corresponding to the sequence patterns captured by the shallow, medium, and deep hidden layers of the deep learning encoding model, respectively.

[0019] like Figure 1 As shown, TAGF-Net adopts a dual-branch parallel architecture, consisting of five parts: an input layer, a hierarchical feature extraction layer, a hierarchical gating fusion layer, a hierarchical attention aggregation layer, and a classification output layer. The overall data flow is as follows: Figure 1 As shown, the model's two core input branches are the DNABERT pre-trained encoding branch and the hierarchical handcrafted feature branch. The two are naturally aligned at the shallow, middle, and deep semantic levels, providing a semantic basis for hierarchical fusion.

[0020] The DNABERT pre-trained encoding branch is responsible for extracting deep representations with hierarchical semantics from the original RNA sequence, providing feature input for the pre-trained representation branch in subsequent fusion. Considering that the general BERT model is designed for natural language scenarios and has limited ability to capture the biological semantics and sequence patterns of nucleic acid sequences, DNABERT is used as the pre-trained encoding base. DNABERT is a BERT-like pre-trained language model specifically optimized for genome sequences, i.e., a deep learning encoding model. It adopts the standard Transformer encoder architecture and completes self-supervised pre-training on human reference genome sequence data. It can accurately capture the long-range contextual dependencies and intrinsic biological semantics of nucleic acid sequences, making it more suitable for RNA sequence classification tasks than the general BERT.

[0021] The sequence is first converted into a token sequence by a tokenizer, and then fed into the DNABERT model for deep encoding. This model consists of 12 stacked Transformer encoder layers, whose hidden layer representations exhibit a significant hierarchical semantic evolution: low-level encoders primarily capture primary sequence patterns such as base composition and short sequence associations; mid-level encoders learn intermediate structural features such as open reading frames and sequence motifs; and high-level encoders encode advanced semantics such as global functional attributes and translation potential. Based on this pattern, the hidden layer outputs of layers 2-4 of DNABERT are aggregated into shallow sequence representations using mean pooling; the hidden layer outputs of layers 5-8 are aggregated into mid-level structural representations; and the hidden layer outputs of layers 9-12 are aggregated into deep functional representations. These three sets of representations correspond to semantic abstraction levels from primary to advanced, forming a one-to-one semantic correspondence with the handcrafted feature layering system constructed above.

[0022] The input to the hierarchical handcrafted feature branch is the aforementioned multi-scale hierarchical handcrafted features, which are divided into three groups according to semantic hierarchy: shallow features containing 57-dimensional base composition and sequence statistical features, carrying primary sequence composition information; mid-level features containing 25-dimensional ORF structure and codon preference features, carrying mid-level reading frame structure information; and deep features containing 98-dimensional global peptide physicochemical descriptors and pseudo-amino acid composition features, carrying high-level functional semantics from the perspective of the global physicochemical properties and amino acid composition distribution of the translation product. This feature system is precisely matched with the three-layer hidden representation of DNABERT in terms of semantic abstraction, achieving hierarchical semantic alignment of "handcrafted features - pre-trained representation" and avoiding the semantic misalignment problem that is prone to occur in cross-level feature fusion.

[0023] The overall inference process of the model is as follows: the three-layer features output from the two branches are fed into three independent vector-gated fusion units to complete the fine-grained adaptive fusion of the two types of features within the same layer, and output the hierarchical fusion representation with three dimensions of 256; then the three hierarchical fusion representations are input into the hierarchical attention aggregation module, which adaptively allocates the discrimination weights of each layer and aggregates them to obtain the global sequence representation; finally, the global representation is mapped into a class probability vector through a fully connected classification head, and the classification results of lncRNA and mRNA are output.

[0024] In a preferred embodiment of the present invention, the shallow features include at least one of the following: global base content, dinucleotide frequency, k-mer frequency, coding potential score, and short-range periodicity energy; the middle features are determined based on the open reading frame regions of the RNA sequence; and the deep features are determined based on the amino acid sequence translated from the open reading frame regions.

[0025] In this embodiment, shallow features express both global sequence composition and local sequential patterns. Shallow features are the most basic sequence statistical features, calculated directly from the original base sequence without needing to identify high-level structures such as open reading frames. They primarily capture global base composition preferences, short-range adjacent base associations, and primary coding potential fingerprints. This level of features has the finest semantic granularity, involving only the statistical regularities of local short sequences, and highly corresponds to the local k-mer patterns and short-range base dependencies captured by the shallow network of the deep learning coding model. This paper designs a total of 6 sets of shallow features, totaling 57 dimensions.

[0026] GC content (S1): refers to the ratio of guanine (G) to cytosine (C) bases in the entire sequence, also known as global base content, and is a core indicator reflecting the basic compositional characteristics of the sequence. Previous studies have shown that the GC content of human protein-coding sequences is generally higher than that of long non-coding sequences, and that GC content is closely related to the secondary structural stability and post-transcriptional regulatory characteristics of RNA molecules, and can serve as a basic discriminant feature for distinguishing between two types of transcripts.

[0027] Dinucleotide frequency (S2): The relative frequency of 16 consecutive dinucleotides (AA, AC, ..., TT) in a statistical sequence, characterizing the combination preference of adjacent bases. Dinucleotide distribution is related to sequence characteristics such as CpG island enrichment and DNA methylation regulation, and can reflect the base composition pattern of transcripts under evolutionary and functional constraints.

[0028] Compressed k-mer spectra (S3): k-mer frequencies are a refined representation of sequence composition, but the feature dimension increases exponentially with the value of k (64 dimensions for k=3, 256 dimensions for k=4), easily leading to the curse of dimensionality and feature redundancy. This paper selects nucleotide frequencies at k=3 and k=4 as the original features and uses principal component analysis (PCA) to reduce their dimensionality to 32 dimensions. This compresses the feature dimension and reduces computational overhead while preserving the main variation information of local sequence patterns. To avoid data leakage, the PCA transformation matrix is ​​only fitted and generated on the training set; the validation and test sets reuse this matrix for transformation.

[0029] Fickett Score (S4): A classic coding potential assessment metric, also known as coding potential score, calculates coding propensity by statistically analyzing the differences in the positional distribution of four bases in three reading frames. It can quickly estimate the coding potential of a sequence without requiring complete ORF identification and is one of the core features of early coding prediction tools.

[0030] Global hexagonal log-likelihood ratio (S5): Based on a Markov model of hexagons (6 consecutive bases), this calculates the log-likelihood ratio of the same sequence under the coding sequence model and the non-coding sequence model. This feature reflects the coding preference of the sequence from a statistical perspective of codon combination; a higher value indicates that the sequence is more likely to have coding capabilities. Both the coding and non-coding hexagonal background models were trained solely on the training set data.

[0031] Short-range periodicity energy (S6): Protein-coding sequences generally exhibit a significant 3-nt periodicity due to the triplet structure of codons, while non-coding sequences do not have this constraint. This paper performs discrete Fourier transform on the sequences and extracts the first 6 energy features in the frequency band near period 3, capturing the unique periodic pattern of coding sequences from the frequency domain perspective, further supplementing the discriminative dimension of coding potential.

[0032] In a preferred embodiment of the present invention, the extraction of the mid-layer features includes: determining open reading frame regions from the RNA sequence based on open reading frame identification rules, wherein the open reading frame identification rules include at least one of standard open reading frame rules, start anchoring rules, end anchoring rules, and longest open reading frame region rules; The standard open reading frame rule is a complete encoded region that starts from the start codon and ends at the first stop codon within the same reading frame; The start anchoring rule starts only from the first ATG in the sequence and extends to the end of the sequence, ignoring the stop codon; The termination anchoring rule begins at the start of the sequence or after the previous stop codon and ends at the first stop codon encountered. The longest open reading frame region rule is to take the longer sequence segment between the start anchor rule and the end anchor rule.

[0033] In this embodiment, the mid-level features express the open reading frame (ORF) family and codon usage preferences. These mid-level features revolve around the structural properties of the ORF, characterizing its length distribution, region coding potential, and codon usage preferences by identifying the structure of the ORF frames in the sequence. This hierarchical feature layer involves medium-span sequence structure patterns and has a higher level of semantic abstraction than shallow sequence statistics.

[0034] Considering the large number of transcripts with incomplete 5' or 3' ends in real-world transcriptome data, a single standard ORF definition is insufficient to comprehensively characterize the coding structure of all sequences, easily leading to feature loss or bias. Therefore, this paper defines four complementary ORF identification rules to cover transcript scenarios with varying degrees of integrity: T0 (Standard ORF): The complete encoded region starting from the start codon (ATG) and ending at the first stop codon (TAG / TAA / TGA) within the same read frame. It is the most classic ORF definition. T1 (Initiation Anchored RF): Starting from the first ATG in the sequence, it extends to the end of the sequence, ignoring the stop codon, and adapts to incomplete transcripts with 3' truncation. T2 (Termination Anchored RF): Starts from the beginning of the sequence or after the previous stop codon and ends at the first stop codon encountered, adapting to incomplete transcripts with a 5' end truncation; T3 (Longest ORF-like Region): The longer sequence segment between T1 and T2 is selected as the final ORF-like region, used for feature completion of sequences without a standard ORF.

[0035] Based on the four types of ORF rules mentioned above, this paper extracts eight sets of mid-level features, totaling 25 dimensions.

[0036] Longest standard ORF length (M1): The longest nucleotide length of all ORFs in the sequence that conform to the T0 definition. It is one of the most discriminative features for distinguishing between coding and non-coding RNAs. Typically, mRNAs have complete ORFs that can be hundreds to thousands of bases long, while the standard ORFs of lncRNAs are generally shorter, often less than 100 codons.

[0037] ORF coverage (M2): The ratio of the longest standard ORF length to the total transcript length. It normalizes the absolute length of the ORF, eliminates numerical bias caused by differences in the total transcript length, and more fairly reflects the proportion of the coding region in the whole sequence.

[0038] Longest ORF Region Hexagram Score (M3): Calculates the average of the hexagram coding potential scores within the longest standard ORF region. Compared to the global hexagram likelihood ratio, this feature focuses on coding candidate regions and can more accurately reflect the coding strength of core regions.

[0039] Start codon offset (M4): The position index of the first ATG codon in the sequence relative to the 5' end. Functional mRNAs usually have a 5' untranslated region (5'UTR) of moderate length before the start codon, and its position distribution is somewhat regular; while the position of ATG in lncRNAs is more random.

[0040] Stop codon position offset (M5): The position index of the last stop codon in the sequence relative to the 3' end, corresponding to the length characteristics of the 3' untranslated region (3'UTR) of the mRNA, supplementing the positional distribution characteristics of the ORF from the terminal structure.

[0041] Effective codon count (M6, ENC): A classic indicator of codon usage preference, ranging from 20 to 61. A smaller value indicates a stronger codon usage preference. Coding sequences, due to translation selection pressure, typically exhibit significant codon usage preference; while non-coding sequences, without translation constraints, show codon usage more closely resembling a random distribution.

[0042] Reading frame phase entropy (M7): The nucleotide distribution under each of the three reading frames is statistically analyzed, and their Shannon entropy is calculated. The coding sequence exhibits a significant phase bias at the three codon bases, resulting in large differences in base distribution and low entropy values ​​across the three reading frames; the non-coding sequence has no phase constraint, leading to higher entropy values. This feature quantifies the balanced use of reading frames from the perspective of information entropy.

[0043] Longest ORF Codon Frequency (M8): This statistic tracks the relative frequency of adjacent codon pairs within the longest ORF, characterizing the combinational preferences for amino acid transitions. Compared to single codon frequencies, codon frequencies can capture local selection constraints during translational extension, containing richer information about the coding structure.

[0044] In a preferred embodiment of the present invention, the deep features include peptide physicochemical descriptors calculated based on the amino acid sequence. The peptide physicochemical descriptors are used to characterize the molecular properties of the translation products of the amino acid sequence, including at least one of the following: molecular scale properties, charge and acid-base properties, polarity and interaction properties, structural stability properties, and biochemical functional properties.

[0045] The deep features also include pseudo-amino acid composition features calculated based on the amino acid sequence. The pseudo-amino acid composition features include a first component and a second component. The first component is used to characterize the relative frequency of occurrence of each type of amino acid, and the second component is used to characterize the correlation strength of core physicochemical properties between amino acids with different intervals.

[0046] In this embodiment, deep features further map sequence information to the translation product level, calculating the molecular properties and sequence arrangement patterns of proteins based on the amino acid sequences obtained from ORF translation. These features cannot be directly observed from the base sequence and need to be deduced after translation from the standard genetic code; their semantic abstraction level is the highest among the three types of features. The core design principle is that real protein-coding sequences are subject to long-term evolutionary selection pressure, and the physicochemical properties, amino acid arrangement patterns, and structural constraints of their translation products are maintained within a reasonable range suitable for protein folding and functional execution. In contrast, the random translation products of non-coding RNA lack this evolutionary constraint, resulting in a more discrete distribution of molecular properties and a lack of regularity in sequence arrangement.

[0047] This paper selects two types of features, global peptide physicochemical descriptors and pseudo-amino acid composition, to jointly construct a deep feature system: the former describes the physicochemical phenotypic constraints of the translation product at the molecular level, while the latter supplements the sequential information of amino acid arrangement at the sequence level. The two complement each other and fully cover the global attributes and local sequential characteristics of the translation product, with a total feature dimension of 98.

[0048] Global peptide physicochemical descriptors quantitatively characterize the physicochemical properties of translation products at the molecular level, representing the most direct set of features reflecting protein functional constraints. This paper selects 48 classic peptide physicochemical descriptors, which can be divided into five categories according to their biological meaning: Molecular scale parameters include basic indicators such as total molecular weight of side chains, total number of atoms, ratio of main chain to side chain atoms, and number of rotatable bonds, which characterize the basic molecular scale and skeletal flexibility of the translation product. Charge and acid-base properties: including isoelectric point, net charge at physiological pH, and the total proportion of positive / negative charge residues, reflecting the ionization characteristics and physiological adaptability of proteins; Polarity and interaction properties: including parameters such as total polar surface area, total number of hydrogen bond donors / acceptors, and lipid solubility coefficient, characterizing the potential tendency of proteins to participate in molecular interactions and subcellular localization; Structural stability properties include indicators such as total average hydrophilicity (GRAVY), aromaticity score, instability index, and proportion of rigid residues, which reflect the potential of the translation product to form a stable spatial conformation. Biochemical functional attributes: including the proportion of cysteine ​​residues, the proportion of hydrophobic core residues, etc., supplementing the structural and functional constraints of proteins from the functional group level.

[0049] Traditional global physicochemical indicators can only characterize the overall properties of molecules, completely losing the information on the amino acid sequence and failing to reflect the structural constraints of residue arrangement. To overcome this deficiency, the pseudo-amino acid composition (PseAAC) feature is introduced. While retaining the amino acid composition information, it deeply integrates the physicochemical properties of residues with the sequence sequence information by introducing a sequence position weighting factor, thereby characterizing the compositional preferences and local sequential patterns of the translation product.

[0050] The PseAAC feature vector consists of two parts: the first 20 dimensions represent the relative frequencies of 20 standard amino acids, corresponding to static component information; the latter λ dimension represents the sequence order correlation factor, obtained by calculating the correlation strength of three core physicochemical properties—hydrophobicity, hydrophilicity, and side chain quality—between residues at different intervals. To control the feature dimensions while retaining sufficient short-order sequence information, this paper sets the sequence order λ=30, ultimately forming a 50-dimensional PseAAC feature vector.

[0051] From a biological perspective, the amino acid arrangement of functional proteins is strictly constrained by three-dimensional structural folding and functional sites, and the physicochemical properties of adjacent residues exhibit stable cooperative arrangement patterns. In contrast, the random translation products of non-coding sequences lack evolutionary selection pressure, and their residue arrangements are closer to random distribution. PseAAC can effectively uncover the statistical differences in sequence arrangement between the two types of translation products, providing a more refined sequential discrimination basis for coding potential identification than static composition.

[0052] All deep features are calculated based on the amino acid sequence obtained by translating the longest standard ORF (T0) of the sequence; if the sequence does not have a standard ORF that meets the requirements, the T3 type ORF region is used as a substitute to ensure that all samples can output complete feature vectors.

[0053] Because the dimensions and numerical ranges of handcrafted features of different levels and types vary significantly (e.g., ORF lengths can reach thousands of bases, while GC content ranges from 0 to 1), directly inputting them into the model will cause features with larger numerical values ​​to dominate the optimization process, weakening the discriminative contribution of features with smaller numerical values. Therefore, before inputting features into the model, all 180-dimensional handcrafted features need to be standardized. This paper uses the Z-score standardization method to transform each feature dimension as follows: ; in Let be the mean of this feature dimension on the training set. This represents the corresponding standard deviation.

[0054] After standardization, all features follow a distribution with a mean of 0 and a standard deviation of 1, eliminating the interference of dimensional differences on model training. To strictly prevent data leakage, the mean and standard deviation are obtained only from the training set statistics. The validation and test sets directly reuse the statistical parameters of the training set for transformation and do not participate in the statistical calculation. In addition, for extreme samples without a standard ORF, a T3-class ORF is used as a substitute calculation region to avoid feature loss. At the same time, reasonable numerical cutoff ranges are set for all feature values ​​to suppress the impact of extreme outliers on model stability and ensure the robustness of feature input.

[0055] In summary, the constructed multi-scale hierarchical handcrafted feature system comprises 180 dimensions, achieving progressive semantic abstraction from primary base composition statistics and intermediate ORF structural patterns to advanced translational product physicochemical properties. This feature system covers the classic discriminative dimensions in lncRNA and mRNA classification tasks, and through its hierarchical design, it forms a natural semantic correspondence with the representation learning hierarchy of the BERT model, providing biologically interpretable structured input for subsequent hierarchical feature fusion models based on gating mechanisms. Furthermore, the pure sequence computation characteristic allows the entire feature system to be unrestricted by species annotation, supporting the evaluation of the model's cross-species generalization performance.

[0056] In a preferred embodiment of the present invention, the adaptive fusion of the handcrafted features and the hidden representation through dimension-wise gating weights specifically includes: The handcrafted features and the hidden representations are mapped to a common latent space of the same dimension to obtain the mapped handcrafted features and the mapped hidden representations. Based on the mapped hand-crafted features and the mapped hidden representation, a vector-level gating weight consistent with the feature dimension is generated; The mapped handmade features and the mapped hidden representation are fused element-wise using the vector-level gating weights to obtain the hierarchical fused representation.

[0057] In this embodiment, to fully integrate the biological interpretability of handcrafted features with the deep representation capabilities of pre-trained language models, a Tiered Alignment Gated Fusion Network (TAGF-Net) is proposed. The model's core design concept is "semantic hierarchical alignment - intra-hierarchical gating fusion - hierarchical attention aggregation." It finely fuses the multi-scale hierarchical handcrafted features constructed earlier with the multi-layered hidden representations of the DNABERT pre-trained model optimized for nucleic acid sequences at the corresponding semantic levels, ultimately achieving accurate identification of lncRNA and mRNA through a classification head. The network structure design is described in detail from three aspects: the overall model framework, the intra-hierarchical vector gating fusion unit, and the hierarchical attention aggregation and classification head.

[0058] The hierarchical in-vector gating fusion unit is a core innovative module of TAGF-Net, designed to address the problems of semantic misalignment, coarse granularity, and insufficient adaptability in traditional fusion methods. Under the premise of semantic alignment, this unit achieves fine-grained fusion of handcrafted features and pre-trained representations through a dimension-wise gating mechanism, automatically learning the trust weights of the two classes of information in each feature dimension. Each semantic level corresponds to an independent fusion unit, and the parameters of each unit are not shared, to adapt to the feature distribution and discriminative characteristics of different levels.

[0059] The computation process of a single fusion unit consists of three steps: feature space mapping, vector gating computation, and gating fusion output, as detailed below: Due to the significant heterogeneity between hierarchical handcrafted features and DNABERT hidden layer features in terms of dimensionality, numerical distribution, and representation space, a unified feature space mapping is first required. This mapping must first be performed linearly to map them to a common latent space of the same dimension, providing a unified representation basis for subsequent fusion. Let the i-th semantic level ( The handcrafted feature vectors (corresponding to shallow, middle, and deep layers respectively) are: ,in , , The hidden layer features of the corresponding level in DNABERT are obtained as vectors after mean pooling. The two types of features are linearly transformed using independent fully connected layers, and layer normalization (LayerNorm) and ReLU activation functions are added to achieve spatial alignment and distribution standardization, as shown in the following equation: ; ; in, , The weight matrix is ​​a learnable matrix. , This corresponds to the bias term. After mapping, the dimensions of both types of features are unified to 256, denoted as... and .

[0060] In a preferred embodiment of the present invention, the vector-level gating weights are generated by linear transformation and Sigmoid activation based on the concatenation result of the mapped handmade features and the mapped hidden representation. Each dimension of the vector-level gating weights corresponds to a dimension of the fused features and is used to adjust the fusion ratio of the handmade features and the hidden representation in that dimension.

[0061] In this embodiment, based on the mapped two types of features, the model automatically generates a gating vector with the same dimension as the features, achieving dimension-wise fusion weight adjustment. The gating vector is obtained by concatenating the two types of features and then performing a linear transformation and sigmoid activation: ; in, This indicates a vector concatenation operation, resulting in a concatenated vector with a dimension of 512. For the gated weight matrix, For bias terms; Using the Sigmoid activation function, the gate vector is... Each dimension's value is constrained to Interval.

[0062] Each element of the gating vector corresponds to a dimension of the fused features. A value closer to 1 indicates that the dimension relies more on the discriminative information of the handcrafted features, while a value closer to 0 indicates that it relies more on the pre-trained representation of DNABERT. This design differs from traditional scalar gating, which assigns a single weight to the entire feature vector, and achieves more fine-grained adaptive adjustment.

[0063] Using the generated gated vectors, the two types of mapping features are fused element-wise to obtain the final fused representation at this level. : ; in This represents the element-wise Hadamard product operation.

[0064] This integration mechanism has three significant advantages: Semantic consistency: Fusion operations are strictly limited to the same semantic level, following the cognitive logic of RNA sequences from base composition to structure to function, avoiding semantic confusion caused by direct fusion of cross-level features; Adaptive selection: The model can automatically learn the discriminative power of different feature dimensions, assigning lower weights to noisy dimensions and higher weights to strong discriminative dimensions, which is equivalent to having a built-in feature selection mechanism. Generalization adaptability: The gating weights are learned entirely by data, without the need for manual preset of the fusion ratio. In scenarios with distribution shifts such as cross-species and different sequence lengths, the dependence of the two types of features can be dynamically adjusted to improve the model's generalization ability.

[0065] In a preferred embodiment of the present invention, the weighted aggregation of the hierarchical fusion representations of each semantic level specifically includes: Attention scores for each level of fusion representation are calculated using an attention-aware network. Normalize the attention scores to obtain the hierarchical weights corresponding to each semantic level; The global sequence representation is obtained by weighting and summing the fusion representations of each level using the hierarchical weights.

[0066] In this embodiment, the three hierarchical representations obtained through hierarchical gating fusion carry biological semantics of different granularities, but the contribution of each level to the classification task varies, and there are complex complementary relationships between the levels. Directly using simple aggregation methods such as concatenation and mean pooling cannot fully explore the collaborative discriminative potential of multi-level features. Therefore, a lightweight hierarchical attention mechanism is introduced to model the importance differences between levels, adaptively aggregate global discriminative information, and then complete the final classification through a classification head.

[0067] The core idea of ​​the hierarchical attention aggregation module is to learn an adaptive global importance weight for the fused representation at each semantic level. The weight is automatically determined by the contribution of the feature at that level to the classification task. Finally, a global sequence representation integrating multi-scale semantics is obtained through weighted summation. This mechanism has a small parameter size, stable training, and the weights have clear biological interpretability.

[0068] Let the fusion characterization of shallow, intermediate, and deep layers be respectively ,and First, using a shared attention-aware network, the attention score for each level of representation is calculated: ; in, For attention transformation matrix, For the corresponding bias term, Here, we have the attention weight vector, and the three parameters are learnable parameters shared across all levels. The attention scores are then normalized using the Softmax function to obtain the final weight coefficients for each level. ; Weighting coefficient The range of values ​​is And satisfy The value directly reflects the contribution of the i-th semantic level to the final classification decision. Finally, the three-level fused representations are weighted and summed using the hierarchical weights as coefficients, and a layer normalization operation is added to stabilize the feature distribution, resulting in a global sequence representation with dimension 256. : ; To improve classification performance and suppress overfitting, a two-layer fully connected network is used to construct the classification head, and batch normalization, activation functions, and Dropout regularization strategies are embedded. The specific structure of the classification head is as follows: The first fully connected layer maps the 256-dimensional global representation to a 64-dimensional feature space, and then passes through batch normalization, ReLU activation function and dropout layer (dropout rate set to 0.3) to achieve non-linear transformation and regularization constraints of the features. The second fully connected layer maps the 64-dimensional features into a 2-dimensional output vector, corresponding to the two categories of lncRNA and mRNA. Finally, the predicted probability of each category is obtained by normalization through the Softmax function.

[0069] During the model training phase, the cross-entropy loss function is used to optimize the network parameters. The DNABERT fine-tuning parameters, gated fusion unit parameters, and classification head parameters are updated end-to-end through backpropagation to achieve joint optimization of the entire network.

[0070] The transcriptome data in this study were all obtained from internationally authoritative genome annotation databases and non-coding RNA databases. The core databases are described below: GENCODE Database: GENCODE is a high-precision genome annotation database within the ENCODE (Encyclopedia of DNA Elements) project framework. It provides manually validated transcriptome annotation information for model organisms such as humans and mice, covering various transcriptome types including protein-coding genes, long non-coding RNAs, and small non-coding RNAs. Its annotation workflow integrates high-throughput transcriptome sequencing data alignment, computational prediction, and multiple quality controls through manual review. The accuracy of transcript structural boundaries and classification annotations is widely recognized in the field, making it one of the most crucial data sources for current lncRNA research. The human lncRNA transcript data in this study were obtained from GENCODE Human v38.

[0071] The NCBI RefSeq database, constructed and maintained by the National Center for Biotechnology Information (NCBI), is the most authoritative reference sequence repository in the life sciences. This database provides multi-dimensionally validated genomic, transcriptomic, and protein reference sequences. Its mRNA transcripts include complete coding sequence (CDS) structural annotations, and the data quality is stable and regularly updated, making it the standard reference dataset for protein-coding transcript research. The human mRNA transcript data in this study were obtained from the RefSeq database.

[0072] NONCODE Database: NONCODE is a comprehensive, specialized database focusing on non-coding RNA. Currently updated to version v6, it covers lncRNA data resources from 16 animal species and 23 plant species. Its data integrates published academic literature and multiple public databases, with standardized search, identification, and annotation processes ensuring quality control throughout the entire process. In addition to basic information such as sequence, genomic location, and exon structure, it also includes multi-dimensional annotation information such as expression profiles, evolutionary conservation, functional prediction, and disease associations, providing rich data support for cross-species non-coding RNA research.

[0073] RNAcentral database: RNAcentral is the world's largest integrated database of non-coding RNAs, aggregating non-coding RNA sequences and annotation information from 44 public databases. The current version covers transcriptome data from over 250 species and includes millions of RNA secondary structure prediction results. This database supports multi-dimensional searching by RNA type, species, data source, and other dimensions, providing a comprehensive data foundation for cross-species lncRNA generalization assessment.

[0074] Human benchmark datasets are the core carriers for model training, hyperparameter tuning, and performance evaluation within the same species. Their construction process includes three core steps: raw data collection, sequence preprocessing, and dataset partitioning.

[0075] This study used human transcripts as the baseline research object, obtained human lncRNA transcript sequences from GENCODE v38 and human mRNA transcript sequences from the NCBI RefSeq database to form the initial human dataset.

[0076] To ensure data standardization, eliminate systematic bias, and balance sample distribution, a step-by-step preprocessing procedure is performed on the initial dataset. The specific process is as follows: Length screening: Based on the general classification definition of lncRNA, transcript sequences with a length of less than 200 nt in the original data were filtered out, and short sequences that did not meet the length standard were excluded.

[0077] Base normalization: The original RNA sequence uses the A / U / C / G base representation system, while current mainstream sequence feature extraction tools, sequence encoding algorithms, and deep learning models are generally designed based on the DNA base system. To unify the sequence encoding format, ensure consistency in subsequent k-mer frequency statistics, sequence vectorization, and model input, and avoid encoding bias introduced by mixing U / T, this study uniformly replaced uracil (U) with thymine (T) in all RNA sequences, so that all sequences use the standard A / T / C / G DNA base representation.

[0078] Furthermore, the original annotation data contains various degenerate mixed base symbols such as R, Y, M, K, S, W, H, B, V, and D. These symbols represent base sites that were not fully determined during sequencing or annotation, which can interfere with the statistical stability of sequence features. Therefore, this study will remove transcripts containing degenerate bases to preserve sequence integrity, reduce the impact of uncertain bases on feature calculations, and ensure the standardization of input data format.

[0079] To strictly ensure sample independence and avoid data leakage, this paper uses the CD-HIT tool to perform sequence redundancy removal on all transcripts, removing homologous transcripts with sequence similarity higher than 90%.

[0080] Class balance processing: After redundancy removal, the number of mRNA transcript samples is significantly greater than that of lncRNAs. This class imbalance can lead to model bias towards the majority class during training, weakening the ability to identify the minority class (lncRNAs) and thus overestimating the overall classification performance of the model. To eliminate the interference of class bias, this study performs random downsampling on the majority mRNA set, randomly selecting mRNA transcripts from it to make its sample size completely consistent with that of lncRNAs, thus constructing a class-balanced human benchmark dataset.

[0081] To support the entire experimental process of model training, validation, and independent testing, this study uses stratified random sampling to divide the class-balanced human benchmark dataset into training, validation, and test sets in an 8:1:1 ratio. Detailed statistics of the human benchmark dataset are shown in Table 1.

[0082] Table 1 Statistics of the Human Benchmark Dataset

[0083] Evaluating model performance solely on datasets of the same species is susceptible to species-specific sequence bias and cannot fully validate the model's true generalization ability. To comprehensively assess the model's cross-species recognition capability and eliminate performance overestimation caused by species-specific bias, an independent cross-species validation set covering eight vertebrate species was constructed. All cross-species data were not used in the model's training, validation, or hyperparameter tuning processes, but only for the final generalization performance test.

[0084] The cross-species datasets were processed using the same preprocessing standards as the human datasets, including length filtering (retaining transcripts ≥200 nt), base normalization, and class balancing to ensure data quality and consistency in input format. The validation set covers mammalian and amphibian species, including cattle (64,906 transcripts), gorillas (33,667 transcripts), rhesus monkeys (12,006 transcripts), mice (41,588 transcripts), chimpanzees (4,506 transcripts), rats (20,903 transcripts), pigs (13,379 transcripts), and Xenopus laevis (8,669 transcripts).

[0085] The experiment used the human benchmark dataset constructed above as the basis for performance evaluation. This dataset contains a total of 96,942 transcripts, of which 48,471 are lncRNA and 48,471 are mRNA. The number of samples in the two classes is completely balanced, avoiding the interference of class distribution bias on the evaluation results.

[0086] To ensure the reliability and generalization of model evaluation, the dataset was randomly stratified in an 8:1:1 ratio to ensure that the class ratio within each subset remained consistent with the full dataset. After the split, the training set contained 77,554 transcripts for iterative optimization of model parameters; the validation set contained 9,694 transcripts for hyperparameter tuning and early stopping strategy triggering; and the test set contained 9,694 transcripts for final model performance evaluation. All comparative and ablation experiments were performed on the exact same dataset split to ensure the comparability of experimental results.

[0087] To comprehensively and objectively evaluate the classification performance of the model, five common evaluation metrics for binary classification tasks were selected: Accuracy (ACC), Sensitivity (SEN), Specificity (SPE), Matthews Correlation Coefficient (MCC), and Area Under Curve (AUC). The model was comprehensively evaluated from multiple dimensions, including overall accuracy, positive example recognition ability, negative example discrimination ability, and classification robustness. The definitions and calculation formulas of each metric are as follows: Accuracy (ACC): The proportion of correctly predicted samples out of the total number of samples, reflecting the overall classification accuracy of the model.

[0088] .

[0089] Sensitivity (SEN): The proportion of samples that are actually positive (lncRNA) and are correctly predicted, which measures the model's ability to identify and recall lncRNA samples.

[0090] .

[0091] Specificity (SPE): The proportion of samples that are correctly predicted to be negative (mRNA) is measured by the model’s ability to distinguish between mRNA samples.

[0092] .

[0093] Matthews correlation coefficient (MCC): A correlation index that comprehensively considers the four categories of results in the confusion matrix. It exhibits good robustness in both class-balanced and class-imbalanced scenarios, and its value ranges from [value range missing]. The closer the value is to 1, the better the model's classification performance.

[0094] .

[0095] Area Under the Curve (AUC): The area under the receiver operating characteristic (ROC) curve, comprehensively reflecting the model's overall discriminative ability at different classification thresholds, with a value range of [value missing]. The closer the value is to 1, the stronger the overall classification performance of the model.

[0096] Wherein, TP (true positive) represents the number of samples that are actually lncRNA and correctly predicted as lncRNA, TN (true negative) represents the number of samples that are actually mRNA and correctly predicted as mRNA, FP (false positive) represents the number of samples that are actually mRNA but incorrectly predicted as lncRNA, and FN (false negative) represents the number of samples that are actually lncRNA but incorrectly predicted as mRNA.

[0097] The proposed TAGF-Net is implemented based on the PyTorch deep learning framework. Model training employs an end-to-end joint optimization approach. The DNABERT encoder uses pre-trained weights for initialization and fine-tuning, while the gated fusion units, hierarchical attention modules, and classification heads are randomly initialized. The training process uses the AdamW optimizer for parameter updates, with an initial learning rate set to... Combined with a linear learning rate decay strategy and weight decay ( Regularization was applied; the batch size was set to 32, and the maximum number of training epochs was set to 20. To prevent overfitting, an early stopping mechanism was introduced during training. When the AUC value on the validation set no longer improved for three consecutive epochs, training was automatically terminated, and the model parameters of the optimal epochs were retained. All experiments were repeated three times in a unified hardware and software environment, and the average value of each metric was taken as the final result to eliminate the influence of random factors on the experimental results and ensure the reproducibility of the experiment and the reliability of the conclusions.

[0098] Experiment Example 1: To verify the comprehensive performance advantages of TAGF-Net in the human lncRNA recognition task, five representative mainstream lncRNA recognition methods from two categories were selected as baselines for comparison, specifically including: Traditional bioinformatics tools include CPC2, CPAT, and PLEK. These methods are based on manually designed sequence features and rules for discrimination or are constructed using traditional machine learning algorithms. They are classic baseline methods in the field of lncRNA identification.

[0099] Deep learning methods: LncADeep, lncRNA_Mdeep. These methods automatically learn sequence features based on deep neural networks and represent the current mainstream technology level in the field of intelligent lncRNA recognition.

[0100] All comparison methods were tested on the exact same human benchmark set to ensure consistency between experimental conditions and evaluation criteria. The final performance comparison results are shown in Table 2.

[0101] Table 2 Performance comparison of different methods on human benchmark sets

[0102] To more intuitively compare the overall discriminative performance of each method, ROC curves for all methods were plotted on a human benchmark set. The results are as follows: Figure 2 As shown, the ROC curve (Receiving Operator Characteristic curve) uses the false positive rate on the horizontal axis and the true positive rate on the vertical axis to measure the overall performance of a classification model at different thresholds. The closer the area under the curve (AUC) value is to 1, the stronger the overall discriminative ability of the model.

[0103] Combined with Table 2 Figure 2 The results lead to the following analytical conclusions: Traditional tools suffer from significant false positive issues: CPC2 and PLEK have specificities below 0.62, and while CPAT is the best among traditional tools (ACC=0.880), it is still limited by the upper limit of manual features.

[0104] Deep learning methods significantly improve specificity: lncRNA_Mdeep (ACC=0.931, AUC=0.977) and LncADeep (ACC=0.973, AUC=0.988) are both significantly better than traditional methods.

[0105] The proposed TAGF-Net achieves the best overall classification performance: TAGF-Net's AUC reaches 0.996, an improvement of 0.008 compared to the second-best LncADeep, and its corresponding ROC curve is closest to the upper left corner, making its overall discriminative ability the best among all methods. While maintaining accuracy and MCC on par with existing best deep learning methods, TAGF-Net improves specificity to 0.967, significantly higher than traditional methods, further reducing false positives in encoding transcripts; at the same time, sensitivity remains at a high level of 0.979, resulting in better classification balance.

[0106] This performance advantage stems from the design of the hierarchical alignment gating fusion mechanism: hierarchical semantic alignment ensures that handcrafted features and pre-trained representations are fused within the same semantic level, avoiding information loss caused by cross-level semantic misalignment; the vector-level gating mechanism can adaptively select the effective dimensions of the two types of features, fully integrating the biological prior knowledge of handcrafted features with the long-program list feature capabilities of the pre-trained model, forming complementary gains. Experimental results show that TAGF-Net has good competitiveness in both classification accuracy and balance, and can provide effective technical support for the accurate annotation of lncRNAs in the human transcriptome.

[0107] Experiment Example 2: To verify the design necessity and performance contribution of each core module in TAGF-Net, five ablation variants were constructed using the controlled variable method, and control experiments were conducted on a human benchmark set. All variants maintained consistent training hyperparameters, data partitioning, and evaluation criteria. By comparing the performance with the complete model, the actual gains of the hierarchical alignment architecture, vector-level gating mechanism, adaptive fusion strategy, hierarchical attention aggregation, and handcrafted feature branches were quantified one by one.

[0108] The specific settings and validation objectives for each group of ablation variants are as follows: Abl-A: This variant removes the hierarchical alignment architecture, concatenates all hand-crafted features into a single vector, and connects it to a single gated fusion unit along with the output of the last hidden layer of DNABERT. It no longer divides the features into three groups for parallel fusion based on semantic hierarchy. This variant is used to verify the necessity of the hierarchical semantic alignment design in mitigating feature semantic misalignment and improving fusion quality.

[0109] Abl-B: Vector gating is replaced with scalar gating. The three-layer hierarchical alignment structure is retained, but the dimension-wise vector gating in each layer is replaced with a single scalar gating. This means each layer learns only one global fusion scaling coefficient and shares it across all feature dimensions. This variant was used to verify the performance advantage of fine-grained vector-level gating over coarse-grained scalar gating.

[0110] Abl-C: Removes the gating adaptation mechanism, retains the hierarchical alignment structure, eliminates gating weight learning, and directly obtains the fused representation by taking the arithmetic mean of the mapped handcrafted features and the pre-trained representation within each layer. This variant was used to verify the effect of the adaptive feature selection capability of the gating mechanism on improving fusion performance.

[0111] Abl-D: This variant removes hierarchical attention aggregation, retains the hierarchical gating fusion module, and directly concatenates the three-layer fusion representations into a 768-dimensional vector before inputting it into the classification head, replacing the original hierarchical attention weighted aggregation. This variant is used to verify the hierarchical attention mechanism's ability to adaptively weight and collaboratively mine multi-level semantics.

[0112] Abl-E: This variant removes the handcrafted feature branches, retaining only the three-layer pooling representation and hierarchical attention aggregation module of the DNABERT branch, completely removing the handcrafted feature input to form a pure pre-trained model baseline. This variant is used to validate the prior biological knowledge gain that multi-scale handcrafted features bring to the model.

[0113] Complete Model: TAGF-Net: The proposed complete architecture of a hierarchical aligned gating fusion network, serving as a performance benchmark.

[0114] The experimental results and analysis of the various performance indicators of the ablation experiment are shown in Table 3.

[0115] Table 3 Comparison of Ablation Test Performance

[0116] Based on the data in the table, the following analytical conclusions can be drawn: First, the handcrafted feature branches provide significant performance gains to the model and are an important component of the model architecture. Comparing Abl-E with the full model, removing the handcrafted features resulted in a decrease of ACC by 0.011, MCC by 0.022, and AUC by 0.006. This indicates that relying solely on the sequence representation of the pre-trained model cannot fully cover the biological prior knowledge related to encoding potential. Multi-scale hierarchical handcrafted features can provide effective information supplementation to the pre-trained representation, and the fusion of the two can form complementary gains.

[0117] Second, hierarchical semantic alignment is the core foundation for ensuring the quality of fusion. Comparing Abl-A with the complete model, after removing the hierarchical structure and adopting global single alignment fusion, the ACC decreased by 0.008 and the MCC decreased by 0.016, verifying that direct fusion of cross-level features can cause semantic misalignment, leading to mutual interference between features of different abstract granularities; while fusion according to the one-to-one correspondence of the semantic hierarchy of shallow-medium-deep layers can significantly improve the consistency and effectiveness of feature fusion.

[0118] Third, the gating adaptive fusion mechanism is key to fine-grained feature selection. Comparing Abl-C and the complete model, the ACC decreases by 0.006 when using fixed average fusion, indicating that indiscriminate feature fusion introduces noise and invalid information equally into the final representation, limiting model performance. Further comparing Abl-B and the complete model, scalar gating can only adjust the fusion ratio at the hierarchical level and cannot perform dimensional information selection, still showing an ACC difference of 0.004 compared to vector gating. This result proves that the fine-grained adaptive weight allocation of vector-level gating can accurately strengthen high-discriminative feature dimensions and suppress noise dimensions, which is one of the core innovations of the model in achieving high-precision classification.

[0119] Fourth, hierarchical attention aggregation can effectively tap the synergistic potential of multi-level features. Compared with the complete model, after replacing attention aggregation with simple concatenation, the ACC decreased by 0.003 and the MCC decreased by 0.006, indicating that different semantic levels contribute differently to classification decisions. By adaptively allocating weights through the attention mechanism, the information dominance of high-discriminative levels can be strengthened, the interference of redundant levels can be weakened, and the feature dimension can be compressed from 768 dimensions to 256 dimensions, improving performance while reducing the computational complexity of the classification head and the risk of overfitting.

[0120] In summary, all four core design features of TAGF-Net have brought about significant performance improvements. Each module performs its own function and works together to support the model's optimal classification results, fully verifying the rationality and innovation of the model architecture design.

[0121] Experimental Example 3: lncRNA sequences exhibit strong evolutionary conservation across species, but also show some differences in sequence distribution. The model's cross-species generalization ability is an important indicator for evaluating its practical application value.

[0122] Based on the constructed independent validation sets of 8 species, the recognition performance of TAGF-Net, which was trained only on the human dataset, was tested on different species to verify the model's cross-species generalization ability and the universality of sequence pattern learning.

[0123] The experiment used independent validation sets from eight species: cattle, gorillas, macaques, mice, chimpanzees, rats, pigs, and Xenopus laevis. None of these samples were used for model training or tuning. After training the model to obtain optimal parameters on the human benchmark dataset, zero-shot predictions were performed directly on the test sets of each species without any cross-species fine-tuning. Comparison models included traditional bioinformatics tools CPC2, CPAT, and PLEK, as well as deep learning methods LncADeep and lncRNA_Mdeep. All methods were evaluated using accuracy (ACC, %) on the same test set, and the results are shown in Table 4.

[0124] Table 4. Accuracy Comparison of Cross-Species Independent Validation Sets

[0125] As shown in Table 4, TAGF-Net achieved the highest accuracy on the independent validation sets of all 8 species, with an average accuracy of 95.6%, which is significantly better than all the comparison methods, demonstrating excellent and stable cross-species generalization ability.

[0126] Among traditional tools, CPC2 and CPAT performed similarly overall, with average accuracies of 91.8% and 92.0%, respectively. However, they showed significant fluctuations in primate species such as gorillas and chimpanzees. PLEK, limited by fixed artificial features, had the worst generalization stability, especially in pig species, with an accuracy of only 76.3%, making it difficult to adapt to differences in sequence distribution among species.

[0127] Deep learning methods significantly outperform traditional tools overall. lncRNA_Mdeep achieved an average accuracy of 92.8% and maintained stable recognition across most species; LncADeep, as a current mainstream deep learning model, achieved an average accuracy of 94.8% and performed exceptionally well in mammalian species.

[0128] The proposed TAGF-Net achieves stronger robustness while maintaining high accuracy: the accuracy exceeds 96% on species with high evolutionary conservation such as mice, cattle, and rats; on tropical clawed frogs with greater evolutionary distance, it still achieves the best performance of 98.9% by capturing conserved sequence patterns; even on species with smaller sample sizes and sparser feature distribution such as gorillas and chimpanzees, the model still maintains its leading position.

[0129] The results demonstrate that TAGF-Net, through its hierarchical alignment-gated fusion mechanism, effectively integrates handcrafted biological features with the universal sequence representation of DNABERT. It dynamically adjusts feature dependencies under species distribution shift scenarios, utilizing conserved coding structure features while avoiding the generalization limitations of fixed rules. Therefore, the model is not only applicable to human transcriptome annotation but can also provide reliable support for lncRNA identification in various model organisms and non-model organisms.

[0130] This invention aligns the semantic hierarchy of handcrafted features at three scales (shallow, medium, and deep) with the representations in the deep learning hidden layers, and utilizes vector-level gated fusion units to adaptively fuse the two types of information dimension by dimension. Compared with existing methods that use handcrafted features or deep learning models alone, this invention achieves better accuracy and Matthews correlation coefficient (MCC) on human benchmark datasets, significantly improving recognition accuracy.

[0131] This invention is based on a multi-scale handcrafted feature system (such as GC content, ORF structure, translation product physicochemical properties, etc.) based on pure sequence computation, without relying on species-specific annotation information; combined with a hierarchical gating fusion mechanism, it enables the model to maintain stable recognition performance on cross-species datasets other than the training species (such as humans), overcoming the shortcomings of existing deep learning methods in terms of insufficient generalization ability in cross-species scenarios.

[0132] The handcrafted feature layers of this invention (shallow sequence composition, mid-layer ORF structure, and deep translation product attributes) all have clear biological and physical meanings. The gating weights of each layer can intuitively reflect the degree of trust between handcrafted knowledge and deep representations in different feature dimensions. The hierarchical attention weights can quantify the discriminative contribution of each semantic level, making the model decision-making process clearly interpretable.

[0133] Compared to traditional experimental methods such as Northern blot and RT-PCR, this invention uses a single RNA sequence as the only input and can classify a large number of sequences into lncRNA / mRNA within seconds through an end-to-end computational model. This significantly shortens the identification cycle, reduces manpower and reagent costs, and is suitable for rapid functional annotation of large-scale transcriptome data.

[0134] Manual feature extraction, deep learning encoding, gating fusion, and hierarchical attention aggregation are all incorporated into the same network framework for end-to-end training, avoiding information loss caused by staged optimization, enabling the entire model parameters to converge collaboratively, and further improving recognition performance.

[0135] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for lncRNA recognition based on a hierarchical alignment-gated fusion network, characterized in that, The method includes: An RNA sequence is obtained, and multi-scale hierarchical handcrafted features are extracted from the RNA sequence. The multi-scale hierarchical handcrafted features are divided into shallow features, medium features, and deep features according to the semantic abstraction level. The RNA sequence is input into a deep learning coding model, and hidden representations output by multiple hidden layers of different depths in the deep learning coding model are extracted. The hidden representations include shallow representations, medium representations, and deep representations corresponding to the shallow features, medium features, and deep features, respectively. The handcrafted features and hidden representations belonging to the same semantic level are gated and fused. The handcrafted features and hidden representations are adaptively fused by gating weights of each dimension to obtain hierarchical fusion representations of each semantic level. We perform weighted aggregation on the hierarchical fusion representations of each semantic level to obtain a global sequence representation; The lncRNA in the RNA sequence is identified based on the global sequence characterization.

2. The lncRNA recognition method based on a hierarchical alignment-gated fusion network according to claim 1, characterized in that, The shallow features include at least one of the following: global base content, dinucleotide frequency, k-mer frequency, coding potential score, and short-range periodicity energy; the middle features are determined based on the open reading frame regions of the RNA sequence; and the deep features are determined based on the amino acid sequence translated from the open reading frame regions.

3. The lncRNA recognition method based on a hierarchical alignment-gated fusion network according to claim 2, characterized in that, The extraction of the mid-layer features includes: determining open reading frame regions from the RNA sequence based on open reading frame identification rules, wherein the open reading frame identification rules include at least one of the standard open reading frame rules, start anchoring rules, end anchoring rules, and longest open reading frame region rules; The standard open reading frame rule is a complete encoded region that starts from the start codon and ends at the first stop codon within the same reading frame; The start anchoring rule starts only from the first ATG in the sequence and extends to the end of the sequence, ignoring the stop codon; The termination anchoring rule begins at the start of the sequence or after the previous stop codon and ends at the first stop codon encountered. The longest open reading frame region rule is to take the longer sequence segment between the start anchor rule and the end anchor rule.

4. The lncRNA recognition method based on a hierarchical alignment-gated fusion network according to claim 3, characterized in that, The deep features include peptide physicochemical descriptors calculated based on the amino acid sequence. These peptide physicochemical descriptors are used to characterize the molecular properties of the translation products of the amino acid sequence, including at least one of the following: molecular scale properties, charge and acid-base properties, polarity and interaction properties, structural stability properties, and biochemical functional properties.

5. The lncRNA recognition method based on a hierarchical alignment-gated fusion network according to claim 4, characterized in that, The deep features also include pseudo-amino acid composition features calculated based on the amino acid sequence. The pseudo-amino acid composition features include a first component and a second component. The first component is used to characterize the relative frequency of occurrence of each type of amino acid, and the second component is used to characterize the correlation strength of core physicochemical properties between amino acids with different intervals.

6. The lncRNA recognition method based on a hierarchical alignment-gated fusion network according to claim 1, characterized in that, The adaptive fusion of the handcrafted features and the hidden representation through dimension-wise gating weights specifically includes: The handcrafted features and the hidden representations are mapped to a common latent space of the same dimension to obtain the mapped handcrafted features and the mapped hidden representations. Based on the mapped hand-crafted features and the mapped hidden representation, a vector-level gating weight consistent with the feature dimension is generated; The mapped handmade features and the mapped hidden representation are fused element-wise using the vector-level gating weights to obtain the hierarchical fused representation.

7. The lncRNA recognition method based on a hierarchical alignment-gated fusion network according to claim 6, characterized in that, The vector-level gating weights are generated based on the concatenation result of the mapped handmade features and the mapped hidden representation through linear transformation and Sigmoid activation. Each dimension of the vector-level gating weights corresponds to a dimension of the fused features and is used to adjust the fusion ratio of the handmade features and the hidden representation in that dimension.

8. The lncRNA recognition method based on a hierarchical alignment-gated fusion network according to claim 1, characterized in that, The weighted aggregation of the hierarchical fusion representations at each semantic level specifically includes: Attention scores for each level of fusion representation are calculated using an attention-aware network. Normalize the attention scores to obtain the hierarchical weights corresponding to each semantic level; The global sequence representation is obtained by weighting and summing the fusion representations of each level using the hierarchical weights.