基于剪枝预训练模型与人工特征编码融合的4mC位点识别方法

By fusing the pruned pre-trained model DNABert with artificial feature encoding, the problem of insufficient DNA sequence feature representation was solved, the recognition accuracy of the 4mC site was improved, and more accurate prediction results were achieved.

CN117216656BActive Publication Date: 2026-07-17GUANGDONG UNIV OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG UNIV OF TECH
Filing Date
2023-09-07
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing DNA sequence feature characterization capabilities are insufficient, resulting in low prediction accuracy in 4mC site identification, especially with overfitting and computational resource consumption issues on small datasets.

Method used

We employ a method that combines the pruned pre-trained model DNABert with artificial feature encoding. We expand the feature representation space through Kmer encoding and CKSNAP encoding, and combine a bidirectional LSTM network and an attention fusion module to extract deep and shallow feature information. Finally, we use a feedforward neural network for classification and prediction.

Benefits of technology

It significantly improves the identification accuracy of 4mC sites on six independent benchmark test datasets, outperforming existing models and achieving more accurate prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117216656B_ABST
    Figure CN117216656B_ABST
Patent Text Reader

Abstract

本发明提供了一种基于剪枝预训练模型与人工特征编码融合的4mC位点识别方法,包括以下步骤:S1:获取DNA‑4mC的核苷酸序列集L,通过Kmer编码将序列集L转化为数值向量集M;S2:将预训练模型DNABert进行剪枝压缩操作得到的模型DNABert‑Pruning作为基准模型,训练得到序列集L的深层特征信息;S3:根据CKSNAP编码特征方式扩充各核苷酸在序列集L中的特征表示空间,通过双向LSTM网络训练得到序列集L的浅层特征信息;S4:将上述训练得到的浅层信息与深层信息特征同时输入到特征融合注意力模块中,得到更为准确的融合特征表征;S5:对融合后的表征特征使用前馈神经网络和Sigmoid函数输出识别预测,计算其分类评分,本发明可以更准确地预测DNA‑4mC的位点识别信息。
Need to check novelty before this filing date? Find Prior Art