Methods, electronic devices, computer-readable storage media, and systems for predicting microbial phenotypes based on genomes

By combining DNABERT-2 with K-mer statistical methods and CCA-PLS analysis, and integrating the deep semantic and local structural features of the genome sequence, this method solves the problems of high cost, low efficiency, and insufficient accuracy in the prediction of probiotic acid and bile salt tolerance in existing technologies, and achieves efficient and accurate strain screening and prediction.

CN122417147APending Publication Date: 2026-07-17INNER MONGOLIA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610571371.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-28
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for screening and predicting the acid and bile salt tolerance of probiotics suffer from high cost, low efficiency, poor stability, and insufficient prediction accuracy and generalization ability. In particular, traditional methods cannot effectively integrate global semantics and local statistical features, leading to feature redundancy and noise amplification.

Method used

By combining a DNABERT-2 pre-trained Transformer model with the K-mer statistical method, and fusing deep semantic features and local structural features of the genome sequence through the CCA-PLS analysis method, a global-local multi-scale feature system was formed to predict the acid and bile salt tolerance of microorganisms.

Benefits of technology

It achieves high-throughput, low-cost strain screening, improves prediction accuracy and generalization ability, can complete screening without in vitro experiments, significantly enhances the model's discriminative ability and stability across datasets, and is suitable for parallel screening of multiple phenotypes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122417147A_ABST
    Figure CN122417147A_ABST
Patent Text Reader

Abstract

本发明涉及一种基于基因组预测微生物表型的方法、电子设备、计算机可读存储介质和系统,通过DNABERT‑2预训练Transformer模型对基因组序列进行编码,生成序列级深度语义特征向量;同时提取局部K‑mer统计特征,形成局部结构特征向量;采用CCA‑PLS分析方法对上述两种特征进行融合与降维,得到融合特征矩阵;最后将融合特征输入至训练好的分类预测模型中,输出微生物的耐受性表型预测结果;本方案通过结合全局语义嵌入与局部结构特征,利用多视图特征融合技术有效提升了微生物表型预测的准确性与鲁棒性,可广泛应用于耐酸性、耐胆盐性、耐热性、产酸能力及抗氧化能力等核心工业表型的高通量预测筛选。
Need to check novelty before this filing date? Find Prior Art