Enzyme mining method and system and storage medium

By using the VenusRXN model and optimizing cross-modal feature alignment and fusion loss values, the problem of inaccurate association between enzymes and catalytic reactions in existing enzyme mining methods is solved, achieving high-precision and high-efficiency enzyme mining, which is suitable for enzyme mining of novel reactions.

CN121415892APending Publication Date: 2026-01-27SHANGHAI MATWINGS TECHNOLOGY CO LTD
0 Cites 1 Cited by

Patent Information

Application Number
CN202511618142.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing enzyme discovery methods rely on protein similarity and functional annotation, which makes it difficult to accurately associate enzymes with their catalytic chemical reactions, and their generalization ability is insufficient, making it impossible to effectively discover enzymes for novel reactions.

Method used

The VenusRXN model is used to mine enzymes with specific catalytic functions from protein databases based on a given chemical reaction or template enzyme through a reaction encoder and a protein encoder. The model is optimized by cross-modal feature alignment and fusion loss value to reduce the dependence on protein similarity and functional annotation.

Benefits of technology

It improves the accuracy and speed of enzyme mining, can generalize to novel reactions, reduces dependence on protein similarity and functional annotation, and enhances the practicality of enzyme mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121415892A_ABST
    Figure CN121415892A_ABST
Patent Text Reader

Abstract

The invention provides an enzyme mining method, system and medium. The enzyme mining method comprises the following steps: inputting to-be-queried data; inputting the data to be queried into a pre-trained VenusRXN model; when the input to-be-queried data is a reaction, determining a reaction CLS embedding vector by using a reaction encoder, mapping the reaction CLS embedding vector into a reaction query embedding vector through a linear layer corresponding to the reaction encoder, calculating the similarity between the reaction query embedding vector and each protein embedding vector, and sorting the similarity; outputting a protein sequence corresponding to the plurality of protein embedding vectors with the highest similarity; or when the input to-be-queried data is the template enzyme, determining a template enzyme CLS embedding vector by using the protein encoder, mapping the template enzyme CLS embedding vector into a template enzyme query embedding vector through a linear layer corresponding to the protein encoder, calculating the similarity between the template enzyme query embedding vector and each protein embedding vector, and sorting the similarity to obtain the template enzyme query embedding vector. And outputting the protein sequences corresponding to the plurality of protein embedding vectors with the highest similarity.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Cited By

  • Enzyme function prediction method based on deep contrast learning

    CN122090963A