Biomolecular sequence searching method and apparatus, device, and storage medium

By using deep learning and efficient similarity search algorithms, biomolecular sequences are converted into vector representations and a vector database is constructed. This solves the problems of low computational efficiency and insufficient accuracy of existing tools in large-scale databases, and achieves fast and accurate homologous sequence identification.

WO2026114095A1PCT designated stage Publication Date: 2026-06-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-11-20
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing homology search tools are computationally inefficient and struggle to accurately capture sequence similarities when dealing with large-scale biomolecular sequences, failing to meet the needs of modern biological research.

Method used

Using deep learning methods, biomolecular sequences are converted into vector representations through sequence coding models. Then, an efficient similarity search algorithm is used to construct a biomolecular sequence vector database, which stores the vector representations of known sequences and performs similarity calculations and screening of homologous sequences.

Benefits of technology

It improves the accuracy and efficiency of homologous biological sequence searching, enabling rapid identification of homologous sequences in large-scale databases and meeting the needs of biological research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025136467_04062026_PF_FP_ABST
    Figure CN2025136467_04062026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a biomolecular sequence searching method and apparatus, a device, and a storage medium. The method comprises: acquiring an unknown biomolecular sequence; inputting the unknown biomolecular sequence into a sequence encoding model for sequence encoding, to obtain an unknown biomolecular sequence representation; performing similarity searching on different candidate vector sets in a biomolecular sequence vector database by means of the unknown biomolecular sequence representation, wherein candidate biomolecular sequences of candidate vectors in the candidate vector sets have a same sequence length; on the basis of the similarity searching result, searching each candidate vector set for a preset number of known biomolecular sequence representations; and screening known biomolecular sequences corresponding to the plurality of known biomolecular sequence representations for homologous biomolecular sequences corresponding to the unknown biomolecular sequence.
Need to check novelty before this filing date? Find Prior Art