Automatic construction method and system of active polypeptide

Through the combination of large language models and biological simulation tools, active peptides are automatically generated, solving the problem of cumbersome screening process and labor-consuming material resources in the existing technology, and achieving efficient and automated peptide acquisition.

CN120452523APending Publication Date: 2025-08-08COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510478641.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing active peptide screening methods are cumbersome, labor-consuming and material-intensive, and low-efficiency, making it difficult to efficiently screen out specifically bound polypeptides from metabolites and polypeptide libraries.

Method used

Using the Big Language Model (LLM) combined with biological simulation tools, the active peptides are automatically generated through the interaction analysis of target proteins and metabolites, reducing blindness and improving screening efficiency.

Benefits of technology

The automatic generation of active peptides is achieved, which significantly improves the efficiency of peptide acquisition, simplifies the screening process, saves time and labor costs, and is suitable for batch large-scale metabolite active peptide construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452523A_ABST
    Figure CN120452523A_ABST
Patent Text Reader

Abstract

The invention provides an automatic construction method and system of an active polypeptide, and the method comprises the following steps: firstly, determining a target metabolite, and obtaining a target protein having binding activity with the target metabolite; then, obtaining interaction information of the target protein and the target metabolite and secondary structure information of target protein residues; then, inputting the interaction information into a large language model LLM to obtain information of key residues of the target protein; and finally, inputting the structural information of the target protein, the information of the key residues, the secondary structural information and the constraint information into LLM to obtain an active polypeptide capable of being specifically combined with the target metabolite, namely the potential peptide aptamer. According to the method, by combining a biological simulation tool and a large language model, automatic batch generation of the active polypeptide is realized, the acquisition process of the polypeptide is greatly simplified, and a new strategy and a new way are provided for metabolite peptide aptamer design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to interdisciplinary subjects such as computer science, structural biology and nutrition, and in particular to a method and system for automatically generating metabolite-bound active polypeptides using a large language model. Background Art

[0002] Metabolites are small molecule compounds involved in metabolism within organisms, primarily including carbohydrates, lipids, amino acids, nucleotides, organic acids, vitamins, and hormones. These metabolites directly participate in biological metabolic activities such as signal transduction and epigenetic regulation, reflecting biological events that have already occurred and can be used to detect the physiological or pathological states of biological systems. Studies have found that metabolites are closely associated with a variety of diseases, including cancer, diabetes, Alzheimer's disease, and cardiovascular disease. With the development of metabolomics technology, an increasing number of metabolite markers have been discovered and widely used in fields such as growth monitoring, health assessment, and disease early warning. In addition, metabolite detection has shown broad application potential in the field of precision medicine and plays a key role in promoting the upgrading of global health management systems.

[0003] Aptamers are single-stranded DNA, RNA, or peptide molecules that can specifically bind to target molecules. Due to their increased chemical stability and ease of synthesis and modification, they are considered potential alternatives to antibodies. Aptamers possess excellent biocompatibility, high specificity, and high sensitivity, and have been widely used in biosensing, drug delivery, disease diagnosis, and environmental monitoring in recent years. Aptamers are short peptides that specifically bind to target substances. First proposed in 1996 by Roger Brent of the Fred Hutchison Cancer Research Center, they are typically composed of 5-20 amino acids. Aptamers have rapidly developed as recognition elements and hold great potential due to their high affinity, strong specificity, excellent biocompatibility, extremely low detection limits, and ease of modification. Identifying active peptides that specifically bind to specific metabolites is a key step in peptide aptamer design. However, the cumbersome and labor-intensive peptide screening process has hindered research on peptide aptamers.

[0004] In recent years, the development of molecular docking simulation technology has significantly improved the screening efficiency of peptide aptamers, making it possible to screen high-binding metabolite peptide aptamers. The molecular docking algorithm is a process in which two or more molecules recognize each other through geometric matching and energy matching. It can help researchers predict the binding mode of receptors and ligands, determine the mechanism of action, and is the basis for designing new drugs in drug development. However, when applying molecular docking technology to peptide aptamer screening, due to the wide variety of plant and animal metabolites, for example, the Human Metabolome Database (HMDB) includes more than 200,000 human-related metabolites. From the perspective of peptides, the structure of a peptide composed of 5-20 amino acids can reach 10 6 -10 12 Each peptide has multiple structural conformations. To find an active peptide that specifically binds to a specific metabolite from such a large library of metabolites and peptides, traditional docking methods require traversing all metabolite and peptide libraries, making the screening process as difficult as finding a needle in a haystack.

[0005] In addition, due to the wide variety of metabolites and peptides, researchers need to face a large number of molecular docking results. For each molecular docking result, they need to use professional software to analyze the interaction pattern between the peptide and the metabolite. The types of interactions include hydrogen bonds, salt bridges, water bridges, and π stacking. Each interaction also has multiple different parameters. It is necessary to comprehensively consider multiple factors to screen out the key interactions. After obtaining the key interactions, researchers also need to conduct an overall evaluation of the binding pattern between the peptide and the metabolite in order to finally obtain the target peptide. As can be seen from the above, the active peptide screening method in the existing technology is cumbersome and consumes a lot of manpower and material resources, which is a major obstacle to the application and research of peptide aptamers. Summary of the Invention

[0006] The present invention aims to overcome the problems of existing active peptide screening methods, such as large screening scales, cumbersome processes, low efficiency, and high labor and material consumption. The present invention rationally designs active peptides based on metabolite structures and utilizes large language models (LLMs) to automatically generate active peptides, significantly improving the efficiency of obtaining active peptides.

[0007] Based on this, the present invention provides an automated construction method and system for active polypeptides.

[0008] According to a first aspect, the present invention provides a method for automatically constructing an active polypeptide, comprising the following steps:

[0009] Identify target metabolites;

[0010] Obtaining a target protein that has binding activity with the target metabolite;

[0011] Obtaining interaction information between the target protein and the target metabolite complex through an interaction analysis tool;

[0012] Obtaining secondary structure information of the target protein residues through a secondary structure analysis tool;

[0013] Inputting the interaction information into the large language model (LLM) to obtain information about key residues of the target protein;

[0014] The structural information, key residue information, secondary structure information and constraint information of the target protein are input into LLM to obtain active polypeptides that can specifically bind to the target metabolite, namely potential peptide aptamers.

[0015] In some embodiments, obtaining the target protein having binding activity with the target metabolite specifically includes: obtaining the target protein having binding activity with the target metabolite by searching a protein-ligand database.

[0016] In some embodiments, obtaining a target protein having binding activity with the target metabolite specifically includes: using reverse docking technology, using the target metabolite as a ligand, and performing molecular docking within the active pocket of a protein in a protein library to obtain multiple possible binding conformations and their corresponding binding free energies, and thereby screening out several proteins with the lowest binding free energy with the target metabolite as target proteins.

[0017] In some embodiments, the binding free energy between the target protein and the target metabolite is less than a specific threshold, such as -6 kcal / mol.

[0018] In some embodiments, the number of target proteins does not exceed 5.

[0019] In some embodiments, the interaction information includes: the serial number and type of residues in the target protein that interact with the target metabolite, and the type and strength of the interaction;

[0020] In some embodiments, the key residue information includes: the name, number, interaction type and importance description of the key residue.

[0021] In some embodiments, the constraint information includes: the obtained polypeptide contains at least one key residue and retains the residue secondary structure characteristics to the greatest extent.

[0022] According to a second aspect, the present invention further provides an active polypeptide construction system, comprising:

[0023] a determination module for determining target metabolites;

[0024] A target protein acquisition module, used to obtain a target protein that has binding activity with the target metabolite;

[0025] An interaction information acquisition module, used to obtain the interaction information between the target protein and the target metabolite complex;

[0026] A secondary structure information acquisition module, used to obtain the secondary structure information of the target protein residues;

[0027] A key residue extraction module is used to input the interaction information into the large language model (LLM) to obtain information on key residues of the target protein;

[0028] The peptide building module is used to input the structural information of the target protein, the information of key residues, the secondary structure information of the residues and the constraint information into the LLM to obtain an active peptide that can specifically bind to the target metabolite, i.e., a potential peptide aptamer.

[0029] In some embodiments, the target protein acquisition module is specifically configured as follows: using reverse docking technology, the target metabolite is used as a ligand to perform molecular docking with proteins in a protein library within the active pocket to obtain multiple possible binding conformations and their corresponding binding free energies, and based on this, several proteins with the lowest binding free energy with the target metabolite are screened out as target proteins.

[0030] In some embodiments, the constraint information includes: the obtained polypeptide contains at least one key residue and retains the residue secondary structure characteristics to the greatest extent.

[0031] The method for automatically generating active polypeptides provided by the present invention designs polypeptides for target metabolite structures, which reduces blindness and significantly improves screening efficiency compared to screening polypeptides from a polypeptide library. The present invention realizes automated batch generation of active polypeptides by combining biosimulation tools with large language models. The construction and screening process of polypeptides does not require human participation, which greatly simplifies the process of obtaining polypeptides. The entire process from positioning the target complex to complex interaction analysis, confirmation of key residues, and construction of active polypeptides in combination with secondary structures is automated, realizing the automated generation of active polypeptides, saving time and labor costs, and significantly improving the efficiency of polypeptide acquisition. It is particularly suitable for the construction of metabolite active polypeptides on a large scale in batches, and provides new strategies and approaches for the design of metabolite peptide aptamers. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 A flow chart of the automated construction method for active polypeptides provided by the present invention;

[0034] Figure 2 This is a diagram of the architecture of the automated construction method for active polypeptides provided by the present invention;

[0035] Figure 3 This is a schematic diagram of the structure of the active polypeptide construction system provided by the present invention. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0037] The present invention belongs to the field of theoretical design of molecular structures. Unless otherwise specified, the determination, acquisition, generation, construction, etc. of metabolites, polypeptides, proteins and other substances mentioned below all refer to the determination of their molecular three-dimensional structures, rather than the acquisition, generation, and construction of the above substances at the physical level.

[0038] The present invention aims to overcome the problems of existing active polypeptide screening methods such as large screening scale, complicated process, low efficiency, and high consumption of manpower and material resources.

[0039] Existing docking methods require traversing all metabolite and peptide libraries, and the screening difficulty is like finding a needle in a haystack. However, if active peptides can be rationally designed based on the metabolite structure, it will be more targeted and can significantly reduce the screening level, thereby improving the screening efficiency.

[0040] Based on this, the present invention is based on the structure of the protein-metabolite complex, and through in-depth analysis of its interaction pattern, generates polypeptides with binding activity. Specifically, the present invention uses molecular simulation technology to analyze the key binding sites and mechanisms of action between proteins and metabolites, and designs short peptide structures that can efficiently bind to target metabolites through large language models and prompt word engineering. The active polypeptide generation method provided by the present invention can not only efficiently and automatically generate active polypeptides, but also significantly improve their binding affinity and specificity, providing an innovative solution for metabolite detection and related applications.

[0041] The following is a detailed description of the method for automatically constructing active polypeptides provided by the present invention in conjunction with the accompanying drawings. Figure 1 As shown in the architecture diagram Figure 2 As shown. Figure 1 The technical solution of the present invention mainly includes two stages: one is the confirmation of the parent target protein, and the other is the use of LLM to achieve short peptide interception of the parent target. Each stage is described in detail below.

[0042] Phase 1: Confirmation of the parental target protein. This phase is divided into two steps:

[0043] S1. Identify target metabolites.

[0044] The target metabolite refers to a metabolite that the active polypeptide to be constructed can specifically bind to. The target metabolite can be an animal or plant metabolite, such as glucose, alanine, ATP, oleic acid, vitamin C, citric acid, histamine, etc. The molecular structure of the target metabolite can be obtained from the HMDB database or other metabolite databases.

[0045] The present invention rationally designs active polypeptides based on the metabolite structure, which is more targeted, can significantly reduce the screening level and improve the screening efficiency.

[0046] S2. Obtaining a target protein that has binding activity with the target metabolite.

[0047] This step can be achieved in two ways:

[0048] Method 1: By calling the API, the protein-ligand database is searched to obtain proteins that have binding interactions with known metabolites as target proteins.

[0049] This method is applicable to the case where the target metabolite is a known metabolite and the target protein information of the target metabolite exists in the protein-ligand database.

[0050] The protein-ligand database may be Protein Data Bank (PDB), PDBbind, etc.

[0051] The number of target proteins obtained through retrieval does not exceed 5 to balance accuracy and computational efficiency.

[0052] Method 2: Using reverse docking technology, the target metabolite is used as a ligand to perform molecular docking with proteins in the protein library in the pocket to obtain multiple possible binding conformations and their corresponding binding free energies. Based on this, several proteins with the lowest binding free energy with the target metabolite are screened out as target proteins.

[0053] This method can be implemented using molecular docking software such as AutoDock Vina. Specifically, the target metabolite is used as the ligand, and the proteins in the protein library are used as the receptor library. Using AutoDock Vina software, the ligand and receptor are docked in the pocket. The software will provide possible binding conformations and their corresponding binding free energies. Based on these, several proteins with the lowest binding free energies with the target metabolite can be automatically selected as target proteins according to preset conditions.

[0054] The pre-set conditions include: (1) the binding energy between the target protein and the target metabolite is less than -6 kcal / mol to ensure that the obtained ligand has good binding activity with the receptor. (2) For each target metabolite, the number of target proteins obtained does not exceed 5 to balance accuracy and computational efficiency.

[0055] The protein library can be the HMDB protein library, the PDB protein library, etc.

[0056] The reverse docking technique is a computer-based drug design method. Unlike traditional molecular docking methods, it simultaneously docks a small molecule ligand with multiple different biomacromolecules to rapidly identify potential targets with high affinity for the small molecule, thus providing new insights and directions for drug discovery. This method utilizes reverse docking, starting from the target metabolite structure, to identify target proteins through molecular docking within the binding pocket. Compared with traditional peptide screening methods, this method is more targeted and can significantly improve screening efficiency.

[0057] The “molecular docking in the pocket” refers to a molecular simulation technology used in drug design, which aims to predict the binding mode and orientation of ligands (metabolite small molecules) in specific pockets (active sites) of receptors (proteins). Specifically, the ligand is placed in the pocket area of the receptor, and the possible binding conformations are obtained by searching and adjusting the conformation of the ligand. The binding energy of each binding conformation is then evaluated and the conformations are screened accordingly. This technology needs to first determine the docking pocket of the receptor. There are usually two ways to determine the binding pocket. One way is that if the receptor contains an endogenous ligand, the area where the endogenous ligand is located is selected as the active pocket. Another way is to use pocket prediction tools such as DoGSite, DeepSite and other tools to predict the active pocket.

[0058] After the active pocket is confirmed, molecular docking software such as AutoDock Vina, QuickVina, etc. can be used to perform molecular docking calculations within the binding pocket.

[0059] The above is a specific implementation method of using reverse docking technology to screen target proteins.

[0060] The above are two implementation methods of step S2. In practical applications, the two methods can be used alone or in combination. When used in combination, the protein-ligand database is first searched. If no target protein is found, the reverse docking technology is used to search for the target protein.

[0061] The above is the first stage of the automated active peptide generation method provided by the present invention. This stage screens parent target proteins that have binding activity with the target metabolite based on the target metabolite. The next second stage analyzes the interaction pattern between the target protein and the target metabolite. Based on the analysis results, short peptides of the parent target are extracted to generate peptides with binding activity. The second stage is described in detail below.

[0062] Phase 2: LLM is used to extract short peptides from the parent target. This phase includes four steps, S3-S6:

[0063] S3. Obtaining interaction information between the target protein and the target metabolite complex.

[0064] The target protein and target metabolite obtained in S2 are molecularly docked to obtain a target-metabolite complex. This step analyzes the interaction between the target protein and the target metabolite based on the target-metabolite complex.

[0065] This step can be achieved using target-metabolite complex interaction analysis tools, such as the Protein-Ligand Interaction Profiler (PLIP). PLIP can analyze the interaction of a given target-metabolite complex structure, analyzing the residue numbers and types (e.g., 58A, TYR, donor atoms, acceptor atoms, etc.) that interact with the metabolite, the types of interactions (e.g., hydrogen bonds, salt bridges, water bridges, and π stacking), and the strength of interactions (e.g., hydrogen bond donor-acceptor distances and angles), and then compile this interaction information into a binding feature analysis data file (affinity_data).

[0066] S4. Obtaining the secondary structure information of the target protein residues.

[0067] This step can be achieved using Pymol software. The secondary structure of the residue can be alpha-helix, beta-sheet, loop, or random coil. The secondary structure information of each residue in the target protein can be obtained by Pymol analysis. The format of the secondary structure information of the residue output by Pymol is as follows:

[0068] 'L':'1-5,10-11,16-18','S':'6-9,12-15','H':'19-23,29-35'.

[0069] Where H stands for alpha-helix, L stands for loop, S stands for beta-sheet, and L:1-5 indicates that the secondary structure corresponding to residues 1-5 is a loop.

[0070] The above introduces the methods for obtaining interaction information and secondary structure information. The following describes how to use LLM to determine key interactions.

[0071] S5. Input the interaction information into the large language model (LLM) to obtain information on key residues of the target protein.

[0072] In this step, the affinity_data file, extracted from the PLIP tool in S3, along with the prompts, is input into LLM to identify key residues. LLM then polishes and describes the interaction information, analyzes the most critical interactions based on this information, identifies key residues, and extracts these key residues and their corresponding positions, generating the key residue numbers, residue names, residue numbers, interaction categories, and importance descriptions.

[0073] The prompts can include physicochemical knowledge relevant to interaction significance assessment. This knowledge helps LLM understand interaction information and assists in extracting key residues. This knowledge can be derived from empirical experience, textbooks, relevant literature in the field, or knowledge graphs in the field. LLM can also gather this knowledge from publicly available online information based on the prompts. This knowledge might include: shorter-range interactions contribute significantly to energy, and near-linear hydrogen bonds provide strong binding energy.

[0074] The LLM can be DeepSeek R1, GPT-4, or a medicinal chemistry domain-specific LLM fine-tuned with instructions from the medicinal chemistry domain. Domain-specific LLMs understand domain-specific language patterns, proper nouns, and concepts, can solve complex problems within that domain, and demonstrate exceptional performance within that domain, enabling accurate analysis of interaction information.

[0075] The following is an example of the prompt word template used in this step:

[0076]

[0077] The above describes the specific process of using LLM to analyze interaction information and thus obtain key residue information. The following describes how to construct active peptides based on the obtained information.

[0078] S6. Input the target protein's structural information, key residue information, secondary structure information, and constraint information into LLM to obtain an active polypeptide that can specifically bind to the target metabolite.

[0079] The structural information of the target protein is obtained in S2 and can be input into LLM in the form of a pdb file.

[0080] If multiple target proteins were obtained in step S2, each target protein corresponds to a corresponding pdb file. Only one target protein pdb file, along with its corresponding key residue information, secondary structure information, and constraint information, is input into the LLM at a time. Multiple inputs are performed multiple times to obtain multiple active peptides, all of which are potential peptide aptamers that meet the requirements. If the number of obtained peptides exceeds a certain threshold (e.g., 20 peptides), they are sorted by the number of interactions, with peptides with a large number of interactions being preferentially selected as active peptides.

[0081] The information of the key residues is obtained in S5, including the sequence number of the key residue, residue name, residue number, interaction category and importance description.

[0082] The secondary structure information is obtained in S4.

[0083] The constraint information is given in the form of prompt words, specifically including: limiting the length of the output polypeptide sequence to 8-20 amino acid residues, the obtained polypeptide contains at least one key residue, and retains the secondary structure information to the greatest extent.

[0084] The following is an example of the prompt word template used in this step:

[0085]

[0086] The above details the specific implementation process of the automated construction method of the active polypeptide provided by the present invention.

[0087] It can be seen that compared with the prior art, the method for automatically generating active polypeptides provided by the present invention has the following beneficial effects:

[0088] 1. The present invention designs polypeptides based on the target metabolite structure, which reduces blindness and significantly improves the efficiency of polypeptide acquisition compared to screening polypeptides from a polypeptide library.

[0089] 2. By combining biosimulation tools with large language models, automated batch generation of active peptides is achieved. The peptide construction and screening process does not require human intervention, which greatly simplifies the peptide acquisition process, saves time and labor costs, and significantly improves the efficiency of peptide acquisition. It is especially suitable for the large-scale construction of metabolite active peptides.

[0090] 3. The polypeptides generated by the present invention contain key residues that bind to metabolites, while also retaining secondary structure information to the greatest extent, significantly improving the efficiency and success rate of metabolite peptide aptamer discovery, and providing new strategies and approaches for metabolite peptide aptamer design.

[0091] On the other hand, the present invention also provides an active polypeptide construction system 300, which can be used to automatically and batch-produce active polypeptides according to target metabolites. The structural diagram of the system is shown in FIG. Figure 3 As shown, the system includes:

[0092] A determination module 301 is used to determine a target metabolite;

[0093] A target protein acquisition module 302 is used to acquire a target protein that has binding activity with the target metabolite;

[0094] An interaction information acquisition module 303 is used to acquire interaction information between the target protein and the target metabolite complex;

[0095] A secondary structure information acquisition module 304 is used to obtain the secondary structure information of the target protein residues;

[0096] A key residue extraction module 305 is used to input the interaction information into a large language model (LLM) to obtain information on key residues of the target protein;

[0097] The polypeptide construction module 306 is used to input the target protein's structural information, key residue information, residue secondary structure information and constraint information into the LLM to obtain an active polypeptide that can specifically bind to the target metabolite, i.e., a potential peptide aptamer.

[0098] In some embodiments, the target protein acquisition module is specifically configured as follows: using reverse docking technology, the target metabolite is used as a ligand to perform molecular docking in the pocket with the protein in the protein library to obtain multiple possible binding conformations and their corresponding binding free energies, and based on this, several proteins with the lowest binding free energy with the target metabolite are screened out as target proteins.

[0099] In some embodiments, the constraint information includes: the obtained polypeptide contains at least one key residue and retains the residue secondary structure characteristics to the greatest extent.

[0100] It should be noted that the above system can execute the above active polypeptide generation method. The functions of each module can be found in the above introduction to the method and will not be described in detail.

[0101] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components.

[0102] Here, we provide a specific implementation of this system, designed as a web server with a front-end using Vue.js and NGL visualization software, and a back-end based on Flask and LangChain. This system integrates the peptide design process and supports batch processing of short peptides. Users simply upload one or more metabolites, and the system automatically generates and provides downloads for multiple corresponding active peptide structures, achieving efficient automation and batch processing of peptide design.

[0103] (1) Front-end design

[0104] Component-based design breaks down complex user interfaces into multiple reusable components, improving development efficiency and code maintainability. It integrates multiple visualization tools, such as ECharts for displaying peptide design and analysis results and NGL for displaying the three-dimensional structure of protein-metabolite complexes, to help users intuitively understand molecular interactions.

[0105] (2) Backend design

[0106] A RESTful API was developed using the FLASK framework to support users uploading metabolite data and obtaining peptide design results. Combined with task queue tools like Celery, this allows for asynchronous processing of peptide design tasks, improving the system's concurrent processing capabilities. The LangChain framework was combined with large language models (such as DeepSeek R1) to generate peptide structures that efficiently bind to target metabolites.

[0107] (3) Data processing and analysis

[0108] It involves the data management of the protein library and the management of the generated peptide library. The protein library is mainly used for reverse docking to find target proteins. The results of peptide generation are stored in the database, which also includes information storage such as context history.

[0109] The above is a specific embodiment of the active polypeptide construction system provided by the present invention.

[0110] In the description of the embodiments of the present application, words such as "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of the present application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.

[0111] In the description of the embodiments of this application, the term "and / or" is merely a description of an association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent the following three situations: A exists alone, B exists alone, and A and B exist at the same time. In addition, unless otherwise specified, the term "plurality" means two or more.

[0112] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly identifying the technical features being referred to. Thus, features specified as "first" or "second" may explicitly or implicitly include one or more of such features. The terms "include," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0113] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for automatically constructing an active polypeptide, characterized in that: The following steps are involved: Identify target metabolites; Obtaining a target protein that has binding activity with the target metabolite; Obtaining interaction information between the target protein and the target metabolite complex; Obtaining secondary structure information of the target protein residues; Inputting the interaction information into the large language model (LLM) to obtain information about key residues of the target protein; The structural information, key residue information, secondary structure information and constraint information of the target protein are input into LLM to obtain an active polypeptide that can specifically bind to the target metabolite.

2. The method according to claim 1, characterized in that The method of obtaining a target protein having binding activity with the target metabolite specifically includes: using a reverse docking technique, taking the target metabolite as a ligand, and performing molecular docking within the active pocket of a protein in a protein library to obtain multiple possible binding conformations and their corresponding binding free energies, and thereby screening out several proteins with the lowest binding free energy with the target metabolite as target proteins.

3. The method according to claim 1 or 2, characterized in that The number of target proteins does not exceed 5.

4. The method according to claim 1 or 2, characterized in that The binding free energy between the target protein and the target metabolite is less than a specific threshold.

5. The method according to claim 1, wherein The interaction information includes: the serial number and type of residues in the target protein that interact with the target metabolite, and the type and strength of the interaction.

6. The method according to claim 1, characterized in that The information of the key residues includes: the name, number, interaction type and importance description of the key residues.

7. The method according to claim 1, characterized in that The constraint information includes: the obtained polypeptide contains at least one key residue and retains the residue secondary structure characteristics to the greatest extent.

8. An active polypeptide construction system comprising: a determination module for determining target metabolites; A target protein acquisition module, used to obtain a target protein that has binding activity with the target metabolite; An interaction information acquisition module, used to obtain the interaction information between the target protein and the target metabolite complex; A secondary structure information acquisition module, used to obtain the secondary structure information of the target protein residues; A key residue extraction module is used to input the interaction information into the large language model (LLM) to obtain information on key residues of the target protein; The peptide building module is used to input the structural information of the target protein, the information of key residues, the secondary structure information of the residues and the constraint information into the LLM to obtain an active peptide that can specifically bind to the target metabolite.

9. The system according to claim 8, characterized in that The target protein acquisition module is specifically configured as follows: using reverse docking technology, the target metabolite is used as a ligand to perform molecular docking within the active pocket of proteins in the protein library, thereby obtaining multiple possible binding conformations and their corresponding binding free energies, and based on this, several proteins with the lowest binding free energy with the target metabolite are screened out as target proteins.

10. The system according to claim 8, wherein: The constraint information includes: the obtained polypeptide contains at least one key residue and retains the residue secondary structure characteristics to the greatest extent.

Citation Information

Cited By

  • Autonomous evolutionary drug discovery and delivery collaboration method and system based on large language model

    CN121747689A

  • An autonomous evolutionary drug discovery and delivery collaborative method and system based on a large language model

    CN121747689B