A method for screening for glucagon-like peptide-1 receptor agonists

CN122889044APending Publication Date: 2026-10-09BEIJING GAOPAI INTELLIGENT TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611206709.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-10
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0004]然而,现有GLP1R多肽激动剂筛选技术存在以下缺陷:其一,依赖类药规则与传统人工分子描述符表征多肽,难以捕获多肽序列上下文、修饰结构等专属特征,分子对接难以适配多肽高柔性,亲和力预测误差大;其二,缺少理化性质、聚集风险、靶点亲和力一体化分级筛选链路,高聚集风险多肽直接进入体外检测,易造成实验数据失真并大幅增加研发耗材与人力成本;其三,传统浅层模型或单一预训练网络无法融合多肽分子序列、蛋白一级序列、蛋白三维拓扑多源信息,缺少多肽-蛋白交叉交互建模模块,难以自动识别关键结合残基,对新型多肽骨架预测泛化性及准确度不足

Benefits of technology

本申请实施例提供的一种胰高血糖素样肽-1受体激动剂筛选方法,方法包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122889044A_ABST
    Figure CN122889044A_ABST
Patent Text Reader

Abstract

The application provides a screening method for a glucagon-like peptide-1 receptor agonist, which comprises the following steps: constructing a polypeptide aggregation prediction model and an affinity prediction model, screening a first batch of candidate molecules according to molecular descriptors of polypeptide candidates to be predicted; inputting the SMILES sequence of each first batch of candidate molecules into the polypeptide aggregation prediction model to determine the aggregation, filtering the high-aggregation molecules in the first batch of candidate molecules by using the aggregation to obtain a second batch of candidate molecules; inputting the SMILES sequence of the second batch of candidate molecules, the GLP-1R protein sequence and the GLP-1R protein structure into the affinity prediction model to determine the binding affinity prediction result; and screening the GLP-1R agonist molecules according to the binding affinity prediction result. By using the above-mentioned screening method for the glucagon-like peptide-1 receptor agonist, the problem of large prediction error of the binding affinity and high screening cost in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical technology, and more specifically, to a method for screening glucagon-like peptide-1 receptor agonists. Background Technology

[0002] Glucagon-like peptide-1 receptor (GLP-1) 1R) belongs to the G protein-coupled receptor family and is a core drug target for the treatment of type 2 diabetes and obesity, targeting GLP. 1R peptide agonists, with their advantages of high target selectivity, low off-target toxicity, and mild action, have become a core direction in the current research and development of drugs for lowering blood sugar and reducing weight. Compared to small molecule compounds, peptide molecules can interact with GLP-12 molecules. 1R forms a wide range of interfacial interactions and has a stronger potential for stimulatory activity. However, the high flexibility of the amino acid chain of polypeptide molecules, the complexity of their physicochemical properties, and their tendency to aggregate and precipitate pose a huge challenge to high-throughput screening.

[0003] Traditional drug development processes rely on wet screening to validate candidate molecules one by one, resulting in long screening cycles and high material costs. Therefore, the industry commonly adopts a combined approach of computational virtual screening and in vitro experimental rescreening. Computer models are used to preemptively eliminate candidates with poor drug-like properties and weak binding activity, reducing experimental scale and improving development efficiency. Currently, GLP... The 1R ligand virtual screening system uses drug-likeness rule filtering, molecular docking, SPR affinity detection, and pharmacodynamic validation as its core links. A representative scheme is the patent CN120847037A published by Anhui Medical University. This scheme relies on the Lipinski five rules and Veber rules to complete the initial screening of small molecules, and then uses Libidock and CDOCKER molecular docking to score and screen candidates. Finally, the SPR biosensor system is used to determine the molecular affinity of GLP-9-hydroxyl radicals (GLP-9-hydroxyl radicals). The actual affinity of 1R, combined with pharmacodynamic validation via protein immunoblotting, can reduce the number of candidate molecules to some extent, making it suitable for small molecule compound screening scenarios.

[0004] However, existing GLP The 1R peptide agonist screening technology has the following drawbacks: First, it relies on drug-like rules and traditional artificial molecular descriptors to characterize peptides, making it difficult to capture specific features such as peptide sequence context and modified structures. Molecular docking is also difficult to adapt to the high flexibility of peptides, resulting in large affinity prediction errors. Second, it lacks an integrated hierarchical screening chain that considers physicochemical properties, aggregation risk, and target affinity. Peptides with high aggregation risk are directly included in in vitro detection, which can easily lead to distorted experimental data and significantly increase R&D consumables and labor costs. Third, traditional shallow models or single pre-trained networks cannot integrate multi-source information such as peptide molecular sequence, protein primary sequence, and protein three-dimensional topology. They lack peptide-protein cross-interaction modeling modules, making it difficult to automatically identify key binding residues and resulting in insufficient generalization and accuracy in predicting novel peptide backbones. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a method for screening glucagon-like peptide-1 receptor agonists to overcome at least one of the above-mentioned defects.

[0006] In a first aspect, embodiments of this application provide a method for screening glucagon-like peptide-1 receptor agonists, comprising: Construct peptide aggregation prediction models and affinity prediction models; Determine the molecular descriptors corresponding to the peptide candidates to be predicted to describe the physicochemical properties of the drug, and screen multiple first-batch candidate molecules that meet the physicochemical property requirements from the peptide candidates to be predicted based on the molecular descriptors. The SMILES sequence of each first batch of candidate molecules is input into the peptide aggregation prediction model to determine the aggregation of each first batch of candidate molecules. The aggregation is then used to filter the first batch of candidate molecules for highly aggregated molecules to obtain multiple second batch of candidate molecules. The SMILES sequence, GLP-1R protein sequence, and GLP-1R protein structure information of each second batch of candidate molecules are input into the affinity prediction model to determine the predicted binding affinity between each second batch of candidate molecules and the GLP-1R protein. Multiple second-batch candidate molecules were sorted according to the binding affinity prediction results in order to screen out GLP-1R agonist molecules that meet the affinity requirements.

[0007] In an optional implementation, the step of inputting the SMILES sequence, GLP-1R protein sequence, and GLP-1R protein structure information of each second batch of candidate molecules into the affinity prediction model to determine the predicted binding affinity between each second batch of candidate molecules and the GLP-1R protein includes: determining the target peptide embedding vector corresponding to the SMILES sequence of each second batch of candidate molecules, the target target residue embedding matrix corresponding to the GLP-1R protein sequence, and the target target contact map matrix corresponding to the GLP-1R protein structure information; and performing masking and normalization on the target target contact map matrix. A first-order adjacency weight matrix is ​​obtained, and multi-scale contact propagation is performed on the residue embedding projection matrix using the first-order adjacency weight matrix to obtain multi-scale fused residue features. Topological encoding is performed on the masked target contact map to obtain topological embedding. The multi-scale fused residue features are then updated with weights using the topological embedding and peptide embedding projection vectors to determine the site-weighted residue features. The peptide embedding projection vector is obtained based on the linear projection of the target peptide embedding vector. Multi-head attention mechanism, site-weighted pooling, and multilayer perceptron are used to fuse the site-weighted residue features and peptide embedding projection vectors to determine the binding affinity prediction results.

[0008] In an optional implementation, the steps of fusing site-weighted residue features and peptide embedding projection vectors using a multi-head attention mechanism, site-weighted pooling, and a multilayer perceptron to determine the binding affinity prediction result include: determining the peptide fusion characterization based on the site-weighted residue features and peptide embedding projection vectors; performing layer normalization on the superimposed vectors corresponding to the peptide fusion characterization and peptide embedding projection vectors, and inputting it into a feedforward neural network to obtain the peptide embedding fusion vector; adding type embedding to the peptide embedding fusion vector and the site-reweighted residue features to obtain the peptide characterization and target characterization; and concatenating the target characterization with the peptide characterization, fitting the concatenation result through a multilayer perceptron to obtain the binding affinity prediction result.

[0009] In an optional implementation, the steps of masking and normalizing the target contact map matrix to obtain a first-order adjacency weight matrix, and using the first-order adjacency weight matrix to perform multi-scale contact propagation on the residue embedding projection matrix to obtain multi-scale fused residue features include: masking the target contact map matrix to obtain a masked target contact map; performing row-wise summation and normalization on the masked target contact map to obtain a first-order adjacency weight matrix; using the first-order adjacency weight matrix to perform weighted aggregation on the residue embedding projection matrix to obtain a one-hop neighborhood context and a two-hop neighborhood context for each residue; and determining multi-scale fused residue features based on the residue embedding projection matrix, the one-hop neighborhood context, and the two-hop neighborhood context, wherein the residue embedding projection matrix is ​​obtained based on the linear projection of the target target residue embedding matrix.

[0010] In an optional implementation, the steps of performing topological encoding on the masked target contact map to obtain topological embedding, and using the topological embedding and peptide embedding projection vector to weight and update the multi-scale fused residue features to determine the site-weighted residue features include: summing each residue row-wise on the masked target contact map to obtain the contact degree of the residue, and determining the local clustering coefficient corresponding to the residue based on the contact degree; splicing and linearly projecting the contact degree and local clustering coefficient to obtain the topological embedding of the residue; splicing and multi-layer sensing of the multi-scale fused residue features, one-hop neighborhood context, topological embedding and peptide embedding projection vector to obtain an effective site score; and using the effective site score to weight the multi-scale fused residue features to determine the site-weighted residue features.

[0011] In an optional implementation, the affinity prediction model is constructed through the following process: deduplication of peptide entities and target entities in the sample table to construct peptide sets and target sets; based on the peptide sets and target sets, determining peptide embedding vectors, target residue embedding matrices, and target contact map matrices; and training the affinity prediction model using the peptide embedding vectors, target residue embedding matrices, and target contact map matrices to obtain the trained affinity prediction model.

[0012] In an optional implementation, the target contact map matrix is ​​determined through the following processes: For each target amino acid sequence in the target set, based on the length of the longest target sequence in the sample set, a corresponding effective mask matrix for the target sequence is constructed according to whether a residue exists at each residue index position; the experimental analytical structure of each target entity is obtained and parsed, and the residue spatial metadata of each standard amino acid residue is extracted to construct a chain-level residue recording sequence; each target amino acid sequence is compared with the corresponding chain-level residue recording sequence to determine whether each residue position is successfully paired, and when the residue pairing is successful, the residue spatial metadata is backfilled to the corresponding residue position of the target amino acid sequence to obtain a three-dimensional coordinate matrix and an effective residue mask, wherein the residue spatial metadata includes carbon atom coordinates; based on the carbon atom coordinates of the effectively mapped residues, the residue pair distance matrix corresponding to each target amino acid sequence is determined; the residue pair distances in the residue pair distance matrix are compared with a set distance threshold, and the residue pair distance matrix is ​​converted into a target contact map matrix based on the comparison results.

[0013] In an optional implementation, the peptide embedding vector is determined by the following process: inputting the SMILES sequence of each peptide entity in the peptide set into the peptide adaptation chemistry model to determine the peptide embedding vector of the peptide. The peptide adaptation chemistry model is a model based on a molecular language model and superimposed with a peptide lightweight fine-tuning module.

[0014] In an optional implementation, the target residue embedding matrix is ​​determined by the following process: encoding each target amino acid sequence in the target set using a protein language model to obtain the target residue embedding matrix.

[0015] In an optional implementation, a peptide aggregation prediction model is constructed by the following process: using the SMILES sequence of each peptide in the peptide set as the input column and whether it aggregates as the label column to construct a dataset; and using the dataset to train the molecular property prediction model to obtain the peptide aggregation prediction model.

[0016] The embodiments of this application bring the following beneficial effects: This application provides a method for screening glucagon-like peptide-1 receptor agonists, the method comprising: A peptide aggregation prediction model and an affinity prediction model were constructed. Molecular descriptors for describing the pharmacological and chemical properties of the candidate peptides to be predicted were determined. Based on these descriptors, multiple first-batch candidate molecules meeting the physicochemical property requirements were screened from the candidate peptides. The SMILES sequence of each first-batch candidate molecule was input into the peptide aggregation prediction model to determine the aggregation of each first-batch candidate molecule. High-aggregation molecules were filtered from the first-batch candidate molecules using aggregation to obtain multiple second-batch candidate molecules. The SMILES sequence, GLP-1R protein sequence, and GLP-1R protein structure information of each second-batch candidate molecule were input into the affinity prediction model to determine the binding affinity prediction result between each second-batch candidate molecule and the GLP-1R protein. The multiple second-batch candidate molecules were ranked according to the binding affinity prediction results to screen out GLP-1R agonist molecules meeting the affinity requirements.

[0017] This application's embodiments can perform initial screening of peptide physicochemical properties using molecular descriptors, then eliminate highly aggregated peptides using a peptide aggregation prediction model, and finally use the SMILES sequence, GLP-1R protein sequence, and GLP-1R protein structure of a second batch of candidate molecules to conduct affinity modeling and scoring screening, constructing a three-layer progressive virtual screening link, reducing affinity prediction bias, significantly reducing the number of peptides entering wet experiments, effectively compressing the R&D cycle and experimental costs. Compared with the existing glucagon-like peptide-1 receptor agonist screening method, it solves the problems of large binding affinity prediction errors and high screening costs in the existing technology.

[0018] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart of the glucagon-like peptide-1 receptor agonist screening method provided in the embodiments of this application is shown; Figure 2 A flowchart illustrating the training steps of the affinity prediction model provided in the embodiments of this application is shown; Figure 3 A flowchart illustrating the steps for determining the binding affinity prediction results provided in an embodiment of this application is shown; Figure 4 A flowchart illustrating the steps for determining the multi-scale fusion residue features provided in an embodiment of this application is shown; Figure 5 A flowchart illustrating the steps for determining site-weighted residue features provided in an embodiment of this application is shown. Figure 6 A flowchart of the affinity prediction steps provided in the embodiments of this application is shown; Figure 7 A schematic diagram of the glucagon-like peptide-1 receptor agonist screening device provided in an embodiment of this application is shown. Figure 8 A schematic diagram of the structure of the electronic device provided in the embodiments of this application is shown. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.

[0022] To facilitate understanding of this embodiment, the following describes each of the exemplary steps provided in this embodiment using the glucagon-like peptide-1 receptor agonist screening method provided in this application embodiment as an example of its application to a terminal device.

[0023] Please see Figure 1 , Figure 1 This is a flowchart illustrating a method for screening glucagon-like peptide-1 receptor agonists provided in an embodiment of this application. Figure 1 As shown in the embodiments of this application, the glucagon-like peptide-1 receptor agonist screening method includes: Step S101: Construct a peptide aggregation prediction model and an affinity prediction model.

[0024] The peptide aggregation prediction model is a molecular property prediction model trained with peptide SMILES (Simplified Molecular-Input Line-Entry System) sequences and aggregation tag datasets. The peptide aggregation prediction model takes the peptide SMILES sequence as input and the risk assessment result of peptide aggregation and precipitation as output. It is used to remove peptide candidates that are prone to aggregation failure in batches after physicochemical screening, avoid highly aggregated peptides from entering subsequent affinity prediction, reduce experimental data distortion, and reduce computational and experimental resource consumption.

[0025] In one embodiment, before constructing the peptide aggregation prediction model and the affinity prediction model, the peptide entities and target entities in the sample table are deduplicated to construct the peptide set and the target set.

[0026] Specifically, since the sample table may contain multiple entries for the same peptide or target, directly using the sample table for model training would lead to problems such as unbalanced sample weights, overfitting, and redundant feature calculations. Therefore, we can first deduplicate peptide and target entities separately, extracting unique sets of peptides and targets to ensure a balanced dataset distribution and improve the training stability and prediction generalization ability of both types of prediction models.

[0027] For example, the sample table records peptide identifiers (pep_id), peptide SMILES sequences (pep_smiles), target identifiers (target_id), target amino acid sequences (target_seq), and affinity tags. During deduplication, the sample table can be read, and peptide entities can be deduplicated according to the peptide identifier pep_id, retaining the first occurrence of the SMILES sequence to obtain a peptide set. Simultaneously, target entities can be deduplicated according to the target identifier target_id, retaining the first occurrence of the amino acid sequence to obtain a target set. Note that the target amino acid sequence (target_seq) itself does not include spatial metadata such as three-dimensional coordinates, residue numbering B-factors, etc.

[0028] In one embodiment, when constructing a peptide aggregation prediction model, a dataset is constructed using the SMILES sequence of each peptide in the peptide set as the input column and whether or not it aggregates as the label column; the molecular property prediction model is trained using the dataset to obtain the peptide aggregation prediction model.

[0029] Specifically, the SMILES sequences of peptide entities are extracted from the peptide set as input columns. Label columns are generated based on whether the peptides aggregate, as recorded in the open-source dataset. A dataset is constructed based on the input and label columns, and then divided into training, validation, and test sets using a backbone balancing method, with each set comprising 0.8, 0.1, and 0.1% of the dataset, respectively. The molecular property prediction model refers to the molecular property modeling framework Chemprop. Supervised classification training of Chemprop using a partitioned dataset yields a peptide aggregation prediction model. The input to the peptide aggregation prediction model is the SMILES sequence, and the output is a binary label indicating whether it aggregates.

[0030] In one embodiment, when constructing an affinity prediction model, peptide embedding vectors, target residue embedding matrices, and target contact map matrices can be determined based on peptide sets and target sets; the affinity prediction model can be trained using peptide embedding vectors, target residue embedding matrices, and target contact map matrices to obtain a trained affinity prediction model.

[0031] The following reference Figure 2 This section will introduce the training process of the affinity prediction model.

[0032] Figure 2 A flowchart illustrating the training steps of the affinity prediction model provided in this application embodiment is shown, as follows: Figure 2 As shown, the training steps for the affinity prediction model include: Step S201: Input the SMILES sequence of each polypeptide entity in the polypeptide set into the polypeptide adaptation chemistry model to determine the polypeptide embedding vector of the polypeptide.

[0033] Peptide adaptation chemistry models are based on molecular language models (such as ChemBERTa-77M-MLM) and superimposed with peptide lightweight fine-tuning modules (such as PepDoRA).

[0034] For each peptide entity in the peptide set, a pre-trained molecular language model, ChemBERTa-77M-MLM, is selected as the base model. The peptide lightweight fine-tuning module PepDoRA is loaded onto the base model to calculate the peptide embedding vector. The peptide embedding vector is a standardized digital feature of the peptide, which transforms the SMILES chemical structure of the peptide into a continuous numerical vector of fixed dimensions.

[0035] For example, the SMILES sequence is encoded into a token sequence by a word segmenter. The token sequence is then padded and truncated. The padded and truncated token sequence is input into a peptide adaptation chemistry model. The hidden states of the last layer are taken, and then weighted average pooling is performed on the hidden states of the valid tokens according to the attention mask to obtain a fixed-dimensional peptide embedding vector, denoted as p. Here, the attention mask is a binary vector of the same length as the peptide SMILES and protein amino acid token sequences. 1 represents a real valid token (real peptide atom / amino acid residue), and 0 represents a padded placeholder. It is used in the attention calculation and vector pooling stages of Transformer-type pre-trained models (ChemBERTa, ESM2) to shield the interference of invalid padded positions.

[0036] Each peptide identifier pep_id corresponds to a peptide embedding vector. Based on the correspondence between peptide identifiers and peptide embedding vectors, a peptide embedding dictionary can be constructed to directly determine the peptide embedding vector corresponding to the peptide candidate to be predicted.

[0037] Step S202: Encode each target amino acid sequence in the target set using a protein language model to obtain a target residue embedding matrix.

[0038] For each target amino acid sequence in the target set, a protein language model (such as ESM2 Long) is used to encode the target amino acid sequence.

[0039] For example: Based on the longest target amino acid sequence length in the sample and the maximum positional encoding limit of the model, the maximum encoding length of the protein sequence, max_pro_len, is determined. The target amino acid sequence is input into the protein language model. After the protein language model outputs multiple layers of hidden states, the last layer of residue-level hidden states is extracted. Then, based on the effective mask of the target amino acid sequence, invalid vectors corresponding to placeholders are removed, retaining only the embedding vectors matching the actual amino acid residues, generating a target residue embedding matrix. The size of the target residue embedding matrix is ​​the number of residues × the embedding dimension, denoted as X.

[0040] Each target identifier target_id corresponds to a target residue embedding matrix. Based on the correspondence between the target identifier and the peptide embedding vector, a protein embedding dictionary can be constructed to directly determine the target residue embedding matrix corresponding to the GLP-1R protein sequence.

[0041] Step S203: For each target amino acid sequence in the target set, based on the length of the longest target sequence in the sample set, construct the corresponding effective mask matrix of the target sequence according to whether there is a residue at each residue index position.

[0042] To ensure consistent processing of the input data for each model, we assume that the longest target amino acid sequence in the sample table is M and the embedding dimension of the protein language model is embed_dim. Using the length M of the longest target amino acid sequence as the mask reference, we construct a binary mask matrix (effective mask matrix of target sequence) m of length M for each target identifier target_id corresponding to the target amino acid sequence.

[0043] For example, in the target amino acid sequence, the mask at the first n positions with actual corresponding residues is set to 1, and the mask at the remaining filling positions is set to 0. Here, n is the minimum value between the length of the target amino acid sequence and the maximum encoding length max_pro_len.

[0044] Step S204: Obtain and parse the experimental structure of each target entity, and extract the residue spatial metadata of each standard amino acid residue to construct a chain-level residue recording sequence.

[0045] Based on the UniProt accession number of the target entity, query the PDB ID (Protein Data Bank) associated with the protein in the UniProt knowledge base, and download the experimental resolution structure of the target entity. If the target entity lacks an experimental resolution structure, use the AlphaFold model to predict and obtain the experimental resolution result for the target entity.

[0046] The experimental structure file was parsed to obtain the standard amino acid residues of each protein chain. The residue space metadata of each standard amino acid residue was extracted. The residue space metadata includes the coordinates of the α-carbon atom (Cα), the residue number and atomic spatial alignment information, and the residue validity marker (including the residue deletion marker) to construct the chain-level residue record sequence based on the residue space metadata.

[0047] Step S205: Compare each target amino acid sequence with the corresponding chain-level residue recording sequence to determine whether each residue position is successfully paired. When the residue pairing is successful, fill the residue spatial metadata back into the corresponding residue position of the target amino acid sequence to obtain a three-dimensional coordinate matrix and an effective residue mask.

[0048] Each target amino acid sequence target_seq is globally compared with the amino acid sequence in the corresponding chain-level residue recording sequence.

[0049] A three-dimensional coordinate matrix with dimensions of L rows and 3 columns is set up, where L represents the total number of residues in the target amino acid sequence, and the 3 columns correspond to the x, y, and z three-dimensional spatial coordinates. Based on the global alignment results, each residue position i of the target amino acid sequence `target_seq` is traversed. If a residue in the target amino acid sequence successfully matches a residue in the chain-level residue recording sequence, the residue spatial metadata of the corresponding residue in the chain-level residue recording sequence is extracted and backfilled into the residue index i of the three-dimensional coordinate matrix and the residue index i of the target amino acid sequence `target_seq`, and the mask value at residue index i is marked as 1. If a residue in the target amino acid sequence fails to match a residue in the chain-level residue recording sequence, the coordinate field corresponding to residue position i in the three-dimensional coordinate matrix and the residue index i of the target amino acid sequence `target_seq` are marked as null, and the mask value at residue index i is marked as 0. The three-dimensional coordinate matrix is ​​obtained based on the backfilling results of the residue spatial metadata, and the effective residue mask is obtained based on the mask value at each residue index position.

[0050] Step S206: Based on the carbon atom coordinates of the effectively mapped residues, determine the residue pair distance matrix corresponding to each target amino acid sequence.

[0051] Effective mapping residues are determined based on effective residue masks. Only residues with a mask value of 1 in the effective residue mask are considered effective mapping residues. Effective mapping residues refer to amino acid residues that successfully match the target amino acid sequence target_seq with the chain-level residue recording sequence after global alignment, can be backfilled into the residue index, and have effective α-carbon atom three-dimensional coordinates.

[0052] For each target amino acid sequence, based on the residue index of the effectively mapped residues in the target amino acid sequence target_seq, the α-carbon coordinates of the effectively mapped residues are selected from the three-dimensional coordinate matrix, and the Euclidean distance between any two effectively mapped residues' α-carbon coordinates is determined, generating a residue pair distance matrix. The residue pair distance matrix is ​​an L×L matrix, where L represents the length of the target amino acid sequence target_seq (target_len). Each target amino acid sequence corresponds to an effective residue mask, a three-dimensional coordinate matrix, and a residue pair distance matrix.

[0053] Step S207: Compare the residue pair distances in the residue pair distance matrix with a set distance threshold, and convert the residue pair distance matrix into a target contact map matrix based on the comparison results.

[0054] Using a set distance threshold as a boundary, if the distance between residue pairs is less than the set distance threshold, the contact value at the corresponding position in the distance feature matrix is ​​set to 1; if the distance between residue pairs is greater than or equal to the set distance threshold, the contact value at the corresponding position in the distance feature matrix is ​​set to 0, thus obtaining a distance feature matrix O with determined values. The distance feature matrix is ​​an L×L matrix. The distance feature matrix is ​​then expanded into a target contact map matrix C by filling in zeros. The target contact map matrix is ​​M×M, where M represents the length of the longest target amino acid sequence.

[0055] Similarly, a target contact map dictionary can be constructed based on the target identifier and the target contact map matrix.

[0056] After determining the peptide embedding vector, target residue embedding matrix, and target contact map matrix, the affinity prediction model can be trained using these matrixes. The training process for the affinity prediction model is the same as the usage process. For details on training the affinity prediction model using the peptide embedding vector, target residue embedding matrix, and target contact map matrix, please refer to the subsequent usage process of the affinity prediction model. It will not be repeated here.

[0057] Step S102: Determine the molecular descriptor corresponding to the peptide candidate to be predicted for describing the physicochemical properties of the drug, and screen multiple first-batch candidate molecules that meet the physicochemical property requirements from the peptide candidate to be predicted based on the molecular descriptor.

[0058] The amino acid sequence or structure of the candidate peptide is converted into a SMILES sequence. The SMILES sequence is then input into cheminformatics software (such as RDKit), which is used to calculate molecular descriptors to describe the pharmacological and chemical properties of the peptide. These pharmacological and chemical properties can refer to the peptide's hydrophobicity, hydrophilicity, and lipophilicity-related physicochemical properties.

[0059] Then, based on the physicochemical property requirements, the candidate peptides to be predicted are screened for physicochemical properties, and molecules within the preset range of physicochemical properties are retained as the first batch of candidate molecules. For example, the physicochemical property requirement can be that the topological polar surface area (TPSA) is in the range of 1200 to 2200 square angstroms.

[0060] Step S103: Input the SMILES sequence of each first batch of candidate molecules into the peptide aggregation prediction model to determine the aggregation of each first batch of candidate molecules. Use the aggregation to filter the first batch of candidate molecules for high aggregation molecules to obtain multiple second batch of candidate molecules.

[0061] The SMILES sequences of the first batch of candidate molecules are input into a pre-trained peptide aggregation prediction model. The peptide aggregation prediction model filters the first batch of candidate molecules, removing highly aggregated molecules and retaining the low-aggregated first batch of candidate molecules. The retained low-aggregated first batch of candidate molecules is called the second batch of candidate molecules, which is directly output by the peptide aggregation prediction model.

[0062] Step S104: Input the SMILES sequence, GLP-1R protein sequence and GLP-1R protein structure information of each second batch of candidate molecules into the affinity prediction model to determine the binding affinity prediction result between each second batch of candidate molecules and GLP-1R protein.

[0063] The SMILES sequence, GLP-1R protein sequence, and GLP-1R protein structure information of the second batch of candidate molecules output by the peptide aggregation prediction model are input into the constructed affinity prediction model. The affinity prediction model predicts the binding affinity of each second-frequency candidate molecule to GLP-1R, and obtains the binding affinity prediction results.

[0064] The following reference Figure 3 This section will introduce the process for determining the results of affinity prediction.

[0065] Figure 3 A flowchart illustrating the steps for determining binding affinity prediction results provided in an embodiment of this application is shown, as follows: Figure 3 As shown, the steps for determining the affinity prediction results include: Step S301: Determine the target peptide embedding vector corresponding to the SMILES sequence of each second batch of candidate molecules, the target target residue embedding matrix corresponding to the GLP-1R protein sequence, and the target target contact map matrix corresponding to the GLP-1R protein structure information.

[0066] If a target polypeptide entity matching the second batch of candidate molecules can be found in the polypeptide embedding dictionary, the polypeptide embedding vector corresponding to the target polypeptide entity in the polypeptide embedding dictionary is directly used as the target polypeptide embedding vector; if a target polypeptide entity matching the second batch of candidate molecules cannot be found, the target polypeptide embedding vector corresponding to the second batch of candidate molecules is directly calculated according to step S201.

[0067] Similarly, if a target entity matching the GLP-1R protein sequence can be found in the protein embedding dictionary, the target residue embedding matrix corresponding to the target entity in the protein embedding dictionary is directly used as the target residue embedding matrix; if a target entity matching the GLP-1R protein sequence cannot be found, the target residue embedding matrix corresponding to the GLP-1R protein sequence is directly calculated according to step S202.

[0068] If a target entity matching the GLP-1R protein structure information can be found in the target contact map dictionary, the target contact map matrix corresponding to the target entity in the target contact map dictionary is directly used as the target target contact map matrix; if a target entity matching the GLP-1R protein structure information cannot be found, the target target contact map matrix corresponding to the GLP-1R protein structure information is directly calculated according to steps S203 to S207.

[0069] Step S302: Mask and normalize the target contact map matrix to obtain a first-order adjacency weight matrix. Use the first-order adjacency weight matrix to perform multi-scale contact propagation on the residue embedding projection matrix to obtain multi-scale fused residue features.

[0070] The following reference Figure 4 This section will introduce the process of determining the characteristics of multi-scale fusion residues.

[0071] Figure 4 A flowchart illustrating the steps for determining the multi-scale fusion residue features provided in an embodiment of this application is shown, as follows: Figure 4 As shown, the steps for determining the characteristics of multi-scale fusion residues include: Step S3011: Perform linear projection on the target peptide embedding vector and the target target residue embedding matrix to obtain the peptide embedding projection vector and the residue embedding projection matrix.

[0072] Specifically, the affinity prediction model includes a linear mapping layer, in which the target peptide embedding vector p and the target residue embedding matrix X are linearly projected onto a unified hidden dimension d, i.e.: ; ,in, H represents the peptide embedding vector; H represents the residue embedding projection matrix. , Represents the learnable projection matrix; , This indicates the bias term.

[0073] Step S3012: Mask the target contact map matrix to obtain the masked target contact map. Then, perform row-wise summation and normalization on the masked target contact map to obtain the first-order adjacency weight matrix.

[0074] The target contact map matrix C is a matrix padded with a uniform length M. Therefore, it includes invalid regions such as zero-filled residues. If the target contact map matrix C is used directly for neighborhood aggregation, invalid positions may be mistakenly treated as neighbors and included in the weighting. Therefore, the target contact map matrix can be sequentially masked and row normalized to obtain a first-order adjacency weight matrix defined only on valid residue pairs.

[0075] Specifically, the effective mask matrix of the target sequence in the target contact map matrix C is m. In the effective mask matrix of the target sequence, if... If , then it indicates that the i-th residue is a real and valid residue; if If , it means that the i-th residue is a filling or invalid residue.

[0076] First, the affinity prediction model also includes matrix operation units, masking units, and normalization units. The matrix operation units are used to expand the one-dimensional target sequence effective mask matrix into a two-dimensional residue pair effective mask matrix. ,Right now ,in, This represents the mask value at the position of the i-th residue. This represents the mask value at the position of the j-th residue. This is true if and only if both residue i and residue j are valid residues. If residue i or residue j is a filling or invalid residue, then .

[0077] Then, the residue pairs are effectively masked using masking units. Contact map matrix with target Multiplying them together yields the masked target contact map, i.e. The effective neighbor weights of each residue are summed row by row in the masked target contact map. ,Right now For the binary contact map (distance feature matrix) O, It is approximately equal to the number of effective contact neighbors of residue i.

[0078] Finally, the masked target contact map is normalized row-wise using normalization units to obtain the first-order adjacency weight matrix. ,in, .

[0079] Step S3013: The residue embedding projection matrix is ​​weighted and aggregated using a first-order adjacency weight matrix to obtain the one-hop neighborhood context and two-hop neighborhood context of each residue; based on the residue embedding projection matrix, one-hop neighborhood context and two-hop neighborhood context, the multi-scale fused residue features are determined.

[0080] Specifically, a first-order adjacency weight matrix is ​​used to perform weighted aggregation of the residue embedding projection matrix to obtain the one-hop neighborhood context of each residue. Then, a multi-scale contact propagation step is performed to obtain the two-hop neighborhood context. ,in, ; Residues are embedded into the projection matrix in the feature dimension. , , The spliced ​​matrix is ​​then fused using a multi-layer perceptron (MLP) and residual connections are performed to obtain the final result. , The semicolon ";" indicates concatenation based on feature dimensions. Finally, the target sequence is effectively masked. and Multiplication yields multi-scale fusion residue features , .

[0081] Step S303: Perform topological encoding on the target contact map after masking to obtain topological embedding, and use the topological embedding and peptide embedding projection vector to perform weighted update of multi-scale fusion residue features to determine the site-weighted residue features.

[0082] The following reference Figure 5 This section will introduce the process of determining the residue characteristics after site weighting.

[0083] Figure 5 A flowchart illustrating the steps for determining site-weighted residue features according to an embodiment of this application is shown, as follows: Figure 5 As shown, the steps for determining the site-weighted residue features include: Step S3014: Sum each residue row by row on the target contact map after masking to obtain the contact degree of the residue, and determine the local clustering coefficient corresponding to the residue based on the contact degree.

[0084] On the masked target contact map, the contact degree of the i-th residue in the target amino acid sequence is obtained by summing the residues row by row. , .

[0085] Let the maximum contact degree among all effective residues of the target be . Then the normalized contact degree is: Local clustering coefficient The calculation formula is: ,in, ; This is the target contact map matrix after masking; ² = Indicates two Matrix multiplication; ⊙ represents element-wise multiplication.

[0086] Step S3015: The normalized contact degree and local clustering coefficient are spliced ​​and linearly projected to obtain the topological embedding of the residue.

[0087] Specifically, the topological embedding of the i-th residue is denoted as: , ,in, and This represents the learnable projection matrix.

[0088] Step S3016 involves splicing and multi-layer sensing of multi-scale fused residue features, one-hop neighborhood context, topological embedding, and peptide embedding projection vectors to obtain effective site scores.

[0089] For each residue i in the target amino acid sequence, multi-scale fusion residue features will be used. One-hop neighborhood context Topology embedding and peptide embedding projection vector By concatenating them, we obtain the concatenated vector. , .

[0090] Will Input multilayer perceptron The effective site scores are then obtained by processing the site with the sigmoid activation function. , ,in, σ represents the activation function sigmoid.

[0091] Step S3017: Use effective site scores to weight the multi-scale fusion residue features to determine the site-weighted residue features.

[0092] Specifically, the site-weighted residue features are denoted as: , .

[0093] Step S304: Using multi-head attention mechanism, site-weighted pooling and multilayer perceptron, the site-weighted residue features and peptide embedding projection vector are fused to determine the binding affinity prediction results.

[0094] The following reference Figure 6 Let's introduce the affinity prediction process.

[0095] Figure 6 A flowchart of the affinity prediction steps provided in the embodiments of this application is shown, as follows: Figure 6 As shown, the affinity prediction steps include: Step S3018: Based on the site-weighted residue features and peptide embedding projection vector, determine the peptide polymerization characterization.

[0096] Using peptide embedding projection vectors The query vector is based on site-reweighted residue features. Determine the query index Key and the total information Value to be extracted carried by the i-th residue, i.e. , , ,in, , , It is a learnable projection matrix.

[0097] The affinity prediction model incorporates a multi-head attention mechanism, which can calculate the unnormalized attention score of the h-th attention head for the i-th residue. , ,in, Let h be the feature dimension of the h-th attention head; Let the vector of the Key of the i-th residue in the h-th attention head be . Let be the vector of peptide Query in the h-th attention head, and β be a learnable scalar. Then, for each residue i, calculate the attention scores of all attention heads. Summing is performed to obtain the original total score for residue i. The original total score of all residues i in the entire target amino acid sequence. Softmax normalization is performed to obtain normalized attention weights. Aggregating the information to be extracted (Value) according to the normalized attention weights yields: ,in, Let i be the value vector of residue i.

[0098] Step S3019: Perform layer normalization on the superposition vector corresponding to the polypeptide polymerization characterization and polypeptide embedding projection vector, and input it into the feedforward neural network to obtain the polypeptide embedding fusion vector.

[0099] The superposition vector corresponding to the peptide polymerization characterization and the peptide embedding projection vector can be represented as: By performing layer normalization on the superposition vector, the peptide embedding fusion vector before updating can be obtained. , ,in, Representation layer normalization, It is a learnable projection matrix.

[0100] The peptide vector output by the attention branch By inputting the data into the feedforward neural network, the updated peptide embedding fusion vector can be obtained. .

[0101] Step S3020 involves adding type embedding to the peptide embedding fusion vector and the site-reweighted residue features to obtain peptide characterization and target characterization.

[0102] After incorporating type embeddings into the peptide embedding fusion vector and the site-reweighted residue features, they are spliced ​​into a joint sequence T of length 1+M to obtain peptide characterizations. and target characterization ,in, , .

[0103] , This indicates a learnable type embedding used to distinguish between peptide tokens and residue tokens.

[0104] Step S3021: After polymerizing the target characterization, splice it with the peptide characterization, and fit the splicing result through a multilayer perceptron to obtain the binding affinity prediction result.

[0105] Using effective site scoring and the effective mask matrix of the target sequence After determining and normalizing the pooling weights, we get: ; In the above formula, ε is a positive number to prevent division by zero.

[0106] Then, the target representation is aggregated according to the weights. Obtain target polymerization characterization r, Characterizing peptides The binding affinity prediction results are obtained by splicing the target aggregation characterization r with the target aggregation characterization r and then fitting the data through a multilayer perceptron. , .

[0107] Step S105: Sort multiple second-batch candidate molecules according to the binding affinity prediction results in order to screen out GLP-1R agonist molecules that meet the affinity requirements.

[0108] According to the predicted binding affinity, multiple second-batch candidate molecules are sorted in descending order, and a predetermined number of the top-ranked second-batch candidate molecules are selected as GLP-1R agonist molecules, in order to retain molecules with high binding affinity.

[0109] The glucagon-like peptide-1 receptor agonist screening method provided in this application has the following technical effects: First, the multi-source characterization of peptides and targets is more comprehensive, overcoming the limitations of traditional small molecule screening. Existing technologies rely solely on Lipinski and Veber artificial molecular descriptors and single-molecule docking scoring to characterize molecules, failing to capture the long sequence context and long-distance residue dependencies of peptides. However, this application employs ChemBERTa+PepDoRA to extract global chemical embeddings of peptides and ESM2 Long to generate target residue-level sequence embeddings, automatically learning the peptide amino acid arrangement and implicit information of protein residue context. This eliminates the need for manually designed screening features and is suitable for long-chain peptide GLP. The characterization of 1R agonists requires more complete feature expression and stronger generalization ability.

[0110] Second, complete alignment and utilization of protein three-dimensional structure information eliminates prediction interference caused by structural defects. Multiple experimental PDB structures are associated using UniProt. When no measured structure is available, AlphaFold predicted structures are used to supplement the data. Furthermore, global sequence-structure alignment is used to backfill the Cα three-dimensional coordinates into the target amino acid sequence index, unifying the numbering system of all residues and avoiding the loss of topological information caused by structural truncation or missing residues.

[0111] Third, the multi-module fusion peptide-target affinity model significantly improves the accuracy of interactive modeling. Based on the global features of the peptide, the binding potential of each residue is predicted. The key residue features are amplified by site score weighting. Cross-attention uses the peptide as the query to focus on residues with high binding potential. The site score is additionally introduced as an attention bias. Compared with the ordinary attention mechanism, this model is closer to the actual binding mode of peptide and receptor, and greatly improves the accuracy of affinity prediction.

[0112] By integrating first- and second-order neighborhood spatial features of residues through multi-scale contact propagation, combined with topological embedding encoding of local structural compactness, and further combined with residual connections, layer normalization, and FFN nonlinear transformation, the high-order interaction features of peptides and proteins are fully explored. Finally, the deep fusion of peptide and target features is achieved through dual-path joint encoding. Compared with the single scoring function of traditional molecular docking, the prediction results are more stable and are not affected by the initial conformation of peptides or pocket fitting bias.

[0113] Fourth, the hierarchical and cascaded virtual screening process significantly reduces wet laboratory costs and improves screening efficiency. The first level uses RDKit physicochemical descriptors to directly eliminate molecules whose molecular weight, TPSA, and hydrogen bond number do not meet the standards for peptide drug development. The second level uses a peptide aggregation prediction model to pre-screen peptides with high aggregation risk, preventing easily aggregated molecules from entering SPR and cell pharmacodynamic experiments, thus avoiding data distortion and sample waste. Only physicochemically qualified candidate peptides with low aggregation risk are sent to a high-precision affinity model for sorting, greatly reducing the size of the candidate set for later experiments.

[0114] Fifth, the fully automated virtual screening process replaces the traditional "full-volume docking and batch SPR detection" model. It eliminates the need to prepare samples from massive amounts of peptides individually for in vitro testing. Relying on AI models, it completes the triple evaluation of physicochemical properties, aggregation, and affinity in batches, significantly improving screening throughput and shortening the GLP (Good Laboratory Practice) timeline. The development and iteration cycle of 1R agonist candidate molecules is shortened, saving on consumables and manpower costs such as SPR chips, peptide raw materials, and cell experiments.

[0115] Based on the same inventive concept, this application also provides a glucagon-like peptide-1 receptor agonist screening device corresponding to the glucagon-like peptide-1 receptor agonist screening method. Since the principle of the device in this application is similar to the glucagon-like peptide-1 receptor agonist screening method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0116] Please see Figure 7 , Figure 7 This is a schematic diagram of a glucagon-like peptide-1 receptor agonist screening device provided in an embodiment of this application. Figure 7 As shown, the glucagon-like peptide-1 receptor agonist screening device 400 includes: Model building module 401 is used to build peptide aggregation prediction models and affinity prediction models; The first batch selection module 402 is used to determine the molecular descriptor corresponding to the peptide candidate to be predicted, which is used to describe the physicochemical properties of the drug-like substance. Based on the molecular descriptor, multiple first batch candidate molecules that meet the physicochemical property requirements are screened from the peptide candidate to be predicted. The second batch selection module 403 is used to input the SMILES sequence of each first batch candidate molecule into the peptide aggregation prediction model, determine the aggregation of each first batch candidate molecule, and use the aggregation to filter the first batch candidate molecules for high aggregation molecules to obtain multiple second batch candidate molecules. The affinity prediction module 404 is used to input the SMILES sequence, GLP-1R protein sequence and GLP-1R protein structure information of each second batch of candidate molecules into the affinity prediction model to determine the binding affinity prediction result between each second batch of candidate molecules and GLP-1R protein; The agonist determination module 405 is used to sort multiple second-batch candidate molecules according to the binding affinity prediction results in order to screen out GLP-1R agonist molecules that meet the affinity requirements.

[0117] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 8 As shown, the electronic device 500 includes a processor 510, a memory 520, and a bus 530.

[0118] The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 is running, the processor 510 and the memory 520 communicate via the bus 530. When the machine-readable instructions are executed by the processor 510, they can perform the operations described above. Figure 1 The specific implementation of the glucagon-like peptide-1 receptor agonist screening method in the illustrated method embodiment can be found in the method embodiment, and will not be repeated here.

[0119] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described actions. Figure 1 The specific implementation of the glucagon-like peptide-1 receptor agonist screening method in the illustrated method embodiment can be found in the method embodiment, and will not be repeated here.

[0120] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0121] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0122] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0123] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0124] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0125] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for screening glucagon-like peptide-1 receptor agonists, characterized in that, include: Construct peptide aggregation prediction models and affinity prediction models; Determine the molecular descriptor corresponding to the candidate peptide to be predicted for describing the physicochemical properties of the drug, and screen multiple first-batch candidate molecules that meet the physicochemical property requirements from the candidate peptide to be predicted based on the molecular descriptor. The SMILES sequence of each first batch of candidate molecules is input into the peptide aggregation prediction model to determine the aggregation of each first batch of candidate molecules. The aggregation is then used to filter the first batch of candidate molecules for high aggregation molecules to obtain multiple second batch of candidate molecules. The SMILES sequence, GLP-1R protein sequence, and GLP-1R protein structure information of each second batch of candidate molecules are input into the affinity prediction model to determine the predicted binding affinity between each second batch of candidate molecules and the GLP-1R protein. The multiple second-batch candidate molecules were sorted according to the binding affinity prediction results to screen out GLP-1R agonist molecules that meet the affinity requirements.

2. The method according to claim 1, characterized in that, The steps for inputting the SMILES sequence, GLP-1R protein sequence, and GLP-1R protein structure information of each second batch of candidate molecules into the affinity prediction model to determine the predicted binding affinity between each second batch of candidate molecules and the GLP-1R protein include: Determine the target peptide embedding vector corresponding to the SMILES sequence of each second batch of candidate molecules, the target residue embedding matrix corresponding to the GLP-1R protein sequence, and the target contact map matrix corresponding to the GLP-1R protein structural information; The target contact map matrix is ​​masked and normalized to obtain a first-order adjacency weight matrix. The first-order adjacency weight matrix is ​​then used to perform multi-scale contact propagation on the residue embedding projection matrix to obtain multi-scale fused residue features. Topological encoding is performed on the target contact map after masking to obtain topological embedding. The multi-scale fused residue features are then updated by weighting the topological embedding and peptide embedding projection vector to determine the site-weighted residue features. The peptide embedding projection vector is obtained based on the linear projection of the target peptide embedding vector. Using multi-head attention mechanism, site-weighted pooling, and multilayer perceptron, the site-weighted residue features and the polypeptide embedding projection vector are fused to determine the binding affinity prediction result.

3. The method according to claim 2, characterized in that, The steps for fusing the site-weighted residue features and the peptide embedding projection vector using multi-head attention, site-weighted pooling, and multilayer perceptron to determine the binding affinity prediction results include: Based on the site-weighted residue features and the polypeptide embedding projection vector, the polypeptide polymerization characterization is determined; The superposition vector corresponding to the polypeptide polymerization characterization and the polypeptide embedding projection vector is subjected to layer normalization and input into a feedforward neural network to obtain the polypeptide embedding fusion vector. Type embedding is added to the peptide embedding fusion vector and the site-reweighted residue features to obtain peptide characterization and target characterization; After the target characterization is polymerized, it is spliced ​​with the peptide characterization. The splicing result is then fitted by a multilayer perceptron to obtain the binding affinity prediction result.

4. The method according to claim 2, characterized in that, The steps of masking and normalizing the target contact map matrix to obtain a first-order adjacency weight matrix, and using the first-order adjacency weight matrix to perform multi-scale contact propagation on the residue embedding projection matrix to obtain multi-scale fused residue features include: The target contact map matrix is ​​masked to obtain the masked target contact map. The target contact map after masking is subjected to row-wise summation and normalization to obtain a first-order adjacency weight matrix; The residue embedding projection matrix is ​​weighted and aggregated using the first-order adjacency weight matrix to obtain the one-hop neighborhood context and two-hop neighborhood context of each residue. Based on the residue embedding projection matrix, the one-hop neighborhood context, and the two-hop neighborhood context, multi-scale fused residue features are determined, wherein the residue embedding projection matrix is ​​obtained based on the linear projection of the target residue embedding matrix.

5. The method according to claim 4, characterized in that, The steps of performing topological encoding on the masked target contact map to obtain topological embedding, and using the topological embedding and peptide embedding projection vector to weight and update the multi-scale fused residue features to determine the site-weighted residue features include: The contact degree of each residue is obtained by summing the rows on the target contact map after masking, and the local clustering coefficient corresponding to the residue is determined based on the contact degree. The topological embedding of the residue is obtained by splicing and linearly projecting the contact degree and the local clustering coefficient. The multi-scale fused residue features, the one-hop neighborhood context, the topological embedding, and the peptide embedding projection vector are spliced ​​together and subjected to multi-layer sensing to obtain an effective site score. The effective site score is used to weight the multi-scale fusion residue features to determine the site-weighted residue features.

6. The method according to claim 1, characterized in that, The affinity prediction model is constructed using the following processing: The peptide entities and target entities in the sample table are deduplicated to construct peptide sets and target sets; Based on the set of peptides and the set of targets, determine the peptide embedding vector, the target residue embedding matrix, and the target contact map matrix; The affinity prediction model is trained using the peptide embedding vector, the target residue embedding matrix, and the target contact map matrix to obtain the trained affinity prediction model.

7. The method according to claim 6, characterized in that, The target contact map matrix is ​​determined through the following processing: For each target amino acid sequence in the target set, based on the length of the longest target sequence in the sample set, a corresponding effective mask matrix for the target sequence is constructed according to whether there is a residue at each residue index position; The experimental structure of each target entity was obtained and parsed, and the residue spatial metadata of each standard amino acid residue was extracted to construct a chain-level residue recording sequence. Each target amino acid sequence is compared with the corresponding chain-level residue recording sequence to determine whether each residue position is successfully paired. When the residue pairing is successful, the residue spatial metadata is backfilled into the corresponding residue position of the target amino acid sequence to obtain a three-dimensional coordinate matrix and an effective residue mask. The residue spatial metadata includes carbon atom coordinates. Based on the carbon atom coordinates of the effectively mapped residues, the distance matrix of residue pairs corresponding to each target amino acid sequence is determined; The residue pair distances in the residue pair distance matrix are compared with a set distance threshold, and the residue pair distance matrix is ​​converted into a target contact map matrix based on the comparison results.

8. The method according to claim 6, characterized in that, The peptide embedding vector is determined through the following process: The SMILES sequence of each polypeptide entity in the polypeptide set is input into the polypeptide adaptation chemistry model to determine the polypeptide embedding vector. The polypeptide adaptation chemistry model is a model based on a molecular language model and superimposed with a polypeptide lightweight fine-tuning module.

9. The method according to claim 6, characterized in that, The target residue embedding matrix was determined through the following process: The amino acid sequence of each target in the target set is encoded using a protein language model to obtain a target residue embedding matrix.

10. The method according to claim 6, characterized in that, A peptide aggregation prediction model was constructed using the following processing: A dataset is constructed using the SMILES sequence of each peptide in the peptide set as the input column and whether or not it is clustered as the label column. The molecular property prediction model was trained using the dataset to obtain a polypeptide aggregation prediction model.

Citation Information

Patent Citations

  • Screening method of glucagon-like peptide-1 receptor stimulant and application of screened compound

    CN120847037A