Method, device, equipment and storage medium for cyclopeptide design

By initializing the chemical map and atomic representation of cyclic peptides, the structural and sequence prediction models are updated, which solves the diversified needs and design quality problems of traditional cyclic peptide design methods, and realizes diversified cyclic peptide design and high-quality chemical bond connections.

CN120496677APending Publication Date: 2025-08-15BYTEDANCE TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510571295.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional cyclic peptide design methods cannot meet the diverse design needs and cannot perform atomic modeling. The design quality needs to be improved and the cost is high.

Method used

By initializing the chemical map, atomic representation and sequence information of the cyclic peptide, the structure prediction model and sequence prediction model are updated, and combined with the diffusion model for denoising, the structure or amino acid sequence of the cyclic peptide is determined.

Benefits of technology

Various types of cyclic peptide designs have been realized, allowing the introduction of non-standard amino acids, improving the accuracy of chemical bond connections and the matching degree of cyclic peptides with receptors, and improving the design quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496677A_ABST
    Figure CN120496677A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method, device and equipment for cyclopeptide design and a storage medium. The method comprises the following steps: initializing a first chemical diagram, atomic representation and sequence information of a cyclic peptide to be designed; updating the initialized atomic representation with a structure prediction model based on the initialized first chemical map to obtain an updated atomic representation; updating the initialized sequence information using a sequence prediction model based on the initialized first chemical map and the updated atomic representation to obtain updated sequence information; and determining at least one of a structure or an amino acid sequence of the cyclic peptide based at least on the updated atomic representation and the updated sequence information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, computer-readable storage media, and computer program products for cyclic peptide design. Background Art

[0002] Peptide drugs are a unique class of pharmaceutical compounds composed of precisely arranged amino acid sequences. Compared to small molecule compounds, peptide drugs generally have lower toxicity, stronger biological activity, higher target specificity, and higher cell permeability. However, linear peptide drugs have problems such as short half-life, limited stability, and susceptibility to degradation by hydrolases. Compared to linear peptides, cyclic peptides typically have one or more closed ring structures, which enhance their resistance to digestive enzymes. In addition, cyclic peptides generally have higher stability and affinity, allowing them to bind to receptors in a more stable conformation. Summary of the Invention

[0003] In the first aspect of the present disclosure, a method for cyclic peptide design is provided. The method includes: initializing a first chemical map, an atomic representation, and sequence information of a cyclic peptide to be designed, wherein the first chemical map indicates a plurality of atoms contained in the cyclic peptide and a bonding relationship between the plurality of atoms, wherein the plurality of atoms include a plurality of constrained atoms for cyclization connection and a plurality of free atoms other than the plurality of constrained atoms, the atomic representation indicates the spatial position of the plurality of atoms, and the sequence information indicates the amino acid sequence of the cyclic peptide; based on the initialized first chemical map, updating the initialized atomic representation using a structure prediction model to obtain an updated atomic representation; based on the initialized first chemical map and the updated atomic representation, updating the initialized sequence information using a sequence prediction model to obtain an updated sequence information; and determining at least one of the structure or amino acid sequence of the cyclic peptide based at least on the updated atomic representation and the updated sequence information.

[0004] In the second aspect of the present disclosure, a device for cyclic peptide design is provided. The device includes: a cyclic peptide initialization module, which is configured to initialize the first chemical diagram, atomic representation and sequence information of the cyclic peptide to be designed, the first chemical diagram indicating the bonding relationship between the multiple atoms contained in the cyclic peptide and the multiple atoms, the multiple atoms including multiple constrained atoms for cyclization connection and multiple free atoms other than the multiple constrained atoms, the atomic representation indicating the spatial position of the multiple atoms, and the sequence information indicating the amino acid sequence of the cyclic peptide; an atomic representation update module, which is configured to update the initialized atomic representation based on the initialized first chemical diagram using a structure prediction model to obtain an updated atomic representation; a sequence information update module, which is configured to update the initialized sequence information based on the initialized first chemical diagram and the updated atomic representation using a sequence prediction model to obtain updated sequence information; and a cyclic peptide determination module, which is configured to determine at least one of the structure or amino acid sequence of the cyclic peptide based at least on the updated atomic representation and the updated sequence information.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored on the computer-readable storage medium, and the computer-executable instructions can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided, comprising computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram illustrating an example environment in which embodiments according to the present disclosure may be implemented;

[0011] Figure 2shows a schematic diagram of an example architecture for cyclic peptide design according to some embodiments of the present disclosure;

[0012] Figures 3A to 3E Schematic diagrams of example scenarios for cyclic peptide design according to some embodiments of the present disclosure are respectively shown;

[0013] Figure 4 A schematic diagram illustrating an example of an atom representation according to some embodiments of the present disclosure;

[0014] Figure 5 shows a flow chart of a process for cyclic peptide design according to some embodiments of the present disclosure;

[0015] Figure 6 A schematic structural block diagram showing an example apparatus for cyclic peptide design according to some embodiments of the present disclosure; and

[0016] Figure 7 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0017] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.

[0019] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0020] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0021] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0022] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium and other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.

[0023] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0025] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0026] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs. It typically includes an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the previous layer.

[0027] Generally speaking, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also known as the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also known as input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values obtained through training to determine the corresponding output.

[0028] As mentioned above, peptide drugs are a unique class of pharmaceutical compounds composed of precisely arranged amino acid sequences. Compared with small molecule compounds, peptide drugs generally have lower toxicity, stronger biological activity, higher target specificity, and higher cell permeability. However, linear peptide drugs have problems such as short half-life, limited stability, and susceptibility to degradation by hydrolases. Compared with linear peptides, cyclic peptides usually have one or more closed ring structures, which enhance the resistance of cyclic peptides to digestive enzymes. In addition, cyclic peptides generally have higher stability and affinity, allowing cyclic peptides to bind to receptors in a more stable conformation.

[0029] Traditional cyclic peptide drug discovery relies primarily on chemical synthesis and high-throughput screening, both of which are labor-intensive and costly. With advances in machine learning modeling, some R&D institutions are leveraging these models to design cyclic peptide drugs. However, traditional cyclic peptide design methods are typically limited to designing certain specific types of cyclic peptides, failing to meet diverse design requirements. Furthermore, traditional cyclic peptide design methods are unable to model cyclic peptides at the atomic level, leaving room for improvement in the rationality and accuracy of the designed cyclic peptides, and the quality of cyclic peptide design remains to be enhanced.

[0030] In view of this, an embodiment of the present disclosure proposes an improved scheme for cyclic peptide design. In this scheme, the first chemical map, atomic representation and sequence information of the cyclic peptide to be designed are initialized. The first chemical map indicates the bonding relationship between the multiple atoms contained in the cyclic peptide and the multiple atoms, and the multiple atoms include multiple constrained atoms for cyclization connection and multiple free atoms other than the multiple constrained atoms. The atomic representation indicates the spatial position of the multiple atoms, and the sequence information indicates the amino acid sequence of the cyclic peptide. Based on the initialized first chemical map, the initialized atomic representation is updated using a structure prediction model to obtain an updated atomic representation. Based on the initialized first chemical map and the updated atomic representation, the initialized sequence information is updated using a sequence prediction model to obtain updated sequence information. Thereafter, at least one of the structure or amino acid sequence of the cyclic peptide is determined based at least on the updated atomic representation and the updated sequence information.

[0031] In the embodiments disclosed herein, the atoms of the cyclic peptide are divided into constrained atoms constrained by the cyclization structure of the cyclic peptide and free atoms not constrained by the cyclization structure. The cyclic peptide is modeled at the full atomic level and chemical bond scale using a chemical diagram, which can model the geometric constraints of the cyclization structure and improve the correctness of the chemical bond connection. The spatial position of the atoms and the amino acid sequence are predicted respectively by the structure prediction model and the sequence prediction model, which can be applied to the design of various types of cyclic peptides and allows the introduction of non-standard amino acids, which can meet the diverse needs of cyclic peptide design.

[0032] Various example implementations of this solution are described in detail below in conjunction with the accompanying drawings.

[0033] Sample Environment

[0034] Figure 1 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 In the environment 100 of FIG. 1 , the design system 110 can utilize one or more machine learning models 105 to perform structural design or structural prediction of the peptide 125 . The machine learning model 105 is configured to determine at least one of the structure or amino acid sequence of the peptide 125 based on input information 115 .

[0035] In some embodiments, peptide 125 may include an oligopeptide, a polypeptide, a cyclic peptide, and the like. Peptide 125 may also include a protein composed of one or more peptide chains. In some examples, design system 110 may also be referred to as a peptide design system or a protein design system. In some embodiments, the amino acid sequence of peptide 125 may include the amino acid sequence of one or more peptide chains of peptide 125, and the amino acid sequence may indicate the residue type of each residue (i.e., amino acid residue) in the peptide chain.

[0036] In some embodiments, the structure of peptide 125 can indicate at least one of the main chain conformation or side chain conformation of peptide 125, the main chain conformation can indicate the spatial structure of multiple atoms of the main chain (Backbone) of peptide 125, and the side chain conformation can indicate the spatial structure of multiple atoms of the side chain (Side Chain) of peptide 125. In some examples, the design system 110 can determine the all-atom structure of peptide 125, and the all-atom structure can indicate the spatial structure of multiple atoms of the main chain and side chain of peptide 125. In some examples, the spatial structure of multiple atoms can be indicated by the spatial coordinates of multiple atoms. In some examples, the design system 110 can also regard the atoms in peptide 125 as nodes, the bonds between two atoms as edges, and represent the spatial structure of multiple atoms by the torsion angles between multiple edges. Of course, the above-mentioned method of representing the spatial structure is only exemplary, and any other appropriate method can be selected according to actual needs to represent the spatial structure of multiple atoms. The embodiments of the present disclosure are not limited to this.

[0037] In some embodiments, the peptide 125 can be a cyclic peptide, which can be used as a ligand to bind to a specific receptor. In this case, the input information 125 can include a receptor structure representation of the receptor, the cyclization type of the cyclic peptide, or the number of residues of the cyclic peptide, etc. In some embodiments, the input information 125 can also include at least part of the sequence information, at least part of the main chain information, at least part of the side chain information, etc. of the peptide 125. The sequence information can indicate the amino acid sequence of the peptide chain, the main chain information can indicate the main chain construction of the peptide chain, and the side chain information can indicate the side chain conformation of the peptide chain. Of course, the above-mentioned input information 125 is only exemplary. In actual application scenarios, various information related to the peptide 125 to be designed or predicted can be selected as input information 125 according to actual needs. The embodiments of the present disclosure are not limited to this.

[0038] In some embodiments, the design system 110 can be implemented on any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices can include any type of mobile, fixed, or portable terminal, including mobile phones, desktop computers, laptops, netbooks, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals thereof, or any combination thereof. Servers include, but are not limited to, mainframe computers, edge computing nodes, computing devices in cloud environments, etc.

[0039] The machine learning model 105 can be a model of different types. In some embodiments, one or more machine learning models 105 may include a generative model. In some embodiments, one or more machine learning models 105 may be constructed based on a diffusion model. The diffusion model, also known as a diffusion probability model, is a type of generative model. The model generates data by simulating a diffusion process. This process is inspired by physical processes such as heat diffusion. The diffusion model includes a forward diffusion process and a reverse diffusion process. The diffusion model generates new data samples by simulating a forward diffusion process in which noise is gradually added, and then learning how to reverse this process.

[0040] During forward diffusion, noise is gradually added to the data, making it increasingly random over a series of steps until the data resembles pure noise. This process can be thought of as a Markov chain, where Gaussian noise is added to the data at each step. The forward diffusion process can be expressed as: where x t is the noise data of step t, α t Used to control the amount of noise added. The forward diffusion process is performed during model training, and the data used to add noise are the training samples.

[0041] In the reverse diffusion process (also known as reverse denoising), the model learns to reverse the steps of adding noise. Starting with pure noise, the diffusion model gradually removes the noise to generate data that matches the training distribution. The reverse diffusion process is usually simulated using a neural network that predicts the noise added at each step: where u θ and σ θ are the learned model parameters. After the model training is completed, the model performing the back-diffusion process can first start sampling from the noise distribution and use the model to iteratively denoise until the desired data is obtained.

[0042] In diffusion models, the time step refers to the number of noise addition steps during the forward diffusion process. The total number of steps, T, is typically a preset value, representing the number of steps required to transform the original data into pure noise. At each time step, t, Gaussian noise is added to the data according to a predetermined noise scheme. This process is continuous, and each step depends on the results of the previous step.

[0043] When generating data, the diffusion model inference step (Inference Step) refers to the number of steps required to recover the original data from pure noise during the back diffusion process. The number of inference steps directly affects the quality and speed of generated data. Generally, the more inference steps there are, the higher the quality of the generated data is, but it also increases the computational cost and time. In practical applications, the generation quality and efficiency can be balanced by adjusting the number of inference steps. In some embodiments, the inference step corresponds to the time step, and each inference step can correspond to one or more time steps. For example, if the total time step of the diffusion model is 1000 steps and the inference step is set to 50 steps, then each inference step can correspond to 20 time steps.

[0044] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0045] Example Scenario

[0046] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings. Figure 2 FIG2 is a schematic diagram of an example architecture 200 for cyclic peptide design according to an embodiment of the present disclosure. In the following, for ease of discussion, the example architecture 200 is described from the perspective of the design system 110 , which is merely exemplary.

[0047] In some embodiments of the present disclosure, the design system 110 initializes the atomic representation 212, the first chemical graph 214, and the sequence information 216 of the cyclic peptide to be designed. The first chemical graph 214 indicates the bonding relationship between the multiple atoms contained in the cyclic peptide and the multiple atoms. In some embodiments, the first chemical graph 214 may include multiple nodes corresponding to the multiple atoms contained in the cyclic peptide, and multiple edges corresponding to the multiple chemical bonds between the multiple atoms. The atoms in the cyclic peptide are indicated by the multiple nodes, and the bonding relationship between the multiple atoms is indicated by the multiple edges. In some examples, the first chemical graph 214 may indicate the bonding relationship between the multiple atoms other than hydrogen atoms in the cyclic peptide and the multiple atoms.

[0048] As an example, Figure 3A Schematic diagram of an example scenario 300A for cyclic peptide design according to some embodiments of the present disclosure is shown. Figure 3A As shown, cyclic peptide 320 includes residues 321, 322, 323, 324, and 325. The first chemical graph 214 may include a plurality of nodes corresponding to a plurality of atoms other than hydrogen atoms (e.g., carbon atoms, oxygen atoms, nitrogen atoms, and sulfur atoms, etc.) in residues 321, 322, 323, 324, and 325, and a plurality of edges corresponding to a plurality of chemical bonds between the plurality of atoms.

[0049] The multiple atoms in the cyclic peptide include multiple constrained atoms and multiple free atoms in addition to multiple constrained atoms that are connected for cyclization. Specifically, the cyclic peptide can include two residues connected by cyclization and at least one residue in addition to the two residues connected by cyclization. Cyclization reaction occurs between the two residues connected by cyclization to form a cyclized structure (i.e., the structure of two reactive groups that undergo cyclization). Because the position and bonding relationship of the multiple atoms in the two residues connected by cyclization are subject to the constraint and restriction of the cyclization structure, the two residues connected by cyclization can be referred to as two constrained residues, and the multiple atoms in the two constrained residues can be referred to as multiple constrained atoms. The position and bonding relationship of the multiple atoms in at least one residue in addition to two constrained atoms are not subject to the cyclization structure constraint and restriction, so this at least one residue can be referred to as at least one free residue, and the multiple atoms in this at least one free residue can be referred to as multiple free atoms.

[0050] In some examples, the plurality of constrained atoms may include a plurality of atoms other than hydrogen atoms in the two constrained residues, and the plurality of free atoms may include a plurality of atoms other than hydrogen atoms in the at least one free residue. Figure 3A As shown, cyclic peptide 320 is connected by a carbon-sulfur bond (CS bond) cyclization between residue 322 and residue 325. In this case, residue 322 and residue 325 can be referred to as two constrained residues, and multiple atoms other than hydrogen atoms in residue 322 and residue 325 can be referred to as multiple constrained atoms. Residue 321, residue 323 and residue 324 can be referred to as multiple free residues, and multiple atoms other than hydrogen atoms in residue 321, residue 323 and residue 324 can be referred to as free atoms. It will be understood that constrained atoms can include all or part of the atoms of the constrained residues, and free atoms can also include all or part of the atoms of the free residues. The embodiments of the present disclosure are not limited to this.

[0051] The atomic representation indicates the spatial position of multiple atoms. In some embodiments, the atomic representation can indicate the spatial coordinates of multiple atoms contained in the cyclic peptide. In some examples, the atomic representation can include coordinate information of each residue in the multiple residues contained in the cyclic peptide. Each item of coordinate information includes main chain coordinate information and multiple side chain coordinate information, and the main chain coordinate information can include multiple spatial coordinates corresponding to multiple atoms in the main chain. The multiple side chain coordinate information corresponds to multiple predetermined residue types (e.g., 20 standard amino acids), respectively. Each item of side chain coordinate information can indicate the spatial coordinates of multiple atoms in the side chain of the residue when the corresponding residue belongs to the corresponding predetermined residue type.

[0052] As an example, Figure 4 A schematic diagram illustrating an example 400 of an atomic representation according to some embodiments of the present disclosure is shown. Figure 3A and Figure 4 As shown, the atomic representation of cyclic peptide 320 may include multiple pieces of coordinate information 410, 420, 430, 440, and 450 corresponding to residues 321, 322, 323, 324, and 325. Coordinate information 410 for residue 321 includes main chain coordinate information 411 and multiple pieces of side chain coordinate information 412, 413, and so on. Side chain coordinate information 412 indicates the spatial coordinates of multiple atoms included in the side chain of residue 321, when residue 321 is glycine. The contents of coordinate information 420, 430, 440, and 450 are similar to those of coordinate information 410 and are not described here one by one.

[0053] In some examples, the atomic representation may include coordinate information of each residue in the plurality of residues included in the cyclic peptide. Each item of coordinate information may include main chain coordinate information and side chain coordinate information. The main chain coordinate information may include the spatial coordinates of the plurality of atoms included in the main chain of the corresponding residue. The side chain coordinate information may include the spatial coordinates of the plurality of atoms included in the side chain when the corresponding residue belongs to the initialized residue type. For example, in combination Figure 3A and Figure 4 As shown, if the initialization residue type of residue 321 is glycine, the coordinate information of residue 321 may include main chain coordinate information 411 and side chain coordinate information 412. Of course, the above atomic representation is only exemplary. In actual application scenarios, any other appropriate method can be selected to represent the spatial position of atoms according to actual needs.

[0054] The sequence information indicates the amino acid sequence of the cyclic peptide. In some examples, the sequence information may include an identifier sequence, which may include multiple identifiers corresponding to multiple residues in the cyclic peptide, each identifier indicating the residue type of the corresponding residue. As an example, in combination with Figure 3A As shown, the initialized residue types of residues 321, 322, 323, 324, and 325 can be serine, cysteine, glycine, alanine, and methionine, respectively. In this case, the sequence information can include, for example, "SerCys Gly Ala Met" or "SCGAM".

[0055] In some embodiments, combined Figure 2 As shown, the design system 110 can initialize the atomic representation 212 , the first chemical graph 214 , and the sequence information 216 of the cyclic peptide based on at least one of the receptor structure representation 202 of the receptor targeted by the cyclic peptide, the cyclization type 204 of the cyclic peptide, or the number of residues 206 of the cyclic peptide.

[0056] In some examples, the receptor structure representation 202 at least indicates the structure of the pocket region on the receptor that can be used for ligand binding. In actual application scenarios, the structure of the pocket region can be represented in a variety of ways. As an example, the receptor structure representation 202 can be represented by the atomic coordinates of the receptor, and the atomic coordinate representation can indicate the spatial coordinates of each atom contained in the receptor. As another example, the receptor structure representation 202 can also include a topological representation of the receptor, for example, the topological representation can include multiple nodes corresponding to multiple amino acids in the receptor and multiple edges corresponding to chemical bonds between multiple amino acids. As another example, Figure 3A As shown, the receptor structure representation 202 may include a pocket region 330 (i.e., Figure 3A The three-dimensional shape representation of the receptor structure 202 is a three-dimensional shape representation of the concave area in the pocket. The three-dimensional shape representation can indicate the surface shape of the pocket area. Of course, the above receptor structure representation 202 is only exemplary, and any other appropriate method can be selected to represent the receptor structure according to actual needs. The embodiments of the present disclosure are not limited to this.

[0057] In some examples, the cyclization type 204 of the cyclic peptide may include a head-to-tail cyclization type, a side-to-tail cyclization type, a head-to-side cyclization type, or a side-to-side cyclization type. The head-to-tail cyclization type indicates that the N-terminus (head) of an amino acid at one end of the peptide chain is bonded to the C-terminus (tail) of another amino acid at the other end of the peptide chain to form a cyclic structure. The side-to-tail cyclization type indicates that the side chain of an amino acid in the peptide chain is bonded to the C-terminus (tail) of the peptide chain to form a cyclic structure. The head-to-side cyclization structure indicates that the side chain of an amino acid in the peptide chain is bonded to the N-terminus (head) of the peptide chain to form a cyclic structure. The side-to-side cyclization type indicates that the side chain of an amino acid in the peptide chain is bonded to the side chain of another amino acid in the peptide chain to form a cyclic structure.

[0058] In some examples, the residue number 206 can be a given residue number. In other examples, the design system 110 can also determine the residue number (e.g., a range of residues) of the cyclic peptide based on a structural representation of the receptor. For example, the design system 110 can determine the residue number based on the size of the ligand binding pocket on the receptor. The cyclization type can indicate the cyclization method of the cyclic peptide.

[0059] In some examples, the design system 110 can initialize the atomic representation 212 based on the receptor structure representation 202, the cyclization type 204, and the number of residues 206. As an example, in combination Figure 4 As shown, if the number of residues is “5”, the design system 110 may initialize the atomic representation 212 including coordinate information 410 , 420 , 430 , 440 , and 450 based on the receptor structure representation 202 , the cyclization type 204 , and the number of residues 206 .

[0060] In some examples, the design system 110 may determine the residue types of the two constrained residues in the cyclic peptide based on the cyclization type 204. The design system 110 may initialize the residue type of at least one free residue in the cyclic peptide based on the residue number. Thereafter, the design system 110 may determine the initialized sequence information 216 based on the residue types of the two constrained residues and the residue type of the at least one free residue. As an example, in combination with Figure 3A As shown, if the cyclization type 204 is a side cyclization type, the design system 110 can determine two constrained residues whose side chains can be bonded to the tail end of the peptide chain, for example, cysteine and methionine can achieve side chain bonding to the tail end of the peptide chain through a C-S bond. If the number of residues is N, the design system 110 can initialize the residue types of N-2 free residues. The design system 110 can then determine the initialized sequence information 216, such as "Ser Cys Gly Ala Met" or "SCG AM", based on the residue types of the two constrained residues (i.e., cysteine and methionine) and N-2 free residues.

[0061] In some examples, the design system 110 may initialize the first chemical graph 214 of the cyclic peptide based on the initialized sequence information 216. As an example, the design system 110 may initialize the first chemical graph 214 of the cyclic peptide 320 based on the sequence information 216 including "Ser Cys Gly Ala Met". The first chemical graph 214 may include multiple nodes corresponding to C, N, O, and S atoms other than hydrogen atoms in the cyclic peptide 320, and multiple edges corresponding to chemical bonds between multiple atoms. In the case where the spatial position of the atoms is indicated by atomic representation, the first chemical graph 214 may not include the spatial coordinates of each atom. Of course, the above-mentioned first chemical graph 214 is merely exemplary, and the embodiments of the present disclosure do not limit the content and structure of the first chemical graph.

[0062] In some embodiments of the present disclosure, return to Figure 2As shown, the design system 110 updates the initialized atomic representation 212 using the structure prediction model 210 based on the initialized first chemical graph 214 to obtain an updated atomic representation 218. In some embodiments, the structure prediction model 210 can be constructed based on a diffusion model, and the structure prediction model 210 can be trained to be able to perform denoising on the atomic representation 212 based on the atomic representation 212 and the first chemical graph 214. In other words, the structure prediction model 210 can perform denoising on the atomic representation 212 based on the reverse diffusion process of the diffusion model to obtain an updated atomic representation 218 (also referred to as a denoised atomic representation). It should be noted that the above-mentioned structure prediction model 210 is only exemplary. In actual application, any appropriate model structure can be selected according to actual needs to construct the structure prediction model 210. The embodiment of the present disclosure does not limit the model structure of the structure prediction model 210.

[0063] As an example, combined with Figure 3A and Figure 3B As shown, the design system 110 may use the initialized atomic representation 212 as noise data for the t-th time step of the structure prediction model 210. The structure prediction model 210 may perform a denoising process on the atomic representation 212 based on the first chemical graph 214 to update the spatial coordinates of the plurality of atoms in the cyclic peptide 320. For example, Figure 3A The relative positions between the multiple atoms shown and the receptor 310 are updated as Figure 3B The relative positions of the multiple atoms and the receptor 310 are shown.

[0064] In some embodiments of the present disclosure, continue to combine Figure 2 As shown, the design system 110 updates the initialized sequence information 216 using the sequence prediction model 220 based on the initialized first chemical map 214 and the updated atomic representation 218 to obtain updated sequence information 220. In some embodiments, the sequence prediction model 220 can also be constructed using a diffusion model. The sequence prediction model 220 can be trained to predict the residue type of at least one free residue in the cyclic peptide based on the first chemical map 214, the atomic representation 218, and the sequence information 216.

[0065] In some embodiments, the design system 110 may remove atoms included in the side chain of at least one free residue from the initialized first chemical map 214 based on the initialized sequence information 216 and the cyclization type 204 to obtain a second chemical map. The design system 110 may generate a residue type prediction for the at least one free residue using the sequence prediction model 220 based on the second chemical map and the updated atomic representation 218. Thereafter, the design system 110 may determine updated sequence information 222 based on the predicted residue type of the at least one free residue and the residue types of the two constrained residues. Removing the side chain of the free residue from the first chemical map 214 can prevent the side chain of the free residue from affecting the residue type prediction of the free residue by the sequence prediction model 220, thereby improving the accuracy of the residue type prediction of the free residue.

[0066] As an example, combined with Figures 3B to 3D As shown, the first chemical map 214 can indicate multiple atoms other than hydrogen atoms and the bonding relationships between multiple atoms in the cyclic peptide 320. The design system 110 can remove the carbon atoms and oxygen atoms included in the side chain of the residue 321 from the first chemical map 214. The design system 110 can also remove the carbon atoms included in the side chain of the residue 324 from the first chemical map 214. Since the residue 323 is glycine, and the side chain of glycine has only one hydrogen atom, the design system 110 does not need to remove the atoms included in the side chain of the residue 323 from the first chemical map 214. In this way, the design system 110 can obtain an indication such as Figure 3C A second chemical diagram showing the bonding relationships between the multiple atoms shown in .

[0067] The design system 110 can provide the second chemical map and the atomic representation 218 to the sequence prediction model 220, and the sequence prediction model 220 can predict the residue types of the residue 321, the residue 323, and the residue 324. Figure 3D As shown, according to the prediction result of the sequence prediction model 220, the residue type of residue 321 is serine (Ser), the residue type of residue 323 is threonine (Thr), and the residue type of residue 324 is alanine (Ala). In this case, the design system 110 can update the sequence information 216 including the identification sequence "Ser Cys Gly Ala Met" to the sequence information 222 including the identification sequence "Ser Cys Thr AlaMet".

[0068] In some embodiments, the design system 110 may determine atomic feature representations for the plurality of free atoms using the sequence prediction model 220 based on the second chemical map and the updated atomic representation. The design system 110 may determine a residue feature representation for at least one free residue based on the atomic feature representations for the plurality of free atoms. The design system 110 may then use a decoder to perform feature decoding on the residue feature representation for the at least one free residue to generate a prediction of the residue type of the at least one free residue.

[0069] As an example, the decoder can be constructed based on a multi-layer perceptron (MLP). The sequence prediction model 220 can be configured to output atom potential representations of the plurality of free atoms in a latent space. The design system 110 can determine a residue potential representation of the at least one free residue based on the atom potential representations of the plurality of free atoms. Subsequently, the multi-layer perceptron can be used to predict the residue type of the at least one free residue based on the residue potential representation of the at least one free residue.

[0070] In some embodiments of the present disclosure, refer back to Figure 2 As shown, the design system 110 determines structural information 226 of the cyclic peptide based at least on the updated atomic representation 218 and the updated sequence information 222. The structural information 226 may indicate at least one of the structure or amino acid sequence of the cyclic peptide. The structure of the cyclic peptide indicates the multiple atoms contained in the cyclic peptide, the bonding relationships between the multiple atoms, and the spatial positional relationships between the multiple atoms. The amino acid sequence may indicate the residue types of the multiple residues (including free residues and constrained residues) contained in the cyclic peptide.

[0071] In some embodiments, continued binding Figure 2 As shown, the design system 110 can update the initialized first chemical map 214 based on the updated sequence information 222 to obtain an updated first chemical map 224. Thereafter, the design system 110 can determine the structural information 226 of the cyclic peptide based on the updated atomic representation 214, the updated first chemical map 224, and the updated sequence information 222. The updated first chemical map 224 can indicate the bonding relationships between the multiple atoms and the multiple atoms contained in the cyclic peptide. Moreover, even if non-standard amino acids are introduced into the cyclic peptide, the updated first chemical map 224 can accurately indicate the bonding relationships between the atoms of these non-standard amino acids, which is conducive to improving the accuracy of the chemical bond connections in the cyclic peptide, so that the cyclic peptide design process can tolerate the introduction of non-standard amino acids.

[0072] In some embodiments, the design system 110 can utilize the structure prediction model 210 and the sequence prediction model 220 to perform multiple iterations of iterative updates on the atomic representation 212, first chemical map 214, and sequence information 216 of the cyclic peptide. The design system 110 can determine the structure and / or residue sequence 226 of the cyclic peptide based on the first chemical map 224 and sequence information 222 updated over a predetermined number of iterations. As an example, both the structure prediction model 210 and the sequence prediction model 220 can be constructed based on a diffusion model. The design system 110 can utilize the structure prediction model 210 and the sequence prediction model 220 to perform a denoising process for t time steps to obtain an updated atomic representation 218 and sequence information 222. Subsequently, the design system 110 can update the initialized first chemical map 214 based on the updated sequence information 222 to obtain an updated first chemical map 224. This helps improve the compatibility between the designed cyclic peptide and the receptor, as well as the correctness and rationality of the cyclic peptide's structure and amino acid sequence.

[0073] In some embodiments, for a second iteration cycle following the first iteration cycle in the plurality of iteration cycles, the design system 110 may initialize the first chemical map 214, atomic representation 212, and sequence information 216 of the cyclic peptide in the second iteration cycle based on the first chemical map 224, atomic representation 218, and sequence information 222 updated in the first iteration cycle. That is, the design system 110 may initialize the first chemical map 214, atomic representation 212, and sequence information 216 of the subsequent iteration cycle (i.e., the second iteration cycle) based on the updated first chemical map 224, atomic representation 218, and sequence information 222 obtained in the previous iteration cycle (i.e., the first iteration cycle), so as to perform iterative design on the cyclic peptide using the structure prediction model 210 and the sequence prediction model 220. It will be understood that the first iteration cycle and the second iteration cycle here may be any two adjacent iteration cycles in the plurality of iteration cycles.

[0074] In some embodiments, the design system 110 may add noise to the atomic representation 218 updated by the first iteration cycle to obtain a noisy atomic representation. Thereafter, the design system 110 may initialize the atomic representation 212 of the cyclic peptide in the second iteration cycle based on the atomic representation 218 updated by the first iteration cycle and the noisy atomic representation. Figure 3B and Figure 3E As shown, the atomic representation 218 updated by the previous iteration cycle may indicate that the multiple atoms in the cyclic peptide 320 have the following characteristics: Figure 3B The design system 110 can add noise to the atomic representation 218 to obtain a noisy atomic representation. The noisy atomic representation can indicate that multiple atoms in the cyclic peptide 320 have the following spatial positions: Figure 3EThe design system 110 can then initialize the noisy atomic representation 212 in a subsequent iteration based on the updated atomic representation 218 and the noisy atomic representation, so that the structure prediction model 210 constructed based on the diffusion model can predict the spatial positions of multiple atoms in the cyclic peptide 320 based on the reverse diffusion process.

[0075] It should be noted that the design system 110 may need to utilize the structure prediction model 210 and the sequence prediction model 220 to perform the design of the cyclic peptide through multiple iterations. Therefore, the above process may be repeatedly performed multiple times by the design system 110. The embodiments of the present disclosure do not limit the specific number of iterations.

[0076] In this way, in the embodiments of the present disclosure, the geometric constraints of the cyclized structure can be modeled, which can improve the accuracy of chemical bond connections. Moreover, the improved solutions of the embodiments of the present disclosure are applicable to various types of cyclic peptide designs and allow the introduction of non-standard amino acids, which can meet the diverse needs of cyclic peptide design.

[0077] Example procedures, devices, and equipment

[0078] Figure 5 FIG. 5 is a flow chart illustrating a process 500 for designing cyclic peptides according to some embodiments of the present disclosure. The process 500 may be implemented in the design system 110 and will be described below in conjunction with the environment 100 .

[0079] At block 510, the design system 110 initializes a first chemical diagram, an atomic representation, and sequence information of the cyclic peptide to be designed. The first chemical diagram indicates the plurality of atoms contained in the cyclic peptide and the bonding relationships between the plurality of atoms, the plurality of atoms including a plurality of constrained atoms for cyclization and a plurality of free atoms other than the plurality of constrained atoms, the atomic representation indicates the spatial positions of the plurality of atoms, and the sequence information indicates the amino acid sequence of the cyclic peptide.

[0080] In block 520 , the design system 110 updates the initialized atomic representation using the structure prediction model based on the initialized first chemical graph to obtain an updated atomic representation.

[0081] In block 530 , the design system 110 updates the initialized sequence information using the sequence prediction model based on the initialized first chemical graph and the updated atom representation to obtain updated sequence information.

[0082] At block 540 , the design system 110 determines at least one of a structure or an amino acid sequence of the cyclic peptide based at least on the updated atomic representation and the updated sequence information.

[0083] In some embodiments, initializing the first chemical graph, atomic representation, and sequence information of the cyclic peptide includes: initializing the atomic representation and sequence information of the cyclic peptide based on the number of residues of the cyclic peptide, the cyclization type of the cyclic peptide, and the receptor structure representation of the receptor targeted by the cyclic peptide; and initializing the first chemical graph of the cyclic peptide based on the initialized sequence information.

[0084] In some embodiments, initializing the sequence information of the cyclic peptide includes: determining the residue types of two constrained residues corresponding to multiple constrained atoms based on the cyclization type; initializing the residue type of at least one free residue corresponding to multiple free atoms based on the number of residues; and determining the initialized sequence information based on the residue types of the two constrained residues and the residue type of at least one free residue.

[0085] In some embodiments, the cyclic peptide includes two constrained residues corresponding to multiple constrained atoms and at least one free residue corresponding to multiple free atoms, and updating the initialized sequence information using the sequence prediction model includes: based on the initialized sequence information and the cyclization type, removing the atoms included in the side chain of at least one free residue from the initialized first chemical map to obtain a second chemical map; based on the second chemical map and the updated atomic representation, generating a prediction of the residue type of the at least one free residue using the sequence prediction model; and determining the updated sequence information based on the predicted residue type of the at least one free residue and the residue type of the two constrained residues.

[0086] In some embodiments, generating a prediction of the residue type of at least one free residue using a sequence prediction model includes: determining atomic feature representations of multiple free atoms using a sequence prediction model based on a second chemical map and an updated atomic representation; determining a residue feature representation of at least one free residue based on the atomic feature representations of the multiple free atoms; and performing feature decoding on the residue feature representation of the at least one free residue by a decoder to generate a prediction of the residue type of the at least one free residue.

[0087] In some embodiments, determining at least one of the structure or amino acid sequence of the cyclic peptide comprises: updating the initialized first chemical map based on the updated sequence information to obtain an updated first chemical map; and determining at least one of the structure or amino acid sequence of the cyclic peptide based on the updated atomic representation, the updated first chemical map, and the updated sequence information.

[0088] In some embodiments, the cyclic peptide is designed through multiple iterative cycles, and initializing the first chemical map, original representation and sequence information of the cyclic peptide includes: for a second iterative cycle after the first iterative cycle in the multiple iterative cycles, based on the first chemical map, atomic representation and sequence information of the cyclic peptide updated through the first iterative cycle, initializing the first chemical map, atomic representation and sequence information of the cyclic peptide in the second iterative cycle.

[0089] In some embodiments, initializing the atomic representation of the cyclic peptide in the second iteration cycle includes: adding noise to the atomic representation updated by the first iteration cycle to obtain a noisy atomic representation; and initializing the atomic representation of the cyclic peptide in the second iteration cycle based on the atomic representation updated by the first iteration cycle and the noisy atomic representation.

[0090] In some embodiments, determining at least one of the structure or residue sequence of the cyclic peptide comprises determining at least one of the structure or residue sequence of the cyclic peptide based at least on the first chemical map and sequence information updated over a predetermined number of iteration cycles.

[0091] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 6 : A schematic block diagram of an example apparatus 600 for cyclic peptide design according to certain embodiments of the present disclosure is shown. The apparatus 600 can be implemented as or included in the design system 110. Each module / component in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.

[0092] like Figure 6 As shown, the apparatus 600 includes: a cyclic peptide initialization module 610, configured to initialize a first chemical map, an atomic representation, and sequence information of the cyclic peptide to be designed, wherein the first chemical map indicates a plurality of atoms contained in the cyclic peptide and a bonding relationship between the plurality of atoms, the plurality of atoms including a plurality of constrained atoms for cyclization connection and a plurality of free atoms other than the plurality of constrained atoms, the atomic representation indicates the spatial positions of the plurality of atoms, and the sequence information indicates the amino acid sequence of the cyclic peptide; an atomic representation updating module 620, configured to update the initialized atomic representation using a structure prediction model based on the initialized first chemical map to obtain an updated atomic representation; a sequence information updating module 630, configured to update the initialized sequence information using a sequence prediction model based on the initialized first chemical map and the updated atomic representation to obtain updated sequence information; and a cyclic peptide determination module 640, configured to determine at least one of the structure or amino acid sequence of the cyclic peptide based at least on the updated atomic representation and the updated sequence information.

[0093] In some embodiments, the cyclic peptide initialization module 610 is further configured to: initialize the atomic representation and sequence information of the cyclic peptide based on the number of residues of the cyclic peptide, the cyclization type of the cyclic peptide, and the receptor structure representation of the receptor targeted by the cyclic peptide; and initialize the first chemical graph of the cyclic peptide based on the initialized sequence information.

[0094] In some embodiments, the cyclic peptide initialization module 610 is further configured to: determine the residue types of two constrained residues corresponding to the multiple constrained atoms based on the cyclization type; initialize the residue type of at least one free residue corresponding to the multiple free atoms based on the number of residues; and determine the initialized sequence information based on the residue types of the two constrained residues and the residue type of the at least one free residue.

[0095] In some embodiments, the cyclic peptide includes two constrained residues corresponding to multiple constrained atoms and at least one free residue corresponding to multiple free atoms, and the sequence information update module 630 is further configured to: based on the initialized sequence information and the cyclization type, remove the atoms included in the side chain of the at least one free residue from the initialized first chemical map to obtain a second chemical map; based on the second chemical map and the updated atomic representation, generate a prediction of the residue type of the at least one free residue using a sequence prediction model; and determine the updated sequence information based on the predicted residue type of the at least one free residue and the residue type of the two constrained residues.

[0096] In some embodiments, the sequence information update module 630 is further configured to: determine the atomic feature representation of multiple free atoms based on the second chemical map and the updated atomic representation using a sequence prediction model; determine the residue feature representation of at least one free residue based on the atomic feature representation of multiple free atoms; and perform feature decoding on the residue feature representation of at least one free residue through a decoder to generate a prediction of the residue type of at least one free residue.

[0097] In some embodiments, the cyclic peptide determination module 640 is further configured to: update the initialized first chemical map based on the updated sequence information to obtain an updated first chemical map; and determine at least one of the structure or amino acid sequence of the cyclic peptide based on the updated atomic representation, the updated first chemical map and the updated sequence information.

[0098] In some embodiments, the cyclic peptide is designed through multiple iterative cycles, and the cyclic peptide initialization module 610 is further configured to: for a second iterative cycle after the first iterative cycle in the multiple iterative cycles, initialize the first chemical map, atomic representation and sequence information of the cyclic peptide in the second iterative cycle based on the first chemical map, atomic representation and sequence information updated by the first iterative cycle.

[0099] In some embodiments, the cyclic peptide initialization module 610 is further configured to: add noise to the atomic representation updated by the first iteration cycle to obtain a noisy atomic representation; and initialize the atomic representation of the cyclic peptide in the second iteration cycle based on the atomic representation updated by the first iteration cycle and the noisy atomic representation.

[0100] In some embodiments, the cyclic peptide determination module 640 is further configured to determine at least one of the structure or residue sequence of the cyclic peptide based on at least the first chemical map and sequence information updated over a predetermined number of iteration cycles.

[0101] The units and / or modules included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 600 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0102] Figure 7 1 shows a block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 7 The illustrated electronic device 700 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The electronic device 700 shown may include or be implemented as Figure 1 Design Systems 110, or Figure 6 device 600.

[0103] like Figure 7 As shown, electronic device 700 is in the form of a general electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processor 710 may be a real or virtual processor and is capable of performing various processes according to executable instructions stored in memory 720. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capabilities of electronic device 700.

[0104] The electronic device 700 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 700.

[0105] The electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 7 As shown in FIG, a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more executable instruction modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0106] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 700 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 700 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0107] Input device 750 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 760 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 700 may also communicate with one or more external devices (not shown) via communication unit 740 as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with electronic device 700, or with any device that allows electronic device 700 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0108] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer-executable instruction product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0109] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer-executable instruction products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable executable instructions.

[0110] These computer-executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-executable instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0111] Computer-executable instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0112] The flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the systems, methods and computer-executable instruction products according to multiple implementations of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, executable instruction or part of an instruction, and the module, executable instruction or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0113] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for designing cyclic peptides, comprising: Initializing a first chemical map, atomic representation, and sequence information of a cyclic peptide to be designed, wherein the first chemical map indicates a plurality of atoms contained in the cyclic peptide and bonding relationships between the plurality of atoms, the plurality of atoms including a plurality of constrained atoms for cyclization connection and a plurality of free atoms other than the plurality of constrained atoms, the atomic representation indicates spatial positions of the plurality of atoms, and the sequence information indicates an amino acid sequence of the cyclic peptide; Based on the initialized first chemical graph, updating the initialized atomic representation using a structure prediction model to obtain an updated atomic representation; updating the initialized sequence information using a sequence prediction model based on the initialized first chemical map and the updated atomic representation to obtain updated sequence information; as well as At least one of a structure or an amino acid sequence of the cyclic peptide is determined based at least on the updated atomic representation and the updated sequence information.

2. The method of claim 1 , wherein initializing the first chemical map, atomic representation, and sequence information of the cyclic peptide comprises: Initializing the atomic representation and sequence information of the cyclic peptide based on the number of residues in the cyclic peptide, the cyclization type of the cyclic peptide, and the receptor structure representation of the receptor targeted by the cyclic peptide; as well as Based on the initialized sequence information, a first chemical map of the cyclic peptide is initialized.

3. The method according to claim 2, wherein initializing the sequence information of the cyclic peptide comprises: determining residue types of two constrained residues corresponding to the plurality of constrained atoms based on the cyclization type; Initializing a residue type of at least one free residue corresponding to the plurality of free atoms based on the residue number; and The initialized sequence information is determined based on the residue types of the two constrained residues and the residue type of the at least one free residue.

4. The method according to claim 1, wherein the cyclic peptide comprises two constrained residues corresponding to the plurality of constrained atoms and at least one free residue corresponding to the plurality of free atoms, and updating the initialized sequence information using the sequence prediction model comprises: Based on the initialized sequence information and the cyclization type, removing atoms included in the side chain of the at least one free residue from the initialized first chemical map to obtain a second chemical map; generating, using the sequence prediction model, a prediction of a residue type for the at least one free residue based on the second chemical map and the updated atomic representation; as well as The updated sequence information is determined based on the predicted residue type of the at least one free residue and the residue types of the two constrained residues.

5. The method of claim 4, wherein generating a prediction of the residue type of the at least one free residue using the sequence prediction model comprises: determining atomic feature representations of the plurality of free atoms using the sequence prediction model based on the second chemical map and the updated atomic representation; determining a residue feature representation of the at least one free residue based on the atom feature representations of the plurality of free atoms; and A decoder is used to perform feature decoding on the residue feature representation of the at least one free residue to generate a prediction of the residue type of the at least one free residue.

6. The method of claim 1, wherein determining at least one of the structure or amino acid sequence of the cyclic peptide comprises: updating the initialized first chemical map based on the updated sequence information to obtain an updated first chemical map; as well as At least one of a structure or an amino acid sequence of the cyclic peptide is determined based on the updated atomic representation, the updated first chemical map, and the updated sequence information.

7. The method of claim 6, wherein the cyclic peptide is designed through a plurality of iterative cycles, and initializing the first chemical map, raw representation, and sequence information of the cyclic peptide comprises: For a second iteration cycle following the first iteration cycle in the multiple iteration cycles, the first chemical map, atomic representation and sequence information of the cyclic peptide in the second iteration cycle are initialized based on the first chemical map, atomic representation and sequence information of the cyclic peptide updated through the first iteration cycle.

8. The method according to claim 7, wherein initializing the atomic representation of the cyclic peptide in the second iteration cycle comprises: adding noise to the atomic representation updated by the first iteration cycle to obtain a noisy atomic representation; as well as Based on the atomic representation updated in the first iteration cycle and the atomic representation after noise addition, the atomic representation of the cyclic peptide in the second iteration cycle is initialized.

9. The method of claim 7, wherein determining at least one of the structure or residue sequence of the cyclic peptide comprises: At least one of the structure or residue sequence of the cyclic peptide is determined based on at least the first chemical map and the sequence information updated over a predetermined number of iteration cycles.

10. A device for cyclic peptide design, comprising: a cyclic peptide initialization module, configured to initialize a first chemical map, an atomic representation, and sequence information of a cyclic peptide to be designed, wherein the first chemical map indicates a plurality of atoms contained in the cyclic peptide and bonding relationships between the plurality of atoms, the plurality of atoms including a plurality of constrained atoms for cyclization connection and a plurality of free atoms other than the plurality of constrained atoms, the atomic representation indicates spatial positions of the plurality of atoms, and the sequence information indicates an amino acid sequence of the cyclic peptide; an atomic representation updating module configured to update the initialized atomic representation using a structure prediction model based on the initialized first chemical graph to obtain an updated atomic representation; a sequence information updating module configured to update the initialized sequence information using a sequence prediction model based on the initialized first chemical map and the updated atomic representation to obtain updated sequence information; as well as The cyclic peptide determination module is configured to determine at least one of the structure or amino acid sequence of the cyclic peptide based at least on the updated atomic representation and the updated sequence information.

11. An electronic device comprising: at least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processor. 12 . A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to claim 1 .

13. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.