Methods and electronic devices for constructing polypeptide molecules

By constructing peptide molecules using a generative model and obtaining the coding table using the VQ-VAE model, the secondary structure and amino acid sequence are determined, which solves the problem of insufficient antibacterial activity of peptide molecules in existing technologies and achieves a more efficient antibacterial effect.

CN114155909BActive Publication Date: 2025-10-28BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111467002.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-10-28
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively construct polypeptide molecules with high antibacterial activity. The mechanism of action of antimicrobial peptides requires a reasonable amino acid sequence and appropriate structure, and existing methods have not fully considered the influence of secondary structure.

Method used

A generative model is used to construct peptide molecules. The coding table is obtained using the VQ-VAE model. The secondary structure of the peptide molecule is determined by the first decoder and the amino acid sequence is determined by the second decoder. Feature representation is generated to construct the target peptide molecule.

Benefits of technology

The generated polypeptide molecules have higher antibacterial activity, and the bactericidal effect of the polypeptide molecules is improved by taking into account the influence of secondary structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155909B_ABST
    Figure CN114155909B_ABST
Patent Text Reader

Abstract

According to embodiments of this disclosure, a method, apparatus, device, storage medium, and program product for constructing polypeptide molecules are provided. The method described herein includes: obtaining a set of encoding tables for a generative model, the set of encoding tables including multiple discrete encoding representations; the generative model including a first decoder and a second decoder; the set of encoding tables being used to construct a first input to the first decoder and a second input to the second decoder; the first decoder being used to determine the secondary structure of the polypeptide molecule based on the first input; and the second decoder being used to determine the amino acid sequence of the polypeptide molecule based on the second input; constructing a first feature representation and a second feature representation based on the multiple discrete encoding representations in the set of encoding tables; and using the generative model to determine the structural information of a target polypeptide molecule. According to embodiments of this disclosure, by considering secondary structure during the construction of polypeptide molecules, polypeptide molecules with higher antibacterial activity can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The various implementations of this disclosure relate to the field of computers, and more specifically, to methods, apparatus, devices, and computer storage media for constructing polypeptide molecules. Background Technology

[0002] Peptides are compounds formed by amino acids linked together by peptide bonds. Antimicrobial peptides (AMPs) have shown promising efficacy in broad-spectrum antibiotic and anti-infective therapy. AMPs are an emerging therapeutic agent defined as short proteins of fewer than 50 amino acids with potent antimicrobial activity.

[0003] Unlike traditional drugs, antimicrobial peptides can attach to bacterial membranes and form pores in them, thereby killing the bacteria. This method of physically destroying bacteria is called a "barrel stave." In this bactericidal process, the antimicrobial activity of the antimicrobial peptide is closely related to its secondary structure. Summary of the Invention

[0004] In a first aspect of this disclosure, a method for constructing a polypeptide molecule is provided. The method includes: obtaining a set of encoding tables for a generative model, the set of encoding tables including multiple discrete encoding representations; the generative model including a first decoder and a second decoder; the set of encoding tables being used to construct a first input to the first decoder and a second input to the second decoder; the first decoder being used to determine the secondary structure of the polypeptide molecule based on the first input; and the second decoder being used to determine the amino acid sequence of the polypeptide molecule based on the second input; constructing a first feature representation and a second feature representation based on the multiple discrete encoding representations in the set of encoding tables; using the first decoder to determine a target secondary structure of the target polypeptide molecule according to the first feature representation; and using the second decoder to determine the target amino acid sequence of the target polypeptide molecule according to the second feature representation.

[0005] In a second aspect of this disclosure, an electronic device is provided, comprising: a memory and a processor; wherein the memory is configured to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to the first aspect of this disclosure.

[0006] In a third aspect of this disclosure, a computer-readable storage medium is provided having one or more computer instructions stored thereon, wherein the one or more computer instructions are executed by a processor to implement the method according to a first aspect of this disclosure.

[0007] In a fourth aspect of this disclosure, a computer program product is provided, comprising one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the method according to a first aspect of this disclosure.

[0008] Based on this approach, embodiments of this disclosure can consider secondary structure during the construction of polypeptide molecules, thereby obtaining polypeptide molecules with higher antibacterial activity. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1A and Figure 1B The application comparison of polypeptide molecules with different structures is shown;

[0011] Figure 2 A schematic block diagram of a computing device capable of implementing some embodiments of the present disclosure is shown;

[0012] Figure 3 A schematic diagram of a trained generative model according to some embodiments of the present disclosure is shown;

[0013] Figure 4 Schematic diagrams illustrating the construction of polypeptide molecules using generative models according to some embodiments of the present disclosure are shown; and

[0014] Figure 5 A flowchart illustrating an example method for constructing polypeptide molecules according to some embodiments of this disclosure is shown. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0017] As discussed above, antimicrobial peptides (AMPs), as an emerging class of therapeutic agents, have already shown promising effects in broad-spectrum antibiotic and anti-infective treatments. Specifically, antimicrobial peptides can physically kill bacteria by disrupting bacterial membranes through a "pore" mechanism.

[0018] Since most bacterial surfaces are anionic, positively charged amino acids are more likely to bind to the bacterial membrane, while highly hydrophobic amino acids tend to migrate from the solution environment to the bacterial membrane. However, the mechanism of action of antimicrobial peptides requires not only a suitable sequence but also an appropriate structure. For example, by forming a helical structure, antimicrobial peptides can collect hydrophobic amino acids on one side and hydrophilic amino acids on the other. This ability, known as amphipathicity, helps antimicrobial peptides insert into the membrane and maintain stable pores with other peptide molecules in the membrane, thereby killing bacteria more effectively.

[0019] Figure 1A and Figure 1B A schematic diagram comparing the applications of polypeptide molecules with different structures is shown. It can be seen that, for example... Figure 1A As shown, polypeptide molecule 110A can only attach to bacterial membrane 120A, but it is difficult to form pores. Conversely, as Figure 1B As shown, due to its amphiphilic nature, the helical polypeptide molecule 110B can more easily form stable pores in the bacterial membrane 120B. Therefore, the secondary structure of a polypeptide molecule directly affects its antibacterial activity.

[0020] According to an implementation of this disclosure, a scheme for constructing polypeptide molecules is provided. In this scheme, a set of encoding tables for a generative model can be obtained, wherein the set of encoding tables includes multiple discrete encoding representations. The generative model includes a first decoder and a second decoder. The set of encoding tables is used to construct a first input to the first decoder and a second input to the second decoder. The first decoder is used to determine the secondary structure of the polypeptide molecule based on the first input, and the second decoder is used to determine the amino acid sequence of the polypeptide molecule based on the second input. Exemplarily, the generative model may be, for example, a VQ-VAE model (Vector Quantization-Variational Autoencoder).

[0021] Furthermore, a first feature representation and a second feature representation can be constructed based on multiple discrete encoding representations in a set of encoding tables. The first decoder is used to determine the target secondary structure of the target polypeptide molecule based on the first feature representation, and the second decoder is used to determine the target amino acid sequence of the target polypeptide molecule based on the second feature representation.

[0022] Based on this approach, the feature representation generated by the embodiments of this disclosure can take into account the influence of secondary structure, and can directly generate the amino acid sequence and secondary structure of the target polypeptide molecule using a decoder. Therefore, the embodiments of this disclosure can construct polypeptide molecules with the desired secondary structure, thereby improving the antibacterial activity of the constructed polypeptide molecule.

[0023] The basic principles and several example implementations of this disclosure are illustrated below with reference to the accompanying drawings.

[0024] Example device

[0025] Figure 2 A schematic block diagram of an example computing device 200 that can be used to implement embodiments of the present disclosure is shown. It should be understood that... Figure 2 The device 200 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the implementation described in this disclosure. Figure 2 As shown, the components of device 200 may include, but are not limited to, one or more processors or processing units 210, memory 220, storage device 230, one or more communication units 240, one or more input devices 250, and one or more output devices 260.

[0026] In some embodiments, device 200 can be implemented as various user terminals or service terminals. Service terminals can be servers, large computing devices, etc., provided by various service providers. User terminals include any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, multimedia computers, multimedia tablets, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is also foreseeable that device 200 can support any type of user-facing interface (such as "wearable" circuitry).

[0027] Processing unit 220 can be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 220. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of device 200. Processing unit 220 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0028] Device 200 typically includes multiple computer storage media. Such media can be any available media accessible to device 200, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 220 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Memory 220 may include one or more design modules 225 configured to perform the functions of the various implementations described herein. Design modules 225 can be accessed and executed by processing unit 210 to implement the corresponding functions. Storage device 230 can be a removable or non-removable medium and may include machine-readable media capable of storing information and / or data and accessible within device 200.

[0029] The functionality of the components of device 200 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, device 200 can operate in a networked environment using logical connections to one or more other servers, personal computers (PCs), or other general network nodes. Device 200 can also communicate as needed with one or more external devices (not shown) via communication unit 240, such as database 245, other storage devices, servers, display devices, etc., with one or more devices that enable user interaction with device 200, or with any device that enables device 200 to communicate with one or more other computing devices (e.g., network interface card, modem, etc.). Such communication can be performed via input / output (I / O) interfaces (not shown).

[0030] Input device 250 can be one or more various input devices, such as a mouse, keyboard, trackball, voice input device, camera, etc. Output device 260 can be one or more output devices, such as a monitor, speaker, printer, etc.

[0031] In some embodiments, such as Figure 2 As shown, device 200 can acquire a set of codebooks 270, which may include, for example, multiple trained discrete code representations. Exemplarily, device 200 may receive the set of codebooks 270 via input device 250. Alternatively, device 200 may also read the set of codebooks 270 from storage device 230 or database 245. Alternatively, device 200 may also receive the set of codebooks 270 from other devices via communication unit 240.

[0032] In some embodiments, the construction module 225 can construct a polypeptide molecule based on the coding table 270. Specifically, the construction module 225 can determine the structural information 280 of the polypeptide molecule, which may include the target amino acid sequence 282 and the target secondary structure 284 of the polypeptide molecule. The process of constructing the polypeptide molecule will be described in detail below.

[0033] Training generative models

[0034] In some embodiments, the construction module 225 may utilize a generative model to construct a target peptide molecule and determine the target amino acid sequence 282 and target secondary structure 284 of the target peptide molecule. In some embodiments, the generative model may be, for example, a VQ-VAE model. Reference will be made below. Figure 3 This describes an example process for training and generating model 300.

[0035] like Figure 3 As shown, the generative model 300 may include an encoder 320, a set of encoding tables 350, a generator 360, and a classifier 380. In some embodiments, as will be described in detail below, the generative model 300 may also include a set of pattern selectors 395.

[0036] In some embodiments, encoder 320 may acquire an amino acid sequence 310 of a set of training polypeptide molecules, and then determine a set of amino acid feature representations 330 corresponding to a set of amino acids in the amino acid sequence 310.

[0037] For example, the amino acid sequence 310 of the training polypeptide molecule can be represented as x = {a1, a2, ..., a...} L}, where α belongs to 20 common amino acids, and L represents the length of the amino acid sequence 310. The set of amino acid feature representations 330 generated by encoder 320 can be represented as z = z 1∶L .

[0038] In some embodiments, the generative model 300 can use vector quantization to find the discrete encoding representation corresponding to each amino acid feature representation 320. Exemplarily, the generative model 300 can utilize nearest neighbor search algorithms in the encoding table 350 (e.g., which can be represented as...). Where K represents the size of the encoding table and d represents the dimension of entry e in the encoding table, the search is performed to find the amino acid feature representation 330 generated by encoder 320 (e.g., it can be represented as...). The corresponding encoding table entry is also called the discrete encoding representation (for example, it can be represented as z). q ={z q (a1), ..., z q (a L Therefore, the process can be represented as:

[0039] z q (a i ) = e k k = argmin j∈K ||z e (a i )-e j ||2 (1)

[0040] In some embodiments, the feature representation determined by the generative model 300 through vector quantization can be provided to the generator 360 (also referred to as the second decoder) for generating the reconstructed amino acid sequence 370.

[0041] In some embodiments, the loss function for generating the reconstructed amino acid sequence 370 can be expressed as:

[0042]

[0043] Where sg(·) represents the gradient stopping operator, and β represents the weight coefficient; log p(a i |z q (a i The purpose of this part is to make the reconstructed amino acid sequence 370 close to the amino acid sequence 310 of the training peptide molecule, which is related to the processing of generator 360. Partially representing the difference between the feature representation output by the encoder and the feature representation obtained by looking up the encoding table, its purpose is to make the feature representation output by the encoder close to the feature representation obtained by looking up the encoding table, that is, related to the lookup process of a set of encoding tables 350.

[0044] In some embodiments, the training of the secondary structure of the peptide molecule can also be considered during the training of the generative model 300. For example, the training of the secondary structure of the peptide molecule can be represented as y = {y1, y2, ..., y...} L}, y i ∈{H, B, E, G, I, T, S, -}, where “H” (α-spiral), “B” (β-bridge), “E” (fold), “G” (spiral-3), “I” (spiral-5), “T” (turn), “S” (bend) and “-” (unknown type) represent different secondary structure types.

[0045] In some embodiments, a generative model 300 can be trained based on the secondary structure of a training peptide molecule. Specifically, an encoder 320 and vector quantization can be used to determine the input features z′ to a classifier 380 (also known as a first decoder). q (a i Furthermore, the loss function related to predicting secondary structure can be expressed as:

[0046]

[0047] Similarly, log p(y i |z′ q (a i The purpose of this part is to make the predicted secondary structure determined by classifier 380 close to the secondary structure of the training peptide molecule, that is, to be related to the processing of classifier 380. Partially representing the difference between the feature representation output by the encoder and the feature representation obtained by looking up the encoding table, its purpose is to make the feature representation output by the encoder close to the feature representation obtained by looking up the encoding table, that is, related to the lookup process of a set of encoding tables 350.

[0048] In some embodiments, different input features can be constructed for the generator 360 and the classifier 380. For example... Figure 3 As shown, the generative model 300 may also include a set of pattern selectors 395, which can be configured to extract patterns (also known as combined feature representations) at different scales from a set of amino acid feature representations 330.

[0049] A sequence consisting of 330 amino acid features can be understood as a pattern with a scale of 0; a pattern with a scale of 1 can be understood as the pattern corresponding to each amino acid in the sequence; and a pattern with a scale of n can be understood as the pattern corresponding to all subsequences of length n in the sequence.

[0050] Accordingly, the pattern selector 395 can determine one or more sub-amino acid sequences matching the corresponding length based on a set of amino acids in the amino acid sequence 310, and further determine the corresponding combined feature representation based on the one or more sub-amino acid sequences. Patterns at different scales extracted by a set of pattern selectors 395 can be represented as follows:

[0051]

[0052] Among them, F (n) This represents the processing procedure of a set of selectors 350, h i This represents a set of amino acid features output by encoder 320, denoted as 330.

[0053] Furthermore, the generative model 300 can utilize a set of encoding tables 360 to update multiple combined feature representations generated by a set of pattern selectors 395. To obtain multiple updated combined feature representations Also known as target discrete coding representation.

[0054] In some embodiments, the generative model may generate an input feature representation to the generator 360 based on multiple updated combined feature representations. In some embodiments, the generative model 300 may select a set of combined feature representations (also referred to as a set of discrete coded representations) from multiple updated combined feature representations to construct an input feature representation to the generator 360.

[0055] For example, the input feature representation to generator 360 can be represented as:

[0056]

[0057] Where N r This represents a set of encoding tables selected for constructing the input feature representation to the generator, and || represents a concatenation operation.

[0058] Accordingly, based on this approach, the representation of the loss function (2) can be updated as follows:

[0059]

[0060] Similarly, the representation of the loss function (3) can be updated to obtain L. s Furthermore, the total loss function used to train the generative model 300 can be expressed as:

[0061] L = L r +γL s (7)

[0062] Where γ represents the weighting coefficient. Therefore, embodiments of this disclosure can take into account the influence of secondary structure during the training of the generative model.

[0063] In some embodiments, known AMP peptide molecules can be used to train the generative model 300. Considering the limitations of known AMP peptide molecule datasets, large protein datasets can also be used to pre-train the sequence construction task, and peptide datasets including protein information can be used to pre-train the secondary structure classification task. Furthermore, AMP peptide molecule datasets can be used to fine-tune the generative model.

[0064] It should be understood that the generative model can be trained based on the loss function discussed above using any appropriate VQ-VAE model training method (e.g., using exponential moving average EMA to update the encoding table).

[0065] Constructing polypeptide molecules

[0066] After training the generative model 300, the construction module 225 can further utilize a set of coding tables 350 from the generative model 300 to construct peptide molecules. It should be understood that the construction device used to construct the peptide molecules (e.g., device 200) can be different from or the same as the training device used to train the generative model 300. The following will refer to... Figure 4 This describes an example process for constructing polypeptide molecules.

[0067] like Figure 4 As shown, the construction device can construct feature representations to generator 360 and to classifier 380 based on a set of encoding tables 350 in generative model 300.

[0068] In some embodiments, the building device may determine an index sequence 420. The index sequence may, for example, include multiple index values ​​X. 1- X N Each index value can indicate the discrete encoding representation selected in the corresponding encoding table.

[0069] Furthermore, the construction device can construct feature representations to classifier 380 (also known as first feature representations) and to generator 360 (also known as second feature representations) based on multiple target discrete coding representations selected from a set of coding tables 350. It should be understood that the construction process discussed with reference to formula (5) can be used to construct feature representations to generator 360 and to classifier 380.

[0070] Specifically, the device can construct a first feature representation based on a first set of discrete coding representations among multiple target discrete coding representations, and construct a second feature representation based on a second set of discrete coding representations among multiple target discrete coding representations.

[0071] In some embodiments, the first set of discrete coded representations may differ from the second set of discrete coded representations. For example, the first set of discrete coded representations may correspond to the first to m coded tables, while the second set of discrete coded representations may correspond to the (m+1)th to Nth coded tables.

[0072] In some embodiments, the first set of discrete coded representations may at least partially overlap with the second set of discrete coded representations. For example, the first set of discrete coded representations may correspond to the 1st to the mth coding tables, while the second set of discrete coded representations may correspond to the mth to the Nth coding tables. Both sets of discrete coded representations may include the target discrete coded representation selected from the mth coding table.

[0073] Furthermore, the constructing device can utilize classifier 380 to generate the target secondary structure 284 of the target polypeptide molecule based on a first feature representation. Correspondingly, the constructing device can also utilize generator 360 to generate the target amino acid sequence 282 of the target polypeptide molecule based on a second feature representation.

[0074] Based on this approach, embodiments of this disclosure can provide not only the amino acid sequence of the target polypeptide molecule, but also the secondary structure of the target polypeptide molecule.

[0075] In some embodiments, such as Figure 4 As shown, the index sequence 420 can be generated by the construction device using the random sequence generation model 410. In some embodiments, the random sequence generation model is trained on a set of training index sequences for a set of training polypeptide molecules, wherein the set of training index sequences indicates discrete coding representations selected from multiple coding tables.

[0076] After training the random sequence generation model 410, the construction device can, for example, use the random sequence generation model 410 to generate an index sequence 420 based on the initial input or randomly.

[0077] In some embodiments, the fabrication device may further determine whether the generated target secondary structure 284 satisfies structural constraints. In some embodiments, structural constraints may include, for example, constraints regarding the proportion of random coils in the secondary structure, such as requiring the proportion of random coils to be less than 30%. Alternatively, structural constraints may also include constraints regarding the length of alpha helices in the secondary structure, such as requiring the length of the alpha helix to be greater than 4. Such structural constraints ensure the antibacterial activity of the generated target peptide molecule.

[0078] Furthermore, if the target secondary structure is determined to satisfy the structural constraints, the construction device further utilizes a second decoder to determine the target amino acid sequence 282 of the target polypeptide molecule based on the second feature representation.

[0079] Conversely, if the target secondary structure is determined to satisfy structural constraints, the construction device can discard the index sequence. Additionally, the construction device can also construct new first feature representations and new second feature representations based on multiple discrete coded representations in a set of coding tables. For example, the construction device can utilize a random sequence generation model 410 to generate new random sequences.

[0080] In some embodiments, the building device may also generate multiple index sequences at once and discard index sequences in which the predicted secondary structure does not satisfy the structural constraints.

[0081] Based on the process of constructing polypeptide molecules discussed above, embodiments of this disclosure can enable input features to fully take into account the influence of secondary structure, thereby enabling the construction of polypeptide molecules (e.g., antimicrobial peptides) with superior antibacterial activity.

[0082] Example process

[0083] Figure 5 A flowchart of a method 600 for constructing a polypeptide molecule according to some implementations of the present disclosure is shown. Method 500 can be implemented by computing device 200, for example, it can be implemented at a construction module 225 in memory 220 of computing device 200.

[0084] like Figure 5 As shown in box 510, computing device 200 acquires a set of encoding tables for a generative model. The set of encoding tables includes multiple discrete encoding representations. The generative model includes a first decoder and a second decoder. The set of encoding tables is used to construct a first input to the first decoder and a second input to the second decoder. The first decoder is used to determine the secondary structure of the polypeptide molecule based on the first input, and the second decoder is used to determine the amino acid sequence of the polypeptide molecule based on the second input.

[0085] In box 520, computing device 200 constructs a first feature representation and a second feature representation based on multiple discrete coded representations in a set of coded tables.

[0086] In box 530, computing device 200 uses a first decoder to determine the target secondary structure of the target polypeptide molecule based on a first feature representation.

[0087] In box 540, computing device 200 uses a second decoder to determine the target amino acid sequence of the target polypeptide molecule based on the second feature representation.

[0088] It should be understood that Figure 5 This is not intended to limit the execution order of the steps in each box. For example, the steps in boxes 530 and 540 can be executed in parallel, box 530 can be executed before box 540, or box 540 can be executed before box 530.

[0089] In some embodiments, a set of encoding tables includes multiple encoding tables, each encoding table including a set of discrete encoding representations.

[0090] In some embodiments, constructing a first feature representation and a second feature representation includes: determining an index sequence, the index sequence including a plurality of index values, each index value indicating a target discrete coding representation selected in a corresponding coding table; and constructing a first feature representation and a second feature representation based on the plurality of target discrete coding representations selected in the plurality of coding tables.

[0091] In some embodiments, constructing a first feature representation and a second feature representation based on multiple target discrete coding representations selected from multiple coding tables includes: constructing a first feature representation based on a first set of discrete coding representations among the multiple target discrete coding representations; and constructing a second feature representation based on a second set of discrete coding representations among the multiple target discrete coding representations, wherein the first set of discrete coding representations is different from the second set of discrete coding representations.

[0092] In some embodiments, determining the index sequence includes: using a random sequence generation model trained on a set of training index sequences for a set of training peptide molecules, the set of training index sequences indicating discrete coding representations selected from multiple coding tables.

[0093] In some embodiments, determining the target amino acid sequence of a target polypeptide molecule using a second decoder based on a second feature representation includes: determining whether a target secondary structure satisfies structural constraints, the structural constraints including at least one of the following: constraints regarding the proportion of random coils in the secondary structure, or constraints regarding the length of alpha helices in the secondary structure; and in response to determining that the target secondary structure satisfies the structural constraints, using the second decoder to determine the target amino acid sequence of the target polypeptide molecule based on the second feature representation.

[0094] In some embodiments, method 600 further includes: in response to determining that the target secondary structure satisfies structural constraints, constructing a new first feature representation and a new second feature representation based on multiple discrete coding representations in a set of coding tables.

[0095] In some embodiments, a set of encoding tables includes multiple encoding tables, and the generative model is trained based on the following process: using the encoder of the generative model to determine a set of amino acid feature representations corresponding to a set of amino acids in the training polypeptide molecule; generating multiple combined feature representations corresponding to multiple amino acid sequence lengths based on the set of amino acid feature representations; updating the multiple combined feature representations using the multiple encoding tables corresponding to multiple amino acid sequence lengths; and determining a loss function for training the generative model based on the updated multiple combined amino acid feature representations.

[0096] In some embodiments, generating multiple combined feature representations corresponding to multiple amino acid sequence lengths based on a set of amino acid feature representations includes: for a first length among multiple amino acid sequence lengths, determining a set of sub-amino acid sequences that match the first length based on a set of amino acids; and using a set of amino acid feature representations to determine a combined feature representation corresponding to the set of sub-amino acid sequences.

[0097] In some embodiments, the loss function includes a first portion associated with a first decoder, a second portion associated with a second decoder, and a third portion associated with updates utilizing multiple encoding tables.

[0098] In some embodiments, the first training input to the first decoder is determined by updating the first initial input using multiple encoding tables, the second training input to the second decoder is determined by updating the second initial input using multiple encoding tables, and the third portion is determined based on a first difference between the first initial input and the first training input and a second difference between the second initial input and the second training input.

[0099] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload programmable logic devices (CPLDs), and so on.

[0100] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0101] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0102] Furthermore, although the operations are described in a specific order, this should be understood as requiring that such operations be performed in the specific order shown or in sequential order, or requiring that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of a single implementation may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.

[0103] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for constructing a polypeptide molecule, comprising: A set of encoding tables for a generative model is obtained, the set of encoding tables including multiple discrete encoding representations, the generative model including a first decoder and a second decoder, the set of encoding tables being used to construct a first input to the first decoder and a second input to the second decoder, the first decoder being used to determine the secondary structure of the polypeptide molecule based on the first input, and the second decoder being used to determine the amino acid sequence of the polypeptide molecule based on the second input; Based on the multiple discrete coding representations in the set of coding tables, a first feature representation and a second feature representation are constructed. Using the first decoder, the target secondary structure of the target polypeptide molecule is determined based on the first feature representation; and Using the second decoder, the target amino acid sequence of the target polypeptide molecule is determined based on the second feature representation.

2. The method according to claim 1, wherein the set of encoding tables includes a plurality of encoding tables, each encoding table including a set of discrete encoding representations.

3. The method according to claim 2, wherein constructing the first feature representation and the second feature representation comprises: Determine an index sequence, which includes multiple index values, each index value indicating the target discrete coding representation selected in the corresponding coding table; as well as Based on the multiple target discrete coding representations selected from the multiple coding tables, the first feature representation and the second feature representation are constructed.

4. The method of claim 3, wherein constructing the first feature representation and the second feature representation based on a plurality of target discrete coding representations selected from the plurality of coding tables comprises: The first feature representation is constructed based on the first set of discrete coding representations among the plurality of target discrete coding representations; as well as The second feature representation is constructed based on the second set of discrete coding representations among the plurality of target discrete coding representations, wherein the first set of discrete coding representations is different from the second set of discrete coding representations.

5. The method of claim 3, wherein determining the index sequence comprises: The index sequence is determined using a random sequence generation model trained on a set of training index sequences for a set of training polypeptide molecules, the set of training index sequences indicating the discrete coding representation selected in the plurality of coding tables.

6. The method of claim 1, wherein determining the target amino acid sequence of the target polypeptide molecule using the second decoder based on the second feature representation comprises: Determine whether the target secondary structure satisfies structural constraints, which include at least one of the following: a constraint regarding the proportion of irregular curls in the secondary structure, or a constraint regarding the length of the alpha spirals in the secondary structure; and In response to determining that the target secondary structure satisfies the structural constraints, the second decoder is used to determine the target amino acid sequence of the target polypeptide molecule based on the second feature representation.

7. The method according to claim 6, further comprising: In response to determining that the target secondary structure satisfies the structural constraints, a new first feature representation and a new second feature representation are constructed based on the multiple discrete coding representations in the set of coding tables.

8. The method of claim 1, wherein a set of encoding tables comprises multiple encoding tables, and the generative model is trained based on the following process: The encoder of the generative model is used to determine a set of amino acid feature representations corresponding to a set of amino acids in the training polypeptide molecule; Based on the set of amino acid feature representations, generate multiple combined feature representations corresponding to the lengths of multiple amino acid sequences; The multiple combined feature representations are updated using the multiple coding tables corresponding to the lengths of the multiple amino acid sequences; as well as Based on the updated combined feature representations, a loss function is determined for training the generative model.

9. The method according to claim 8, wherein generating multiple combined feature representations corresponding to the lengths of multiple amino acid sequences based on the set of amino acid feature representations comprises: Regarding the first length among the multiple amino acid sequence lengths, Based on the set of amino acids, a set of sub-amino acid sequences matching the first length are determined; as well as Using a set of amino acid feature representations, a combination feature representation corresponding to the set of sub-amino acid sequences is determined.

10. The method of claim 8, wherein the loss function includes a first portion associated with the first decoder, a second portion associated with the second decoder, and a third portion associated with the update utilizing the plurality of encoding tables.

11. The method of claim 10, wherein the first training input to the first decoder is determined by updating a first initial input using the plurality of encoding tables, the second training input to the second decoder is determined by updating a second initial input using the plurality of encoding tables, and the third portion is determined based on a first difference between the first initial input and the first training input and a second difference between the second initial input and the second training input.

12. An electronic device, comprising: Memory and processor; The memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1 to 11.

13. A computer-readable storage medium having stored thereon one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the method according to any one of claims 1 to 11.

14. A computer program product comprising one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Antibacterial peptide generation and recognition method and system

    CN116206690A