A method and apparatus for constructing a protein sequence prediction model

By constructing a protein sequence prediction model and training an encoder using surface point cloud information and sequence information, the problem of insufficient accuracy in protein sequence prediction in existing technologies is solved, and accurate prediction and functional regulation of amino acids on the protein surface are achieved.

CN119724325BActive Publication Date: 2026-04-28SHANGHAI MOLECULAR HEART INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI MOLECULAR HEART INTELLIGENT TECH CO LTD
Filing Date
2024-12-05
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In the prior art, protein sequence prediction methods based on protein backbones have limited accuracy in predicting amino acids on the protein surface, and cannot effectively design protein sequences to regulate their functions.

Method used

By constructing a protein sequence prediction model, the structural encoder, sequence encoder, and decoder are trained using the surface point cloud information and sequence information of sample proteins, and the amino acid sequence is accurately predicted by combining protein surface feature information.

Benefits of technology

It improves the accuracy of protein sequence prediction, enabling more precise design of amino acids on the protein surface and better regulation of protein interactions with other molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724325B_ABST
    Figure CN119724325B_ABST
Patent Text Reader

Abstract

The application aims to provide a method and device for constructing a protein sequence prediction model, which comprises: determining corresponding surface point cloud information and sample protein sequence information of a sample protein based on sample protein information, wherein the sample protein information comprises protein sequence information and protein structure information, and sequence information belonging to the surface of the sample protein in the sample protein sequence information is masked; and training a corresponding protein sequence prediction model based on the surface point cloud information and the sample protein sequence information of the sample protein, wherein the protein sequence prediction model comprises a structure encoder, a sequence encoder and a corresponding decoder. In the construction of the protein sequence prediction model, the application introduces information of a protein surface for sequence prediction, which can effectively improve the prediction accuracy of the model and improve the surface accuracy of the predicted protein sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics technology, and in particular to a technique for constructing protein sequence prediction models. Background Technology

[0002] Proteins regulate various physiological responses through interactions with other proteins or small molecules on their surface. Precisely designing the amino acids on the protein surface is very helpful in regulating protein function. Currently, protein sequences can be predicted relatively accurately based on the protein backbone. However, because the designed sequences are based on the backbone, their accuracy in predicting the surface amino acids that actually affect the protein is limited, and therefore they cannot be used most effectively for protein sequence design. Summary of the Invention

[0003] One objective of this application is to provide a method and apparatus for constructing protein sequence prediction models.

[0004] According to one aspect of this application, a method for constructing a protein sequence prediction model is provided, the method comprising:

[0005] Based on the sample protein information, the surface point cloud information and sample protein sequence information corresponding to the sample protein are determined. The sample protein information includes protein sequence information and protein structure information. The sequence information belonging to the sample protein surface in the sample protein sequence information is masked.

[0006] Based on the surface point cloud information corresponding to the sample protein and the sample protein sequence information, a corresponding protein sequence prediction model is trained and obtained, wherein the protein sequence prediction model includes a structure encoder, a sequence encoder and a corresponding decoder.

[0007] According to one aspect of this application, a computer device for constructing a protein sequence prediction model is provided, comprising a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of any of the methods described above.

[0008] According to one aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0009] According to one aspect of this application, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of any of the methods described above.

[0010] According to one aspect of this application, a device for constructing a protein sequence prediction model is provided, the device comprising:

[0011] The module is used to determine the surface point cloud information and sample protein sequence information corresponding to the sample protein based on the sample protein information. The sample protein information includes protein sequence information and protein structure information. The sequence information belonging to the surface of the sample protein in the sample protein sequence information is masked.

[0012] The first and second modules are used to train and obtain a corresponding protein sequence prediction model based on the surface point cloud information corresponding to the sample protein and the sample protein sequence information. The protein sequence prediction model includes a structure encoder, a sequence encoder and a corresponding decoder.

[0013] Compared with existing technologies, this application determines the surface point cloud information and sample protein sequence information corresponding to the sample protein based on sample protein information. The sample protein information includes protein sequence information and protein structure information, and the sequence information belonging to the sample protein surface is masked. Based on the surface point cloud information and the sample protein sequence information, a corresponding protein sequence prediction model is trained and obtained. The protein sequence prediction model includes a structure encoder, a sequence encoder, and a corresponding decoder. This application introduces protein surface information for sequence prediction in the construction of the protein sequence prediction model, which can effectively improve the model's prediction accuracy and enhance the surface precision of the predicted protein sequences. Attached Figure Description

[0014] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0015] Figure 1 This diagram illustrates a method for constructing a protein sequence prediction model according to an embodiment of the present application.

[0016] Figure 2 This diagram illustrates a method for training a protein sequence prediction model according to one embodiment of the present application.

[0017] Figure 3 This illustration shows an angle diagram of two points in a surface point cloud according to an embodiment of the present application.

[0018] Figure 4 This diagram illustrates a method for constructing a protein sequence prediction model according to an embodiment of the present application.

[0019] Figure 5This diagram illustrates a device structure for constructing a protein sequence prediction model according to an embodiment of the present application.

[0020] Figure 6 This diagram illustrates a device structure for constructing a protein sequence prediction model according to an embodiment of the present application.

[0021] Figure 7 Exemplary systems that can be used to implement the various embodiments described in this application are shown.

[0022] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation

[0023] The present application will now be described in further detail with reference to the accompanying drawings.

[0024] In a typical configuration of this application, the terminal, the device of the service network, and the trusted party all include one or more processors (e.g., a central processing unit (CPU)), input / output interfaces, network interfaces, and memory.

[0025] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of computer-readable media.

[0026] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0027] The devices referred to in this application include, but are not limited to, user equipment, network equipment, or devices composed of user equipment and network equipment integrated through a network. The user equipment includes, but is not limited to, any mobile electronic product capable of human-computer interaction (e.g., via a touchpad), such as smartphones and tablets. These mobile electronic products can use any operating system, such as Android or iOS. The network equipment includes an electronic device capable of automatically performing numerical calculations and information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices. The network equipment includes, but is not limited to, computers, network hosts, single network servers, multiple network server clusters, or clouds composed of multiple servers. Here, a cloud consists of a large number of computers or network servers based on cloud computing, where cloud computing is a type of distributed computing, consisting of a virtual supercomputer composed of a group of loosely coupled computer clusters. The network includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network, etc. Preferably, the device can also be a program running on the user equipment, network device, or a device formed by integrating user equipment and network device, network device, touch terminal, or network device and touch terminal through a network.

[0028] Of course, those skilled in the art should understand that the above-described devices are merely examples, and other existing or future devices that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0029] In the description of this application, "multiple" means two or more, unless otherwise expressly and specifically defined.

[0030] Figure 1A flowchart of a method for constructing a protein sequence prediction model according to an embodiment of this application is shown. The method includes steps S11 and S12. In step S11, device 1 determines the surface point cloud information and sample protein sequence information corresponding to the sample protein based on sample protein information, wherein the sample protein information includes protein sequence information and protein structure information, and the sequence information belonging to the surface of the sample protein is masked. In step S12, device 1 trains and obtains a corresponding protein sequence prediction model based on the surface point cloud information and the sample protein sequence information, wherein the protein sequence prediction model includes a structure encoder, a sequence encoder, and a corresponding decoder.

[0031] In step S11, device 1 determines the surface point cloud information and sample protein sequence information corresponding to the sample protein based on the sample protein information. The sample protein information includes protein sequence information and protein structure information, and the sequence information belonging to the surface of the sample protein is masked. In some embodiments, device 1 includes, but is not limited to, user devices or network devices with information processing or computing capabilities, such as tablet computers, computers, and servers. In some embodiments, the sample protein information can be obtained from relevant literature or protein databases. The sample protein information includes the protein sequence information and protein structure information corresponding to the sample protein, meaning that the information of the sample protein is completely known. The protein structure information includes the three-dimensional coordinate information of atoms in the sample protein. Device 1 processes the sample protein information, distinguishes between the sample protein surface and interior, extracts the surface point cloud information corresponding to the sample protein, and masks the surface amino acid sequence information for subsequent model training.

[0032] In some embodiments, step S11 includes: device 1 performing surface sampling using a preset virtual sphere based on the protein structure information to determine the surface point cloud information corresponding to the sample protein and the amino acid information located on the surface of the sample protein; and determining the sample protein sequence information based on the amino acid information on the sample protein surface and the protein sequence information. For example, device 1 can set a radius of... A virtual sphere is used to roll around the atoms of the sample protein, recording the points it touches. These points constitute a point cloud on the surface of the sample protein. Device 1 can determine the electrical and polarity characteristics of these points based on the physicochemical features corresponding to the atoms closest to them. The surface point cloud information of the sample protein includes the positional information of each point (e.g., the coordinates of the point or the normal vector of the point towards the outside of the protein surface) and the corresponding physicochemical feature information. Device 1 can determine the amino acids to which these atoms belong based on the atoms closest to them, identifying these amino acids as amino acids on the sample protein surface. Furthermore, based on the amino acid information of the sample protein surface, the amino acids belonging to the sample protein surface in the protein sequence information can be masked to obtain the sample protein sequence information. In some embodiments, the device 1 may use protein structure processing / visualization tools such as PyMol (https: / / pymol.org / ) or PyGAMer (https: / / gamer.readthedocs.io / en / latest / tutorials / notebooks / meshingprotein.html) to sample the surface and generate point clouds.

[0033] In step S12, device 1 trains and obtains a corresponding protein sequence prediction model based on the surface point cloud information corresponding to the sample protein and the sample protein sequence information. The protein sequence prediction model includes a structural encoder, a sequence encoder, and a corresponding decoder. In some embodiments, the structural encoder is used to encode the surface point cloud information corresponding to the sample protein, and the sequence encoder is used to encode the sample protein sequence information. This encoded information is input into the decoder to predict the masked surface amino acid sequence information. The protein sequence prediction model is then optimized by combining the prediction results with known actual sequence information, thereby obtaining a model that can combine protein surface information for protein sequence prediction, making the design of protein surface amino acid sequences more accurate. The protein sequence prediction model described in this scheme is mainly for the case where the protein on one side is known and the protein on the other side is designed in the protein-protein interaction design task. In this design task, the required protein surface shape and corresponding physicochemical characteristics of the interacting protein on the other side, as well as the correspondence between surface point clouds and surface amino acids, can be deduced based on the known protein information. Alternatively, it can be used for tasks requiring extensive modification of known proteins. For example, a protein may bind tightly to receptor A but not to receptor B. To enable the protein to bind to receptor B, the amino acid composition of a specific region of the protein's surface needs to be redesigned based on the known approximate binding conformation. In this task, the internal amino acid sequence information is known, and the corresponding surface shape, physicochemical characteristics, and the correspondence between surface point clouds and surface amino acids can be inferred based on the original characteristics of receptor B and the protein. Therefore, even when the types of amino acids on the protein surface are unknown, the surface information of these proteins can be provided to the model as known information to make more accurate predictions of the amino acid sequence, enabling the predicted protein to interact better with other proteins or small molecules. It should be understood by those skilled in the art that the above model application scenario is merely an example, and other existing or future protein design tasks with the same or similar needs and providing the same or similar model input information, if applicable to this application, should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0034] In some embodiments, reference Figure 2The flowchart shown illustrates the training process for the protein sequence prediction model. Step S12 includes: Step S121, where device 1 determines the corresponding surface coding information of the sample protein using the structure encoder based on the surface point cloud information corresponding to the sample protein; Step S122, where device 1 determines the corresponding sample protein sequence coding information using the sequence encoder based on the sample protein sequence information; Step S123, where device 1 determines the protein sequence prediction information corresponding to the sample protein using the decoder based on the sample protein surface coding information and the sample protein sequence coding information; and Step S124, where device 1 optimizes the protein sequence prediction model based on the protein sequence prediction information and the protein sequence information. The execution order of steps S121 and S122 is not limited; they can be executed simultaneously or sequentially.

[0035] In some embodiments, the structural encoder includes a graph neural network. Device 1 uses the graph neural network to capture the feature information of each point and its neighboring points in the surface point cloud information corresponding to the sample protein, and updates and integrates the point in the surface point cloud information corresponding to the sample protein based on the feature information of each point and its neighboring points to obtain the corresponding sample protein surface encoding information.

[0036] In some embodiments, step S121 includes: step S1211, whereby device 1, based on the surface point cloud information corresponding to the sample protein, uses the structure encoder to determine the surface feature information corresponding to the surface point cloud information corresponding to the sample protein, wherein the surface feature information includes feature information corresponding to each point in the surface point cloud information corresponding to the sample protein; step S1212, whereby device 1, based on the surface feature information, determines the corresponding sample protein surface encoding information through information integration, wherein the sample protein surface encoding information includes feature information corresponding to each amino acid on the surface of the sample protein.

[0037] In some embodiments, the process of determining the surface feature information corresponding to the surface point cloud information of the sample protein using a structural encoder is as follows: For any point i in the surface point cloud information corresponding to the sample protein and its neighboring point j, the information transmitted between point i and j can be determined based on pre-constructed feature information. The feature information of point i is then updated based on the information transmitted between point i and j. In some embodiments, to improve prediction accuracy, an attention mechanism can be used to weight the information transmitted between point i and j, giving higher weight to information that is more strongly associated with and more important than point i, and then combining the weighted information to update the feature information of point i. In some embodiments, as described in the foregoing embodiments, in the protein sequence design / modification task targeted by this solution, although the specific types of amino acids on the protein surface are unknown, the surface shape formed by these surface amino acids and the corresponding physicochemical characteristics of the surface can still be obtained. Therefore, even if the types of amino acids are unknown, the surface amino acids to which a point in the point cloud belongs can be determined according to the principle of closest proximity. Neighboring point j belongs to the same surface amino acid as point i, or the surface amino acid to which neighboring point j belongs is one of the multiple surface amino acids that are closest to the surface amino acid to which point i belongs. The number of the nearest surface amino acids can be set based on actual task requirements; preferably, this number can be set to 7.

[0038] In some embodiments, the feature information corresponding to each point in the surface point cloud information corresponding to the sample protein includes at least one of the following: physicochemical feature information corresponding to each point; curvature feature corresponding to each point; and orientation feature between each point and any neighboring points. In some embodiments, the curvature feature corresponding to each point includes pseudo-curvature vectors, which provide features related to rotation invariance. In some embodiments, the orientation feature between each point and any neighboring points includes the orientation vector between the point and the neighboring points, and angular feature information determined based on the position information corresponding to the point and the neighboring points.

[0039] In some embodiments, a corresponding covariance matrix can be constructed based on the surface amino acid to which any point i belongs in the surface point cloud information and the multiple surface amino acids closest to that surface amino acid.

[0040]

[0041] Where, N (i) Let i be the surface amino acid to which point i belongs, and the set of multiple surface amino acids that are closest to that surface amino acid. x is the centroid of the surface amino acid to which point i belongs and the set of multiple surface amino acids closest to that surface amino acid. j For N(i) The position information corresponding to the midpoint j. Calculate the eigenvalues ​​of this covariance matrix. Based on this eigenvalue, the corresponding pseudo-curvature vector (ψ1,ψ2,ψ3) is obtained, where,

[0042]

[0043] In some embodiments, reference Figure 3 The diagram shown illustrates the angle, based on the surface point cloud information of point i, its neighboring point j, and the normal vector n of these two points towards the outer surface of the protein. i ,n j Determine the normal vector n corresponding to point i. i The angle information between the two points, and the normal vector n corresponding to point j. j With the angle information between the two points and the normal vector n i With normal vector n j Angular feature information such as inter-angle information

[0044]

[0045] Based on the above feature information, the information transmitted between point i and its neighboring point j can be determined.

[0046]

[0047] In some embodiments, an attention mechanism can also be used to perform weighted calculations on the information transmitted between point i and its neighboring point j.

[0048]

[0049] Then, by combining the information transmitted above, the feature information of point i is updated, and the surface feature information corresponding to the surface point cloud information of the corresponding sample protein is obtained.

[0050]

[0051] Among them, f m f h f x Refers to three different multilayer perceptrons (MLPs). This refers to the physicochemical characteristics contained in point i in the l-th layer of the network. It refers to the direction vector from point i to point j in the point cloud of the l-th layer network.

[0052] In some embodiments, due to the flexibility of the atoms on the protein surface, their frequent and minute movements can introduce significant noise into the protein surface determined by the virtual sphere surface sampling. To reduce the noise of the surface point cloud information corresponding to the sample protein and improve prediction accuracy, before step S1211, step S121 further includes: the device 1 smoothing the surface point cloud information corresponding to the sample protein. The smoothed surface point cloud information corresponding to the sample protein is then input into the structure encoder for processing. In some embodiments, the smoothing process can use Gaussian kernel smoothing or nearest neighbor smoothing, etc. Those skilled in the art should understand that the above smoothing methods are merely examples, and other existing or future smoothing methods that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0053] In some embodiments, after updating the feature information corresponding to each point in the surface point cloud information of the sample protein through the structural encoder, the updated feature information corresponding to each point in the surface point cloud information of the sample protein is integrated into the corresponding surface amino acid. The feature information corresponding to the surface amino acid of the sample protein includes the feature information corresponding to one or more points in the surface point cloud information of the sample protein, where the surface amino acid is closest to the one or more points compared to other surface amino acids. Here, as described in the foregoing embodiments, the specific types of surface amino acids are unknown. However, surface amino acids not only have their own feature information but also the feature information of surrounding amino acids, allowing for a more comprehensive consideration in subsequent sequence design, resulting in better design accuracy.

[0054] In some embodiments, step S1212 includes: device 1 determining the surface amino acid to which each point in the surface point cloud information corresponding to the sample protein belongs; and determining the feature information corresponding to each amino acid on the surface of the sample protein based on the surface amino acid to which each point in the surface point cloud information corresponding to the sample protein and the surface feature information. For example, determining the atom closest to a point in the surface point cloud information corresponding to the sample protein. Based on the surface amino acid to which the atom belongs, determining the surface amino acid to which the point belongs. Then, integrating the feature information corresponding to the point into the feature information of the surface amino acid.

[0055] In some embodiments, the sequence encoder includes a protein language model; preferably, a protein language model with a multi-layer self-attention mechanism can be used. The sequence encoder encodes sample protein sequence information that does not contain surface amino acid sequence information to obtain an embedding representation of the input sample protein sequence information, thereby capturing the contextual relationships of amino acids in the sequence.

[0056] In some embodiments, device 1 can use a protein language model as a decoder to predict the amino acid types on the surface of the masked sample protein based on the sample protein surface encoding information and the sample protein sequence encoding information, generating corresponding protein sequence prediction information. Here, the decoder can predict the surface amino acid types not only based on internal amino acid information but also by combining the sample protein surface encoding information, i.e., the feature information corresponding to the surface amino acids to be predicted. With more available information, the prediction accuracy is higher. After obtaining the protein sequence prediction information, a corresponding loss function (e.g., cross-entropy loss function or Focal Loss, etc., suitable for protein sequence design tasks) is used to calculate the loss between the protein sequence prediction information and the actual protein sequence information corresponding to the sample protein. The gradient of the loss function with respect to the model parameters is calculated using the backpropagation algorithm. The protein sequence prediction model parameters are updated using a corresponding optimization algorithm (e.g., gradient descent or Adam, RMSprop optimizer, etc.). Device 1 can repeat the above steps S121-S124 until the model performance no longer improves or reaches a predetermined number of training epochs, thereby obtaining the final protein sequence prediction model.

[0057] In some embodiments, reference Figure 4The flowchart shown further illustrates that the method includes: Step S13, based on one or more target protein information corresponding to the target protein, using the protein sequence prediction model to determine the corresponding target protein sequence information, wherein the target protein information includes target protein surface information and target protein sequence association information. In some embodiments, the protein sequence prediction model can be trained by device 1 and directly applied, or it can be deployed to other devices for protein sequence modification / design. In some embodiments, as described in the preceding embodiments for protein modification / design tasks, for the target protein to be predicted, the sequence length / length range of its surface amino acids can be estimated to obtain one or more target protein information. For example, for tasks involving the modification of existing proteins, the corresponding sequence length / length range can be determined based on the length of its original surface amino acid sequence (e.g., by adding or subtracting a certain number of times from the original length). For protein design tasks, the length of the surface amino acid sequence can be estimated based on a known side of the protein. The target protein surface information includes the predicted surface point cloud information of the target protein. The target protein sequence association information includes the internal amino acid sequence information of the target protein and the length information of the surface amino acid sequence. The length information of the surface amino acid sequence corresponding to different target protein information is different. The most suitable target protein sequence information can be determined from the prediction results of these target protein information based on the protein sequence prediction model. Here, the process of predicting the target protein sequence using the protein sequence prediction model based on the target protein information is similar to the aforementioned steps S121-S123, and therefore will not be repeated here, but is included in the form of a reference.

[0058] In some embodiments, step S13 further includes using a generative simulation method to determine one or more target protein information corresponding to the target protein. For example, for protein design tasks, diffusion-based models such as RFDiffusion (https: / / github.com / RosettaCommons / RFdiffusion, an open-source protein structure generation method) can be used to predict the length of the amino acid sequence on the surface of the target protein, and then, combined with the estimated length / length range (e.g., by adding or subtracting a certain number of times based on the estimated length), the target protein information corresponding to different surface amino acid sequence lengths can be determined.

[0059] In some embodiments, step S13 includes determining candidate target protein sequence information corresponding to each target protein information using the protein sequence prediction model based on one or more target protein information corresponding to the target protein; and determining the corresponding target protein sequence information based on the candidate target protein sequence information. In some embodiments, the protein sequence prediction model also outputs confidence information corresponding to its predicted candidate target protein sequence information. Based on the confidence information, sequences that meet certain conditions (e.g., one or more candidate target protein sequence information with the highest confidence information or greater than a corresponding threshold) can be selected as target protein sequence information from the candidate target protein sequence information. In some embodiments, perplexity can be used as an evaluation metric for the prediction results of the protein sequence prediction model to obtain confidence information about the predicted candidate target protein sequence information.

[0060] Figure 5 This diagram illustrates a device structure for constructing a protein sequence prediction model according to an embodiment of this application. The device 1 includes a primary module 11 and a secondary module 12. The primary module 11 determines the surface point cloud information and protein sequence information corresponding to a sample protein based on sample protein information. The sample protein information includes protein sequence information and protein structure information, and the sequence information belonging to the surface of the sample protein is masked. The secondary module 12 trains and acquires a corresponding protein sequence prediction model based on the surface point cloud information and the sample protein sequence information. The protein sequence prediction model includes a structure encoder, a sequence encoder, and a corresponding decoder. Figure 5 The specific implementation methods corresponding to module 11 and module 12 shown are the same as or similar to the specific embodiments of steps S11 and S12 described above, so they will not be repeated here, but are included by reference.

[0061] In some embodiments, the first and second modules 12 include a first and second unit 121, a first and second unit 122, a first and second third unit 123, and a first and second fourth unit 124. The first and second unit 121 determines the corresponding sample protein surface coding information using the structure encoder based on the surface point cloud information corresponding to the sample protein; the first and second unit 122 determines the corresponding sample protein sequence coding information using the sequence encoder based on the sample protein sequence information; the first and second third unit 123 determines the protein sequence prediction information corresponding to the sample protein using the decoder based on the sample protein surface coding information and the sample protein sequence coding information; and the first and second fourth unit 124 optimizes the protein sequence prediction model based on the protein sequence prediction information and the protein sequence information. Here, the specific implementations of Unit 121, Unit 122, Unit 123, and Unit 124 are the same as or similar to the specific embodiments of steps S121, S122, S123, and S124 mentioned above, and will not be repeated here, but are included by reference.

[0062] In some embodiments, the first-two-one unit 121 includes a second-two-one subunit 1211 and a second-two-one subunit 1212. The second-two-one subunit 1211, based on the surface point cloud information corresponding to the sample protein, uses the structure encoder to determine the surface feature information corresponding to the surface point cloud information of the sample protein, wherein the surface feature information includes feature information corresponding to each point in the surface point cloud information of the sample protein; the second-two-one subunit 1212, based on the surface feature information, determines the corresponding sample protein surface encoding information through information integration, wherein the sample protein surface encoding information includes feature information corresponding to each amino acid on the surface of the sample protein. Here, the specific implementations of the second-two-one subunit 1211 and the second-two-one subunit 1212 are the same as or similar to the specific embodiments of steps S1211 and S1212 described above, and therefore will not be repeated here, but are incorporated herein by reference.

[0063] In some embodiments, reference Figure 6 The device structure diagram shown includes a three-module 13. Based on one or more target protein information corresponding to the target protein, the three-module 13 uses the protein sequence prediction model to determine the corresponding target protein sequence information. The target protein information includes target protein surface information and target protein sequence association information. The specific implementation of the three-module 13 is the same as or similar to the specific embodiment of step S13 described above, and therefore will not be repeated here, but is incorporated herein by reference.

[0064] Figure 7Exemplary systems that can be used to implement the various embodiments described in this application are shown. Figure 7 As shown in some embodiments, system 300 can function as any of the devices described in each of the embodiments. In some embodiments, system 300 may include one or more computer-readable media having instructions (e.g., system memory or NVM / storage device 320) and one or more processors (e.g., one or more processors 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement the module and thus perform the actions described in this application.

[0065] In one embodiment, the system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of the processors 305 and / or any suitable device or component communicating with the system control module 310.

[0066] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.

[0067] System memory 315 can be used, for example, to load and store data and / or instructions for system 300. In one embodiment, system memory 315 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, system memory 315 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).

[0068] In one embodiment, the system control module 310 may include one or more input / output (I / O) controllers to provide interfaces to the NVM / storage device 320 and (one or more) communication interfaces 325.

[0069] For example, NVM / storage device 320 can be used to store data and / or instructions. NVM / storage device 320 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives).

[0070] NVM / storage device 320 may include storage resources that are physically part of a device on which system 300 is mounted, or that can be accessed by the device without necessarily being part of it. For example, NVM / storage device 320 may be accessed via a network through one or more communication interfaces 325.

[0071] One or more communication interfaces 325 may provide the system 300 with an interface to communicate over one or more networks and / or with any other suitable device. The system 300 may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.

[0072] In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 (e.g., memory controller module 330). In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 to form a system-in-package (SiP). In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die. In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die to form a system-on-a-chip (SoC).

[0073] In various embodiments, system 300 may be, but is not limited to, a server, workstation, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, system 300 may have more or fewer components and / or different architectures. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0074] In addition to the methods and devices described in the above embodiments, this application also provides a computer-readable storage medium storing computer code that, when executed, performs the method described in any of the preceding embodiments.

[0075] This application also provides a computer program product that, when executed by a computer device, performs the method described in any of the preceding claims.

[0076] This application also provides a computer device, the computer device comprising:

[0077] One or more processors;

[0078] Memory, used to store one or more computer programs;

[0079] When the one or more computer programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described in any of the preceding methods.

[0080] It should be noted that this application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, and similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0081] Furthermore, a portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0082] Communication media include media through which communication signals containing, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media can include guided transmission media (such as cables and wires (e.g., optical fibers, coaxial cables, etc.)) and wireless (unguided transmission) media capable of propagating energy waves, such as sound, electromagnetic, RF, microwave, and infrared. Computer-readable instructions, data structures, program modules, or other data can be embodied as modulated data signals in, for example, wireless media (such as carrier waves or similar mechanisms embodied as part of spread spectrum technology). The term "modulated data signal" refers to a signal whose one or more characteristics are altered or set in a manner that encodes information in the signal. Modulation can be analog, digital, or a hybrid modulation technique.

[0083] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memories such as random access memory (RAM, DRAM, SRAM); and non-volatile memories such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or hereafter developed capable of storing computer-readable information / data for use by a computer system.

[0084] Herein, one embodiment of this application includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to run a method and / or technical solution based on the foregoing embodiments of this application.

[0085] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.

Claims

1. A method for constructing a protein sequence prediction model, wherein, The method includes: Based on the sample protein information, the surface point cloud information and sample protein sequence information corresponding to the sample protein are determined. The sample protein information includes protein sequence information and protein structure information. The sequence information belonging to the sample protein surface in the sample protein sequence information is masked. The surface point cloud information includes the position information corresponding to each point in the sample protein surface point cloud and the corresponding physicochemical feature information. Based on the surface point cloud information corresponding to the sample protein and the sample protein sequence information, a corresponding protein sequence prediction model is trained and obtained, wherein the protein sequence prediction model includes a structure encoder, a sequence encoder and a corresponding decoder. The process of training a corresponding protein sequence prediction model based on the surface point cloud information and the protein sequence information of the sample protein includes: training a structural encoder, a sequence encoder, and a corresponding decoder; determining the corresponding sample protein surface encoding information using the structural encoder based on the surface point cloud information of the sample protein; determining the corresponding sample protein sequence encoding information using the sequence encoder based on the sample protein sequence information, wherein the sequence encoder is used to encode sample protein sequence information that does not contain surface amino acid sequence information; determining the protein sequence prediction information corresponding to the sample protein using the decoder based on the sample protein surface encoding information and the sample protein sequence encoding information, wherein the decoder is used to predict the amino acid types on the masked sample protein surface; and optimizing the protein sequence prediction model based on the protein sequence prediction information and the protein sequence information.

2. The method according to claim 1, wherein, The step involves determining the surface point cloud information and sample protein sequence information corresponding to the sample protein based on the sample protein information. The sample protein information includes protein sequence information and protein structure information. The masking of sequence information belonging to the sample protein surface includes: Based on the protein structure information, surface sampling is performed using a preset virtual sphere to determine the surface point cloud information corresponding to the sample protein and the amino acid information located on the surface of the sample protein. The protein sequence information of the sample is determined based on the amino acid information on the surface of the sample protein and the protein sequence information.

3. The method according to claim 1, wherein, The step of determining the corresponding sample protein surface encoding information using the structure encoder based on the surface point cloud information corresponding to the sample protein includes: Based on the surface point cloud information corresponding to the sample protein, the surface feature information corresponding to the surface point cloud information corresponding to the sample protein is determined using the structure encoder, wherein the surface feature information includes the feature information corresponding to each point in the surface point cloud information corresponding to the sample protein. Based on the surface feature information, the corresponding sample protein surface coding information is determined through information integration, wherein the sample protein surface coding information includes the feature information corresponding to each amino acid on the sample protein surface.

4. The method according to claim 3, wherein, Before determining the surface feature information corresponding to the surface point cloud information corresponding to the sample protein using the structure encoder based on the surface point cloud information corresponding to the sample protein, wherein the surface feature information includes the feature information corresponding to each point in the surface point cloud information corresponding to the sample protein, the determination of the corresponding sample protein surface encoding information using the structure encoder based on the surface point cloud information corresponding to the sample protein further includes: The surface point cloud information corresponding to the sample protein is smoothed.

5. The method according to claim 3 or 4, wherein, The feature information corresponding to each point in the surface point cloud information of the sample protein includes at least one of the following: the physicochemical feature information corresponding to each point; the curvature feature corresponding to each point; and the directional feature between each point and any neighboring points.

6. The method according to claim 3 or 4, wherein, Based on the surface feature information, the corresponding sample protein surface coding information is determined through information integration. The sample protein surface coding information includes feature information corresponding to each amino acid on the sample protein surface, including: Determine the surface amino acid associated with each point in the surface point cloud information corresponding to the sample protein; Based on the surface amino acid to which each point belongs in the surface point cloud information corresponding to the sample protein and the surface feature information, the feature information corresponding to each amino acid on the surface of the sample protein is determined.

7. The method according to claim 1, wherein, The method further includes: Based on one or more target protein information corresponding to the target protein, the corresponding target protein sequence information is determined using the protein sequence prediction model, wherein the target protein information includes target protein surface information and target protein sequence association information.

8. The method according to claim 7, wherein, The method involves using the protein sequence prediction model to determine the corresponding target protein sequence information based on one or more target protein information corresponding to the target protein. The target protein information includes target protein surface information and target protein sequence association information, and further includes: Generative simulation methods are used to determine one or more target protein information corresponding to the target protein.

9. The method according to claim 7 or 8, wherein, The method involves using the protein sequence prediction model to determine the corresponding target protein sequence information based on one or more target protein information corresponding to the target protein, wherein the target protein information includes target protein surface information and target protein sequence association information, including: Based on one or more target protein information corresponding to the target protein, the candidate target protein sequence information corresponding to each target protein information is determined using the protein sequence prediction model. Based on the candidate target protein sequence information, the corresponding target protein sequence information is determined.

10. A computer device for constructing a protein sequence prediction model, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 9.

11. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 9.

12. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Sequence prediction method and device, medium and electronic equipment

    CN115662517A

  • Method and device for training protein prediction model based on graph neural network

    CN116935952A