A processing method and device for predicting a self-assembling short peptide model

Through an end-to-end self-assembly short peptide prediction model, multiple machine learning models and PACE force field are used to simulate the self-assembly of polypeptide molecules, which solves the problems of long experimental research cycle and high cost, and achieves efficient and accurate short peptide prediction.

CN117316305BActive Publication Date: 2025-09-26PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311246066.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2025-09-26
Estimated Expiration
2043-09-25

AI Technical Summary

Technical Problem

In the existing technology, the research on the self-assembly process of polypeptide molecules relies on experimental methods, which has the problems of long research cycle, unpredictable costs and reliance on manual experience, resulting in a low probability of discovering new self-assembling short peptides.

Method used

An end-to-end self-assembly short peptide prediction model is adopted, multiple machine learning models and the molecular dynamics simulation module of the PACE force field are used to simulate the self-assembly process of polypeptide molecules through a synchronous learning strategy to generate a short peptide prediction report.

Benefits of technology

Shorten the research cycle, reduce research costs, increase the probability of discovering new self-assembling short peptides, reduce dependence on manual experience, and improve prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117316305B_ABST
    Figure CN117316305B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention relates to a method and device for processing a self-assembling short peptide prediction model, the method comprising: constructing a self-assembling short peptide prediction model; receiving a first polypeptide molecule set, a first simulation number threshold, and a first output structure type input by a user; setting the force field type parameter of the molecular dynamics simulation module of the self-assembling short peptide prediction model to the PACE force field, setting the simulation number threshold parameter of the self-assembly index evaluation module to the corresponding first simulation number threshold, and setting the output structure morphology parameter of the self-assembling short peptide output module to the corresponding first output structure type; and performing a self-assembling short peptide chain structure prediction based on the first polypeptide molecule set by the self-assembling short peptide prediction model to obtain a corresponding first short peptide prediction report. The present invention can shorten the research cycle and increase the probability of discovering new self-assembling short peptides.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method and device for processing a self-assembling short peptide prediction model. Background Art

[0002] Peptides, molecules based on a variety of naturally occurring amino acids, are widely used in the study of supramolecular materials. Short peptides, also known as oligopeptides, are peptide molecules with a relatively small number of amino acids. Each peptide molecule is considered to form a new peptide molecule when the number or order of amino acids changes. Self-assembly, also known as supramolecular polymerization, refers to the process by which molecules spontaneously combine to form specific structures and functions without external intervention. This process is driven by non-covalent interactions between molecules, resulting in changes in properties such as the number and order of amino acids in the peptide molecule. Peptide self-assembly refers to the phenomenon in which different peptides self-assemble to form materials with specific morphologies through non-covalent interactions between side chains and backbone atoms. Currently, research on peptide self-assembly relies primarily on experimental methods, which have significant technical drawbacks: long experimental research cycles, unpredictable research costs, and the limitations of human experience, which significantly reduce the probability of discovering new self-assembling short peptides. Summary of the Invention

[0003] The purpose of the present invention is to provide a processing method, device, electronic device and computer-readable storage medium for a self-assembling short peptide prediction model to address the defects of the prior art; provide an end-to-end self-assembling short peptide prediction model to simulate the self-assembly process of polypeptide molecules and output a short peptide prediction report consisting of multiple short peptide prediction records; the self-assembling short peptide prediction model is characterized in that: multiple similar or different types of machine learning models are used to predict the self-assembly ability of each polypeptide molecule in the polypeptide molecule set and a candidate molecule set is formed by one or more polypeptide molecules with high probability, and the candidate molecule set is simulated by a molecular dynamics simulation module whose force field condition is specifically set to a PACE force field. During the simulation process, each machine learning model is trained based on a synchronous learning strategy to reduce the difficulty of training data collection, reduce the difficulty of model training, and improve the accuracy of model prediction. The present invention can achieve the purpose of shortening the research cycle, reducing and quantifying research costs, and can also get rid of excessive reliance on manual experience to achieve the purpose of increasing the probability of discovering new self-assembling short peptides.

[0004] To achieve the above-mentioned object, a first aspect of an embodiment of the present invention provides a method for processing a self-assembling short peptide prediction model, the method comprising:

[0005] Constructing a self-assembling short peptide prediction model; the self-assembling short peptide prediction model includes a molecular dynamics simulation module, a self-assembly index evaluation module and a self-assembling short peptide output module; the molecular dynamics simulation module is preset with a configurable force field type parameter; the self-assembly index evaluation module is preset with a configurable simulation number threshold parameter; the self-assembly short peptide output module is preset with a configurable output structure morphology parameter;

[0006] Receiving a first polypeptide molecule set, a first simulation number threshold, and a first output structure type input by a user; the first polypeptide molecule set includes a plurality of first polypeptide molecules; the molecular structure types of the first polypeptide molecules include one-dimensional structure, two-dimensional structure, and three-dimensional structure; the first output structure type includes one-dimensional structure, two-dimensional structure, and three-dimensional structure;

[0007] The force field type parameter of the molecular dynamics simulation module of the self-assembly short peptide prediction model is set to the PACE force field, the simulation number threshold parameter of the self-assembly index evaluation module is set to the corresponding first simulation number threshold, and the output structure morphology parameter of the self-assembly short peptide output module is set to the corresponding first output structure type;

[0008] The self-assembling short peptide prediction model predicts the structure of the self-assembling short peptide chain based on the first polypeptide molecule set to obtain a corresponding first short peptide prediction report.

[0009] Preferably, the self-assembling short peptide prediction model includes an initialization module, a feature encoding module, a self-training module, a first number N of machine learning models M i , a self-assembly molecule screening module, the molecular dynamics simulation module, the self-assembly index evaluation module and the self-assembly short peptide output module, the first number N is an integer not less than 2, 1≤model index i≤first number N;

[0010] The initialization module is respectively connected to the input end of the self-assembly short peptide prediction model, the feature encoding module, and each of the machine learning models M i , the self-assembly molecular screening module and the molecular dynamics simulation module are connected; each of the machine learning models M i are respectively connected to the self-training module, the feature encoding module and the self-assembly molecular screening module; the self-assembly molecular screening module is connected to the molecular dynamics simulation module; the molecular dynamics simulation module is connected to the self-assembly index evaluation module; the self-assembly index evaluation module is respectively connected to the self-training module and the self-assembly short peptide output module; the self-assembly short peptide output module is connected to the output end of the self-assembly short peptide prediction model;

[0011] The machine learning model M iIt includes a first main control unit and a first prediction model; the first main control unit is respectively connected to the initialization module, the self-training module, the feature encoding module, the first prediction model and the self-assembly molecular screening module; the model structure of the first prediction model includes a linear regression model structure, a kernel function model structure, a support vector machine model structure, a decision tree model structure, a random forest model structure, a gradient boosting model structure, an XGBoost model structure, a LightGBM model structure and a deep learning model structure; every two of the machine learning models M i The first main control unit is the same as each of the two machine learning models M i The first prediction models are not forced to be the same; each of the two machine learning models M i The initial model parameters of the first prediction model and the model parameters obtained after training are different;

[0012] The self-training module is pre-set with a self-training data set; the self-training data set includes a plurality of first self-training data; the first self-training data includes a first training polypeptide molecule and a first label indicator sequence;

[0013] An evaluation record list is preset on the self-assembly index evaluation module; the evaluation record list includes multiple first evaluation records; the first evaluation record includes a first original molecule set field, a first candidate molecule set field, a first polymer set field and a first evaluation index set field.

[0014] Preferably, the self-assembling short peptide prediction model is used to input the polypeptide molecule set {x j} performs self-assembly short peptide structure prediction and outputs the corresponding short peptide prediction report; the polypeptide molecule set {x j} includes multiple polypeptide molecules x j , the molecular index j is an integer greater than 0; the polypeptide molecule x j The structural types include one-dimensional structure, two-dimensional structure and three-dimensional structure; the short peptide prediction report includes multiple first short peptides; the structural types of the first short peptides correspond to the output structural morphological parameters, including one-dimensional structure, two-dimensional structure and three-dimensional structure.

[0015] Furthermore, the initialization module is used to receive the input polypeptide molecule set {x j}; and each of the polypeptide molecules x in the collection j Input the feature encoding module to encode and obtain the corresponding first encoding tensor e j ; and all the first encoding tensors e obtained by j The first coding sequence corresponding to {e j}; and the polypeptide molecule set {x j} as the corresponding first original molecule set X init and the first candidate molecule set X 1,opt ; and the first original molecule set X init , the first candidate molecule set X 1,opt and the first coding sequence {e j} input to the molecular dynamics simulation module; and the first original molecule set X init To each of the said machine learning models M i The first main control unit sends the first original molecule set X init and the first coding sequence {e j}Sending to the self-assembly molecular screening module;

[0016] The feature encoding module is used to perform feature encoding on the input one-dimensional, two-dimensional or three-dimensional molecules and output the corresponding encoding tensor;

[0017] The molecular dynamics simulation module is used to receive the first original molecule set X sent by the initialization module init , the first candidate molecule set X 1,opt and the first coding sequence {e j}, for the first original molecule set X init and initialize the count value of the preset simulation counter to 0; and based on the preset molecular dynamics simulation input file format, the first coding sequence {e j}Convert the simulation input file to obtain the corresponding current simulation input file; call the preset molecular dynamics simulation interface to perform molecular dynamics simulation according to the force field type parameters and the current simulation input file, and continuously time this simulation; and identify whether an aggregation effect occurs in the current simulation process when the simulation timing reaches a preset first short-time threshold. If no aggregation effect occurs, terminate this simulation in advance; if an aggregation effect occurs, continue the simulation and terminate this simulation normally when the simulation timing reaches a preset first long-time threshold; and when this simulation terminates normally, use the count value of the simulation times counter as the corresponding first count value C1; and use the simulation trajectory file output by this simulation as the corresponding first trajectory file; and perform cluster structure identification on the first trajectory file and use each identified cluster structure as the corresponding first aggregate p k , the aggregate index k is an integer greater than 0; and all the first aggregates p obtained this time k The first count value C1, the first trajectory file, the first original molecule set X init , the first candidate molecule set X 1,optand the first polymer set P1 to the self-assembly index evaluation module;

[0018] The self-assembly index evaluation module is used to receive the first count value C1, the first trajectory file, the first original molecule set X init , the first candidate molecule set X 1,opt and the first aggregate set P1, each of the first aggregates p in the first aggregate set P1 is tracked based on the first trajectory file. k Perform multiple self-assembly index calculations to obtain the corresponding first polymer index sequence E k ; and all the first polymer index sequences E obtained k The first evaluation index set GE1 is formed; and the first original molecule set X init , the first candidate molecule set X 1,opt , the first polymer set P1 and the first evaluation index set GE1 are used as the corresponding first original molecule set field, the first candidate molecule set field, the first polymer set field and the first evaluation index set field to form a corresponding first evaluation record and add it to the local evaluation record list; and judge whether the first count value C1 is less than the simulation number threshold parameter; if the first count value C1 is greater than or equal to the simulation number threshold parameter, the evaluation record list is extracted as the corresponding first evaluation record list and sent to the self-assembly short peptide output module, and the local evaluation record list is cleared; if the first count value C1 is less than the simulation number threshold parameter, each of the first polymers p k and the corresponding first polymer index sequence E k As a set of corresponding first training polypeptide molecules and first label indicator sequences, the obtained first training polypeptide molecules and the first label indicator sequences form a corresponding first self-training data set, and all the obtained first self-training data forms a corresponding first batch data set added to the self-training data set of the self-training module; the first evaluation index set GE1 includes multiple first polymer indicator sequences E k ; The first polymer index sequence E k Including multiple first aggregate indicators e k,h , the index h is an integer greater than 0; the first aggregate index e k,h The index types include aggregation strength index, water solubility index and structural order index;

[0019] The self-training module is used to extract all the first self-training data of the current self-training data set to form a corresponding first self-training data set each time a first batch of data sets is added to the self-training data set; and to train each of the machine learning models M according to the first self-training data set. i The first prediction model performs a round of model training; and all the machine learning models M i At the end of this round of training, each of the machine learning models M i The first main control unit sends a corresponding single-round training end flag;

[0020] The machine learning model M i The first main control unit is used to receive the first original molecule set X sent by the initialization module init Save it when

[0021] The machine learning model M i The first main control unit is further configured to, upon receiving the single-round training end flag sent by the self-training module, convert the first original molecule set X init Each of the polypeptide molecules x j Input the feature encoding module to encode and obtain the corresponding second encoding tensor e i,j ; and all the second encoding tensors e obtained by i,j The corresponding second coding sequence {e i,j}; and the second coding sequence {e i,j}Send to the first prediction model, and receive the first prediction indicator set GE sent back by the first prediction model p,i ; and the second coding sequence {e i,j} and the first set of predictive indicators GE p,i Form a corresponding first learning data D i Sending to the self-assembly molecular screening module;

[0022] The machine learning model M i The first prediction model is used to receive the second coding sequence {e i,j}, according to each of the second encoding tensors e in the sequence i,j Perform multi-class self-assembly index prediction to obtain the corresponding first prediction index sequence E i,j , and all the first prediction indicator sequences E are obtained by i,j The first prediction indicator set GE p,i Send back to the first main control unit; the first prediction indicator set GE p,i Including a plurality of the first prediction indicator sequences E i,j; The first prediction indicator sequence E i,j Including multiple first prediction indicators e i,j,h The first prediction indicator e i,j,h The index types include aggregation strength index, water solubility index and structural order index;

[0023] The self-assembly molecular screening module is used to receive the first original molecular set X sent by the initialization module init and the first coding sequence {e j} when saving both;

[0024] The self-assembly molecular screening module is also used to receive all the learning models M i The first learning data D sent i Afterwards, all the first learning data D obtained this time i Combining a corresponding first learning data set; and integrating the first learning data set with the first original molecule set X init Each of the polypeptide molecules x j Corresponding to all the first prediction indicators e i,j,h Extracted to form a corresponding first molecule observation vector b j , and all the first molecular observation vectors b are obtained by j The first observation tensor B is formed; and the first observation tensor B is substituted into the preset acquisition function equation for high probability evaluation to obtain a plurality of first evaluation data v j The first evaluation vector V is composed of the first evaluation data v j The polypeptide molecule x is not less than the preset qualification assessment threshold j Denote as the corresponding first candidate molecule x opt,w , the molecular index w is an integer greater than 0; and all the first candidate molecules x are obtained opt,w The corresponding second candidate molecule set X 2,opt ; and the first coding sequence {e j} and each of the first candidate molecules x opt,w The corresponding first encoding tensor e j As the corresponding third encoding tensor e w , and all the third encoded tensors e are obtained by w The corresponding third coding sequence {e w}; and the second candidate molecule set X 2,opt and the third coding sequence {e w} is sent to the molecular dynamics simulation module; the first evaluation vector V includes a plurality of the first evaluation data v j, the first evaluation data v j With the first molecular observation vector b j One-to-one correspondence;

[0025] The molecular dynamics simulation module is further configured to receive the second candidate molecule set X sent by the self-assembly molecular screening module. 2,opt and the third coding sequence {e w}, the count value of the simulation times counter is increased by 1; and the third coding sequence {e w}Convert the simulation input file to obtain the corresponding current simulation input file; call the molecular dynamics simulation interface to perform molecular dynamics simulation according to the force field type parameters and the current simulation input file, and continuously time this simulation; and identify whether an aggregation effect occurs in the current simulation process when the simulation timing reaches the first short-term threshold. If no aggregation effect occurs, terminate this simulation in advance; if an aggregation effect occurs, continue the simulation and terminate this simulation normally when the simulation timing reaches the first long-term threshold; and when this simulation terminates normally, use the count value of the simulation times counter as the corresponding second count value C2; ​​and use the simulation trajectory file output by this simulation as the corresponding second trajectory file; and perform cluster structure identification on the second trajectory file and use each identified cluster structure as the corresponding second aggregate p r , the aggregate index r is an integer greater than 0; and all the second aggregates p obtained this time r The corresponding second polymer set P2 is formed; the second count value C2, the second trajectory file, the first original molecule set X init , the second candidate molecule set X 2,opt and the second polymer set P2 to the self-assembly index evaluation module;

[0026] The self-assembly index evaluation module is further configured to receive the second count value C2, the second trajectory file, and the first original molecule set X init , the second candidate molecule set X 2,opt and the second aggregate set P2, each second aggregate p in the second aggregate set P2 is generated based on the second trajectory file. r Calculate multiple types of self-assembly indices to obtain the corresponding second polymer index sequence E r ; and all the second polymer index sequences E obtained r The corresponding second evaluation index set GE2 is formed; and the first original molecule set X init , the second candidate molecule set X 2,opt, the second polymer set P2 and the second evaluation index set GE2 are used as the corresponding first original molecule set field, the first candidate molecule set field, the first polymer set field and the first evaluation index set field to form a corresponding first evaluation record and add it to the local evaluation record list; and judge whether the second count value C2 is less than the simulation number threshold parameter; if the second count value C2 is greater than or equal to the simulation number threshold parameter, the evaluation record list is extracted as the corresponding first evaluation record list and sent to the self-assembly short peptide output module, and the local evaluation record list is cleared; if the second count value C2 is less than the simulation number threshold parameter, each second polymer p r and the corresponding second polymer index sequence E r As a set of corresponding first training polypeptide molecules and first label indicator sequences, the obtained first training polypeptide molecules and the first label indicator sequences constitute a corresponding first self-training data, and all the obtained first self-training data constitute a corresponding first batch data set added to the self-training data set of the self-training module; the second evaluation index set GE2 includes multiple second polymer indicator sequences E r The second polymer index sequence E r Including multiple second aggregate indicators e r,h , the index h is an integer greater than 0; the second aggregate index e r,h The index types include aggregation strength index, water solubility index and structural order index;

[0027] The self-assembling short peptide output module is used to poll each of the first evaluation records in the first evaluation record list in sequence when receiving the first evaluation record list; and when polling, the first evaluation record currently polled is used as the corresponding current evaluation record, and the first polymer set field and the first evaluation index set field of the current evaluation record are extracted as the corresponding current polymer set P and current evaluation index set GE; and each polymer index sequence E in the current evaluation index set GE is traversed; and when traversing, the currently traversed polymer index sequence E is used as the corresponding current polymer index sequence, and when all polymer indicators in the current polymer index sequence exceed the preset index threshold of their respective corresponding types, the current polymer is traversed. The combined index sequence is recorded as the corresponding first qualified index sequence, and the polymer p corresponding to the current polymer index sequence in the current polymer set P is recorded as the corresponding first qualified polymer, and when the molecular structure type of the first qualified polymer does not match the output structural morphological parameters, the first qualified polymer is converted into a one-dimensional, two-dimensional or three-dimensional structure corresponding to the structural morphological parameters and the conversion result is used as the new first qualified polymer, and it is identified whether the first qualified polymer is a short peptide. If so, the current first qualified polymer is used as a corresponding first short peptide; and when all the first evaluation records in the first evaluation record list are polled, the corresponding short peptide prediction report is composed of all the obtained first short peptides and output.

[0028] Further preferably, the first self-training data set is used to train each of the machine learning models M i The first prediction model performs a round of model training, specifically including:

[0029] Step 51: the self-training module extracts the first self-training data from the first self-training data set as the corresponding current training data; and uses the currently trained machine learning model M i As the corresponding current machine learning model;

[0030] Step 52: The self-training module extracts the first training polypeptide molecule and the first label indicator sequence from the current training data as the corresponding current training molecular structure and current label indicator sequence;

[0031] Step 53: The self-training module inputs the current training molecular structure into the first main control unit of the current machine learning model; and the first main control unit inputs the current training molecular structure into the feature encoding module for encoding to obtain the corresponding first training encoding tensor e trSend to the first prediction model of the current machine learning model; and the first prediction model encodes the first training tensor e tr Perform multi-class self-assembly index prediction to obtain the corresponding first training prediction index sequence E tr Send back to the first main control unit; and the first main control unit sends the first training prediction indicator sequence E tr Sending back to the self-training module;

[0032] Step 54: the self-training module converts the first training prediction indicator sequence E tr Substituting the current label indicator sequence into a preset first loss function to calculate a corresponding first loss value;

[0033] In step 55, the self-training module determines whether the first loss value satisfies a preset first loss value range. If the first loss value satisfies the first loss value range, the self-training module determines whether the current training data is the last first self-training data in the first self-training data set. If the current training data is the last first self-training data set, the process proceeds to step 56. If the current training data is not the last first self-training data set, the self-training module extracts the next first self-training data in the first self-training data set as the new current training data and returns to step 52 to continue training. If the first loss value does not satisfy the first loss value range, the model parameters of the first prediction model of the current machine learning model are modulated, and when the modulation is completed, the process returns to step 53 to continue training.

[0034] Step 56: The self-training module confirms the machine learning model M corresponding to the current machine learning model i This round of training is over.

[0035] A second aspect of the embodiments of the present invention provides a device for implementing the processing method of the self-assembling short peptide prediction model described in the first aspect, the device comprising: a model construction module, a data receiving module, a model configuration module and a model application module;

[0036] The model construction module is used to construct a self-assembling short peptide prediction model; the self-assembling short peptide prediction model includes a molecular dynamics simulation module, a self-assembly index evaluation module and a self-assembling short peptide output module; the molecular dynamics simulation module is preset with a configurable force field type parameter; the self-assembly index evaluation module is preset with a configurable simulation number threshold parameter; the self-assembly short peptide output module is preset with a configurable output structure morphology parameter;

[0037] The data receiving module is used to receive a first polypeptide molecule set, a first simulation number threshold, and a first output structure type input by a user; the first polypeptide molecule set includes a plurality of first polypeptide molecules; the molecular structure types of the first polypeptide molecules include one-dimensional structure, two-dimensional structure, and three-dimensional structure; the first output structure type includes one-dimensional structure, two-dimensional structure, and three-dimensional structure;

[0038] The model configuration module is used to set the force field type parameter of the molecular dynamics simulation module of the self-assembly short peptide prediction model to the PACE force field, and set the simulation number threshold parameter of the self-assembly index evaluation module to the corresponding first simulation number threshold, and set the output structure morphology parameter of the self-assembly short peptide output module to the corresponding first output structure type;

[0039] The model application module is used to predict the structure of a self-assembled short peptide chain according to the first polypeptide molecule set using the self-assembled short peptide prediction model to obtain a corresponding first short peptide prediction report.

[0040] A third aspect of an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;

[0041] The processor is configured to be coupled to the memory, read and execute instructions in the memory, so as to implement the method steps described in the first aspect above;

[0042] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

[0043] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a computer, the computer executes the instructions of the method described in the first aspect above.

[0044] The embodiment of the present invention provides a processing method, device, electronic device and computer-readable storage medium for a self-assembling short peptide prediction model; an end-to-end self-assembling short peptide prediction model is provided to simulate the self-assembly process of polypeptide molecules and output a short peptide prediction report consisting of multiple short peptide prediction records; the self-assembling short peptide prediction model is characterized in that: multiple similar or different types of machine learning models are used to predict the self-assembly ability of each polypeptide molecule in the polypeptide molecule set and a candidate molecule set is formed by one or more polypeptide molecules with high probability, and the candidate molecule set is simulated by a molecular dynamics simulation module whose force field condition is specifically set to a PACE force field. During the simulation process, each machine learning model is trained based on a synchronous learning strategy to reduce the difficulty of training data collection, reduce the difficulty of model training, and improve the accuracy of model prediction. The present invention can shorten the research cycle, reduce and quantify the research cost, and get rid of excessive reliance on manual experience, thereby increasing the probability of discovering new self-assembling short peptides. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A schematic diagram of a method for processing a self-assembling short peptide prediction model provided in Example 1 of the present invention;

[0046] Figure 2 This is a module structure diagram of the self-assembling short peptide prediction model provided in Example 1 of the present invention;

[0047] Figure 3 A module structure diagram of a processing device for a self-assembling short peptide prediction model provided in Example 2 of the present invention;

[0048] Figure 4 This is a structural diagram of an electronic device provided in Example 3 of the present invention. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the embodiments described herein are merely some, rather than all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.

[0050] The first embodiment of the present invention provides a method for processing a self-assembling short peptide prediction model, such as Figure 1 The schematic diagram of a processing method for a self-assembling short peptide prediction model provided in Example 1 of the present invention is shown. The method mainly includes the following steps:

[0051] Step 1: Construct a self-assembling short peptide prediction model.

[0052] Here, the self-assembly short peptide prediction model of the embodiment of the present invention is as follows Figure 2 The module structure diagram of the self-assembling short peptide prediction model provided in Example 1 of the present invention is shown as follows: it includes an initialization module, a feature encoding module, a self-training module, a first number N of machine learning models M i , self-assembly molecule screening module, molecular dynamics simulation module, self-assembly index evaluation module and self-assembly short peptide output module, the first number N is an integer not less than 2, 1≤model index i≤first number N.

[0053] The connection relationship between each module is as follows Figure 2 As shown: the initialization module is respectively connected to the input end of the self-assembly short peptide prediction model, the feature encoding module, and each machine learning model M i , self-assembly molecular screening module and molecular dynamics simulation module connection; each machine learning model M i They are respectively connected to the self-training module, feature encoding module and self-assembly molecule screening module; the self-assembly molecule screening module is connected to the molecular dynamics simulation module; the molecular dynamics simulation module is connected to the self-assembly index evaluation module; the self-assembly index evaluation module is respectively connected to the self-training module and the self-assembly short peptide output module; the self-assembly short peptide output module is connected to the output end of the self-assembly short peptide prediction model.

[0054] It should be noted that the model structure of the feature encoding module of the self-assembling short peptide prediction model of the embodiment of the present invention is composed of some or all of the following types of models: a molecular physicochemical property encoding model with molecular architecture as input, a sequence pre-training model with one-dimensional molecular string expression as input, a graph neural network pre-training model with two-dimensional molecular topology expression as input, and a Uni-Mol pre-training model with three-dimensional molecular system coordinates as input. Among them:

[0055] 1) The three-dimensional molecular architecture is a structural expression with information such as atomic properties (atom identity, element type, atomic coordinates, atomic charge, cluster identity, etc.) and bond properties (such as bond identity, bond type, etc.). The molecular physicochemical property encoding model of the embodiment of the present invention can be implemented with reference to, for example, a multilayer perception network model and a convolutional neural network model. The molecular physicochemical property encoding model of the embodiment of the present invention is used to encode the physical and electrochemical characteristics of the input three-dimensional molecular architecture to obtain a corresponding encoding tensor, and provide the encoding tensor to other modules using this type of encoding;

[0056] 2) Common one-dimensional molecular string expressions in the embodiments of the present invention include molecular fingerprint sequences, SMILES sequences, and other types. The sequence pre-training model in the embodiments of the present invention can be implemented with reference to the BERT model or a similar NLP language model. The sequence pre-training model corresponding to each sequence type in the embodiments of the present invention is used to encode the physical and electrochemical characteristics of the input molecular string expression to obtain a corresponding encoding tensor, and provide the encoding tensor to other modules using this type of encoding;

[0057] 3) The two-dimensional molecular topological expression of the embodiment of the present invention is a molecular topological graph. The graph neural network pre-training model of the embodiment of the present invention can be implemented with reference to model structures such as the GNN model and the EGNN model. The graph neural network pre-training model of the embodiment of the present invention is used to encode the physical and electrochemical characteristics according to the input molecular topological expression to obtain a corresponding encoding tensor, and provide the encoding tensor to other modules using this type of encoding;

[0058] 4) The three-dimensional molecular system coordinates of the embodiment of the present invention include multiple atomic coordinates, each atomic coordinate corresponds to an atom, each atom has a set of atomic properties, each set of atomic properties has a bonding information (if it is an isolated atom, this information is empty), and each bonding information includes a connecting bond identifier and a connecting bond type. The Uni-Mol pre-trained model of the embodiment of the present invention performs physical and electrochemical feature encoding based on the input molecular system coordinates to obtain a corresponding encoding tensor, and provides the encoding tensor to other modules that use this type of encoding;

[0059] Furthermore, the processing logic of the feature encoding module in the embodiment of the present invention is similar to the processing logic of the base learners module of the Uni-QSAR model. For details, please refer to the technical document "Uni-QSAR: an Auto-ML Tool for Molecular Property Prediction". It should be pointed out that the purpose of this design of the feature encoding module in the embodiment of the present invention is to maximize compatibility with multiple types of input structures.

[0060] It should also be noted that each machine learning model M in the self-assembly short peptide prediction model of the embodiment of the present invention i It includes a first main control unit and a first prediction model; the first main control unit is connected to the initialization module, the self-training module, the feature encoding module, the first prediction model and the self-assembly molecular screening module respectively; the model structure of the first prediction model includes a linear regression model structure, a kernel function model structure, a support vector machine model structure, a decision tree model structure, a random forest model structure, a gradient boosting model structure, an XGBoost model structure, a LightGBM model structure and a deep learning model structure; each two machine learning models M iThe first master control unit is the same, each of the two machine learning models M i The first prediction model is not forced to be the same; every two machine learning models M i The initial model parameters of the first prediction model and the model parameters obtained after training are different.

[0061] It should also be noted that several modules in the self-assembly short peptide prediction model of the embodiment of the present invention provide configurable parameters to the outside, among which: a configurable force field type parameter is preset on the molecular dynamics simulation module; a configurable simulation number threshold parameter is preset on the self-assembly index evaluation module; and a configurable output structure morphology parameter is preset on the self-assembly short peptide output module.

[0062] The self-assembling short peptide prediction model of the embodiment of the present invention is used to predict the input polypeptide molecule set {x j} to predict the structure of self-assembled short peptides and output the corresponding short peptide prediction report; wherein, the polypeptide molecule set {x j} includes multiple polypeptide molecules x j , the molecular index j is an integer greater than 0; the polypeptide molecule x j The structural types include one-dimensional structure, two-dimensional structure and three-dimensional structure; the short peptide prediction report includes multiple first short peptides; the structural type of the first short peptide corresponds to the output structural morphological parameters, including one-dimensional structure, two-dimensional structure and three-dimensional structure.

[0063] The following describes the execution steps of each module in the self-assembling short peptide prediction model according to an embodiment of the present invention when performing a self-assembling short peptide structure prediction process.

[0064] The initialization module is used to receive the input polypeptide molecule set {x j}; and each polypeptide molecule x in the collection j Input feature encoding module to encode and obtain the corresponding first encoding tensor e j ; and all the first encoded tensors e are obtained j The first coding sequence corresponding to {e j}; and the polypeptide molecule collection {x j} as the corresponding first original molecule set X init and the first candidate molecule set X 1,opt ; and the first original molecule set X init , the first candidate molecule set X 1,opt and the first coding sequence {e j} Input to the molecular dynamics simulation module; and the first original molecule set X init Each machine learns a model M i The first master control unit sends the first original molecule set X init and the first coding sequence {ej}Send to the self-assembly molecular screening module.

[0065] The feature encoding module is used to perform feature encoding on one-dimensional, two-dimensional or three-dimensional molecules input by any internal module and output the corresponding encoding tensor.

[0066] The molecular dynamics simulation module is used to receive the first original molecular set X sent by the initialization module init , the first candidate molecule set X 1,opt and the first coding sequence {e j}hour:

[0067] Step A1: For the first original molecule set X init Save;

[0068] Step A2, and initialize the count value of the preset simulation times counter to 0;

[0069] Step A3, and based on the preset molecular dynamics simulation input file format, the first coding sequence {e j}Convert the simulation input file to obtain the corresponding current simulation input file;

[0070] Step A4, calling the preset molecular dynamics simulation interface to perform molecular dynamics simulation according to the force field type parameters and the current simulation input file, and continuously timing this simulation;

[0071] Here, the molecular dynamics simulation interface of the embodiment of the present invention is a processing interface of a molecular dynamics simulation software. The aforementioned molecular dynamics simulation input file format is the input file format corresponding to the interface. As can be seen from the subsequent steps, the embodiment of the present invention sets the force field as the PACE force field by default when performing molecular dynamics simulation. The force field potential function and the corresponding motion system equation of the PACE force field can be referred to the technical document "PACE Force Field for Protein Simulations.1. Full Parameterization of Version 1 and Verification" and other relevant technical documents, which are not further described here. It should be noted that the PACE force field is a multi-scale coarse-grained force field, which has a faster calculation speed than the conventional all-atom force field and has better calculation precision and accuracy than the conventional MARTINI coarse-grained force field.

[0072] Step A5: When the simulation time reaches a preset first short duration threshold, it is determined whether a clustering effect occurs during the current simulation. If no clustering effect occurs, the current simulation is terminated early. If a clustering effect occurs, the simulation is continued and terminated normally when the simulation time reaches a preset first long duration threshold.

[0073] Here, the first short time threshold and the first long time threshold are two pre-set time length parameters, and the first short time threshold is much smaller than the first long time threshold. The self-assembling short peptide prediction model of the embodiment of the present invention can be implemented in a variety of ways to determine whether an aggregation effect occurs. One processing method is: when the simulation timing reaches the first short time threshold, a three-dimensional structure snapshot is taken of the three-dimensional simulation space corresponding to the current simulation process, and polymer recognition is performed on the three-dimensional structure snapshot based on the image processing model, and the recognition result is compared with the molecular structure of each polypeptide at the initial moment of the simulation. If there is a significant difference, it is confirmed that an aggregation effect has occurred. Another processing method is: when the simulation timing reaches the first short time threshold, an aggregation effect analysis is performed based on the continuous trajectory file of the simulated process. If it is found through the trajectory file that the atomic order of each polypeptide molecule at the initial moment of the simulation has changed significantly or interacts with neighboring particles, it is confirmed that an aggregation effect has occurred.

[0074] It should also be noted that, when the self-assembly short peptide prediction model of the embodiment of the present invention confirms that no aggregation effect occurs and terminates the simulation early when the simulation timing reaches the first short time threshold, a preset early termination prompt message can be given, or the self-assembly index evaluation module can be directly transferred to the self-assembly index evaluation module to extract the local evaluation record list as the corresponding first evaluation record list and send it to the self-assembly short peptide output module and clear the local evaluation record list at the same time;

[0075] Step A6: When the simulation is terminated normally, the count value of the simulation counter is used as the corresponding first count value C1; the simulation trajectory file output by the simulation is used as the corresponding first trajectory file; the cluster structure of the first trajectory file is identified and each identified cluster structure is used as the corresponding first aggregate p k , the aggregate index k is an integer greater than 0; and all the first aggregates p obtained this time k Form the corresponding first polymer set P1; and combine the first count value C1, the first trajectory file, the first original molecule set X init , the first candidate molecule set X 1,opt and the first polymer set P1 are sent to the self-assembly index evaluation module.

[0076] Here, when performing cluster structure recognition on the first trajectory file, the self-assembly short peptide prediction model of the embodiment of the present invention will exclude cluster structures that are highly similar to the molecular structures of each polypeptide molecule at the initial moment of the simulation, and will also exclude cluster structures with unstable structural properties.

[0077] An evaluation record list is preset on the self-assembly index evaluation module; the evaluation record list includes multiple first evaluation records; the first evaluation record includes a first original molecule set field, a first candidate molecule set field, a first polymer set field and a first evaluation index set field;

[0078] Here, the molecular dynamics simulation module of the self-assembly short peptide prediction model of the embodiment of the present invention will perform multiple molecular dynamics simulations. After each simulation is completed, the self-assembly index evaluation module will perform a self-assembly index evaluation on the polymer obtained in the current simulation, and generate a corresponding first evaluation record based on the relevant information of the current evaluation and store it in the evaluation record list.

[0079] The self-assembly index evaluation module is used to receive the first count value C1, the first trajectory file, the first original molecule set X init , the first candidate molecule set X 1,opt and the first aggregate set P1:

[0080] Step B1: Based on the first trajectory file, each first aggregate p in the first aggregate set P1 is k Perform multiple self-assembly index calculations to obtain the corresponding first polymer index sequence E k ;

[0081] Among them, the first aggregate indicator sequence E k Including multiple first aggregate indicators e k,h , the index h is an integer greater than 0; the first aggregate index e k,h The index types include aggregation strength (AP), water solubility (logP), and permutation ordering (PO). Here, when the simulation trajectory file is known, the calculation of the corresponding aggregation strength, water solubility, and permutation ordering for each specified polymer object in the molecular system based on the simulation trajectory file can be achieved based on conventional calculation methods provided in the fields of materials science, biochemistry, statistics, etc., which will not be described in detail here.

[0082] Step B2, and all the first aggregate index sequences E obtained k Form the corresponding first evaluation index set GE1;

[0083] The first evaluation index set GE1 includes a plurality of first aggregate index sequences E k ;

[0084] Step B3, and the first original molecule set X init , the first candidate molecule set X 1,opt , the first polymer set P1 and the first evaluation index set GE1 are used as the corresponding first original molecule set field, the first candidate molecule set field, the first polymer set field and the first evaluation index set field to form a corresponding first evaluation record and add it to the local evaluation record list;

[0085] Step B4, determining whether the first count value C1 is less than a simulation times threshold parameter;

[0086] Step B5: if the first count value C1 is greater than or equal to the simulation number threshold parameter, the evaluation record list is extracted and sent as the corresponding first evaluation record list to the self-assembly short peptide output module, and the local evaluation record list is cleared;

[0087] Step B6: If the first count value C1 is less than the simulation times threshold parameter, each first aggregate p k and the corresponding first aggregate index sequence E k As a set of corresponding first training polypeptide molecules and first label indicator sequences, the obtained first training polypeptide molecules and first label indicator sequences constitute a corresponding first self-training data set, and all the obtained first self-training data constitute a corresponding first batch data set added to the self-training data set of the self-training module.

[0088] Here, the self-assembly short peptide prediction model of the embodiment of the present invention trains each machine learning model based on a synchronous learning strategy during the simulation process. The training data set required for training is composed of the self-assembly index evaluation module based on the autonomous collection method as shown above.

[0089] A self-training data set is preset on the self-training module; the self-training data set includes a plurality of first self-training data; the first self-training data includes a first training polypeptide molecule and a first label indicator sequence.

[0090] The self-training module is used to extract all the first self-training data of the current self-training data set to form a corresponding first self-training data set each time a first batch of data sets is added to the self-training data set; and to train each machine learning model M according to the first self-training data set. i The first prediction model performs a round of model training; and all machine learning models M i At the end of this round of training, each machine learning model M i The first main control unit sends a corresponding single-round training end mark.

[0091] Here, each machine learning model M is trained based on the first self-training data set. i The first prediction model performs a round of model training, which includes:

[0092] Step C1, the self-training module extracts the first self-training data of the first self-training data set as the corresponding current training data; and uses the currently trained machine learning model M i As the corresponding current machine learning model;

[0093] Step C2, the self-training module extracts the first training polypeptide molecule and the first label index sequence from the current training data as the corresponding current training molecular structure and current label index sequence;

[0094] Step C3: The self-training module inputs the current training molecular structure into the first main control unit of the current machine learning model; and the first main control unit inputs the current training molecular structure into the feature encoding module for encoding to obtain the corresponding first training encoding tensor e tr Send to the first prediction model of the current machine learning model; and the first prediction model encodes the first training tensor e tr Perform multi-class self-assembly index prediction to obtain the corresponding first training prediction index sequence E tr Send back to the first main control unit; and the first main control unit sends the first training prediction indicator sequence E tr Send back to the self-training module;

[0095] Step C4, the self-training module converts the first training prediction indicator sequence E tr Substitute the current label indicator sequence into the preset first loss function to calculate the corresponding first loss value;

[0096] Here, the first loss function is a preset loss function, which corresponds to the model structure and model parameters of the first prediction model of the current machine learning model;

[0097] In step C5, the self-training module determines whether the first loss value satisfies a preset first loss value range. If the first loss value satisfies the first loss value range, the self-training module determines whether the current training data is the last first self-training data in the first self-training data set. If it is the last, the process proceeds to step C6. If it is not the last, the self-training module extracts the next first self-training data in the first self-training data set as the new current training data and returns to step C2 to continue training. If the first loss value does not satisfy the first loss value range, the model parameters of the first prediction model of the current machine learning model are modulated, and when the modulation is completed, the process returns to step C3 to continue training.

[0098] Here, the first loss value range is a preset loss value range, which corresponds to the loss function of the current model;

[0099] Step C6: The self-training module confirms the machine learning model M corresponding to the current machine learning model i This round of training is over.

[0100] Machine learning model M i The first main control unit is used to receive the first original molecule set X sent by the initialization module init The machine learning model M i The first main control unit is further configured to, upon receiving a single round training end flag sent by the self-training module, convert the first original molecule set X init Each polypeptide molecule x j Input feature encoding module to encode and obtain the corresponding second encoding tensor e i,j ; and all the second encoding tensors e obtained are i,j The corresponding second coding sequence {e i,j}; and the second coding sequence {e i,j}Send to the first prediction model and receive the first prediction indicator set GE sent back by the first prediction model p,i ; and by the second coding sequence {e i,j} and the first set of predictors GE p,i Form a corresponding first learning data D i Send to the self-assembly molecular screening module.

[0101] Machine learning model M i The first prediction model is used to receive the second coded sequence {e i,j}, according to each second encoding tensor e in the sequence i,j Perform multi-class self-assembly index prediction to obtain the corresponding first prediction index sequence E i,j , and all the first prediction indicator sequences E are obtained i,j The first prediction indicator set GE p,i Send back to the first main control unit;

[0102] Among them, the first prediction indicator set GE p,i Including multiple first prediction indicator sequences E i,j ; The first prediction indicator sequence E i,j Including multiple first prediction indicators e i,j,h ; The first prediction indicator e i,j,h The types of indicators include aggregation strength index, water solubility index and structural order index.

[0103] The self-assembly molecular screening module is used to receive the first original molecular set X sent by the initialization module init and the first coding sequence {e j} to save both.

[0104] The self-assembly molecular screening module is also used to receive all the learning models M i The first learning data D sent i after:

[0105] Step D1, all the first learning data D obtained this time i Form a corresponding first learning data set;

[0106] Step D2, and combine the first learning data with the first original molecule set X init Each polypeptide molecule x j All corresponding first prediction indicators e i,j,h Extracted to form a corresponding first molecule observation vector b j , and all the first molecule observation vectors b are obtained j Composed of the corresponding first observation tensor B;

[0107] Step D3, the first observation tensor B is substituted into the preset acquisition function equation to perform high probability evaluation to obtain a plurality of first evaluation data v j The first evaluation vector V composed of

[0108] The first evaluation vector V includes a plurality of first evaluation data v j , the first evaluation data v j With the first molecular observation vector b j One-to-one correspondence;

[0109] Here, the acquisition function is a function used for high availability evaluation in a conventional Bayesian optimization model. The sub-classification methods of the function include the expected improvement (EI) method, the probability improvement (PI) method, and the upper confidence bound (UCB) method. The embodiment of the present invention can construct a corresponding acquisition function equation based on any of the above methods, and the acquisition function equation is used to evaluate the quality of each polypeptide molecule x. j The first molecular observation vector b j The high availability of each molecule (here, the high probability of being able to self-assemble) is evaluated, and the first evaluation data v j Each polypeptide molecule x j the assessment results;

[0110] Step D4, and the first evaluation data v j Peptide molecules that are not lower than the preset qualification assessment threshold x j Denote as the corresponding first candidate molecule x opt,w , the molecule index w is an integer greater than 0;

[0111] Here, the qualified assessment threshold is a pre-set assessment value;

[0112] Step D5, and all the first candidate molecules x obtained opt,w The corresponding second candidate molecule set X 2,opt ;

[0113] Step D6, and the first coding sequence {e j} and each first candidate molecule x opt,w The corresponding first encoding tensor e j As the corresponding third encoding tensor e w , and all the third encoded tensors e are obtained by w The corresponding third coding sequence {e w};

[0114] Step D7, and set the second candidate molecule set X 2,opt and the third coding sequence {e w}Sent to the molecular dynamics simulation module.

[0115] The molecular dynamics simulation module is also used to receive the second candidate molecule set X sent by the self-assembly molecule screening module. 2,opt and the third coding sequence {e w}hour:

[0116] Step E1, adding 1 to the count value of the simulation times counter;

[0117] Step E2, and based on the molecular dynamics simulation input file format, the third coding sequence {e w}Convert the simulation input file to obtain the corresponding current simulation input file;

[0118] Step E3, calling the molecular dynamics simulation interface to perform molecular dynamics simulation according to the force field type parameters and the current simulation input file, and continuously timing this simulation;

[0119] Step E4: When the simulation timer reaches the first short duration threshold, it is determined whether a clustering effect occurs during the current simulation. If no clustering effect occurs, the simulation is terminated early. If a clustering effect occurs, the simulation is continued and terminated normally when the simulation timer reaches the first long duration threshold.

[0120] Step E5: When the simulation is terminated normally, the count value of the simulation counter is used as the corresponding second count value C2; ​​the simulation trajectory file output by the simulation is used as the corresponding second trajectory file; the cluster structure of the second trajectory file is identified and each cluster structure identified is used as the corresponding second aggregate p r , the aggregate index r is an integer greater than 0; and all the second aggregates p obtained this time r The corresponding second polymer set P2 is formed; the second count value C2, the second trajectory file, the first original molecule set X init , the second candidate molecule set X 2,opt And the second polymer set P2 is sent to the self-assembly index evaluation module.

[0121] The self-assembly index evaluation module is further configured to receive the second count value C2, the second trajectory file, the first original molecule set X init , the second candidate molecule set X 2,opt And the second aggregate set P2:

[0122] Step F1: Based on the second trajectory file, each second aggregate p in the second aggregate set P2 is r Calculate multiple types of self-assembly indices to obtain the corresponding second polymer index sequence E r ;

[0123] Among them, the second aggregate indicator sequence E r Including multiple second aggregate indicators e r,h , the index h is an integer greater than 0; the second aggregate index e r,h The index types include aggregation strength index, water solubility index and structural order index;

[0124] Step F2, and all the second aggregate index sequences E are obtained r Composing the corresponding second evaluation index set GE2;

[0125] The second evaluation index set GE2 includes multiple second aggregate index sequences E r ;

[0126] Step F3, and the first original molecule set X init , the second candidate molecule set X 2,opt , the second polymer set P2 and the second evaluation index set GE2 are used as the corresponding first original molecule set field, the first candidate molecule set field, the first polymer set field and the first evaluation index set field to form a corresponding first evaluation record and add it to the local evaluation record list;

[0127] Step F4, and determining whether the second count value C2 is less than the simulation times threshold parameter;

[0128] Step F5: If the second count value C2 is greater than or equal to the simulation number threshold parameter, the evaluation record list is extracted and sent to the self-assembly short peptide output module as the corresponding first evaluation record list, and the local evaluation record list is cleared;

[0129] Step F6: If the second count value C2 is less than the simulation times threshold parameter, each second aggregate p r and the corresponding second aggregate index sequence E r As a set of corresponding first training polypeptide molecules and first label indicator sequences, the obtained first training polypeptide molecules and first label indicator sequences constitute a corresponding first self-training data set, and all the obtained first self-training data constitute a corresponding first batch data set added to the self-training data set of the self-training module.

[0130] The self-assembling short peptide output module is used to, upon receiving the first evaluation record list:

[0131] Polling each first evaluation record in the first evaluation record list in sequence; and during the polling, taking the first evaluation record currently polled as the corresponding current evaluation record, and extracting the first aggregate set field and the first evaluation index set field of the current evaluation record as the corresponding current aggregate set P and current evaluation index set GE;

[0132] and traverse each polymer index sequence E in the current evaluation index set GE; and during the traversal, use the currently traversed polymer index sequence E as the corresponding current polymer index sequence, and when all polymer indicators in the current polymer index sequence exceed the preset index threshold of their respective corresponding types, record the current polymer index sequence as the corresponding first qualified index sequence, and record the polymer p corresponding to the current polymer index sequence in the current polymer set P as the corresponding first qualified polymer, and when the molecular structure type of the first qualified polymer does not match the output structural morphological parameters, perform a one-dimensional, two-dimensional or three-dimensional structure conversion corresponding to the structural morphological parameters on the first qualified polymer and use the conversion result as the new first qualified polymer, and identify whether the first qualified polymer is a short peptide, and if so, use the current first qualified polymer as a corresponding first short peptide;

[0133] When all first evaluation records in the first evaluation record list are polled, a corresponding short peptide prediction report is composed of all the obtained first short peptides and output.

[0134] Step 2, receiving a first polypeptide molecule set, a first simulation number threshold, and a first output structure type input by a user;

[0135] The first polypeptide molecule set includes multiple first polypeptide molecules; the molecular structure types of the first polypeptide molecules include one-dimensional structure, two-dimensional structure and three-dimensional structure; the first output structure types include one-dimensional structure, two-dimensional structure and three-dimensional structure.

[0136] Step 3: Set the force field type parameter of the molecular dynamics simulation module of the self-assembly short peptide prediction model to the PACE force field, set the simulation number threshold parameter of the self-assembly index evaluation module to the corresponding first simulation number threshold, and set the output structure morphology parameter of the self-assembly short peptide output module to the corresponding first output structure type.

[0137] Step 4: The self-assembling short peptide prediction model predicts the structure of the self-assembling short peptide chain based on the first polypeptide molecule set to obtain a corresponding first short peptide prediction report.

[0138] Here, from the previous description of the execution steps of each module within the self-assembling short peptide prediction model when performing self-assembling short peptide structure prediction processing, it can be seen that the self-assembling short peptide prediction model of an embodiment of the present invention will output a short peptide prediction report, i.e., the first short peptide prediction report, after inputting a polypeptide molecule set, i.e., the first polypeptide molecule set.

[0139] Figure 3 This is a module structure diagram of a processing device for a self-assembling short peptide prediction model provided in the second embodiment of the present invention. The device is a terminal device or server that implements the aforementioned method embodiment, or can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiment. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 3 As shown, the device includes: a model building module 201, a data receiving module 202, a model configuration module 203 and a model application module 204.

[0140] The model construction module 201 is used to construct a self-assembling short peptide prediction model; the self-assembling short peptide prediction model includes a molecular dynamics simulation module, a self-assembly index evaluation module and a self-assembly short peptide output module; a configurable force field type parameter is preset on the molecular dynamics simulation module; a configurable simulation number threshold parameter is preset on the self-assembly index evaluation module; and a configurable output structure morphology parameter is preset on the self-assembly short peptide output module.

[0141] The data receiving module 202 is used to receive a first polypeptide molecule set, a first simulation number threshold, and a first output structure type input by a user; the first polypeptide molecule set includes multiple first polypeptide molecules; the molecular structure types of the first polypeptide molecules include one-dimensional structure, two-dimensional structure, and three-dimensional structure; the first output structure type includes one-dimensional structure, two-dimensional structure, and three-dimensional structure.

[0142] The model configuration module 203 is used to set the force field type parameter of the molecular dynamics simulation module of the self-assembly short peptide prediction model to the PACE force field, set the simulation number threshold parameter of the self-assembly index evaluation module to the corresponding first simulation number threshold, and set the output structure morphology parameter of the self-assembly short peptide output module to the corresponding first output structure type.

[0143] The model application module 204 is used to predict the structure of a self-assembled short peptide chain based on the first polypeptide molecule set using the self-assembled short peptide prediction model to obtain a corresponding first short peptide prediction report.

[0144] An embodiment of the present invention provides a processing device for a self-assembling short peptide prediction model, which can execute the method steps in the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.

[0145] It should be noted that the division of the modules of the above devices is merely a division of logical functions. In actual implementation, they can be fully or partially integrated into a physical entity or physically separated. Furthermore, these modules can all be implemented in the form of software called by a processing element; or all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the model building module can be a separate processing element, or it can be integrated into a chip of the above device. In addition, it can be stored in the form of program code in the memory of the above device, and called by a processing element of the above device to perform the functions of the above-mentioned module. The implementation of other modules is similar. In addition, these modules can all or partly be integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed by hardware integrated logic circuits in the processor element or instructions in the form of software.

[0146] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, when a module is implemented by scheduling program code through a processing element, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0147] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the above method embodiments are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The above-mentioned computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above-mentioned computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.) means. The above-mentioned computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The above-mentioned available medium can be a magnetic medium (such as a floppy disk, hard disk, tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0148] Figure 4 This is a schematic diagram of the structure of an electronic device provided in the third embodiment of the present invention. The electronic device can be a terminal device or server that implements the method of the aforementioned embodiment, or it can be a terminal device or server that implements the method of the aforementioned embodiment connected to the aforementioned terminal device or server. Figure 4As shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver 303's transceiver actions. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the aforementioned embodiment method. Preferably, the electronic device involved in the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The above-mentioned communication port 306 is used for connecting and communicating between the electronic device and other peripherals.

[0149] exist Figure 4 The system bus 305 mentioned in the figure can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to realize communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include random access memory (RAM) and may also include non-volatile memory (Non-Volatile Memory), such as at least one disk storage.

[0150] The above-mentioned processors can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0151] It should be noted that an embodiment of the present invention further provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the methods and processing procedures provided in the above embodiments.

[0152] An embodiment of the present invention further provides a chip for executing instructions, which is used to execute the processing steps described in the above method embodiment.

[0153] The embodiment of the present invention provides a processing method, device, electronic device and computer-readable storage medium for a self-assembling short peptide prediction model; an end-to-end self-assembling short peptide prediction model is provided to simulate the self-assembly process of polypeptide molecules and output a short peptide prediction report consisting of multiple short peptide prediction records; the self-assembling short peptide prediction model is characterized in that: multiple similar or different types of machine learning models are used to predict the self-assembly ability of each polypeptide molecule in the polypeptide molecule set and a candidate molecule set is formed by one or more polypeptide molecules with high probability, and the candidate molecule set is simulated by a molecular dynamics simulation module whose force field condition is specifically set to a PACE force field. During the simulation process, each machine learning model is trained based on a synchronous learning strategy to reduce the difficulty of training data collection, reduce the difficulty of model training, and improve the accuracy of model prediction. The present invention can shorten the research cycle, reduce and quantify the research cost, and get rid of excessive reliance on manual experience, thereby increasing the probability of discovering new self-assembling short peptides.

[0154] Professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0155] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0156] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for processing a self-assembling short peptide prediction model, characterized in that: The method comprises: Constructing a self-assembling short peptide prediction model; the self-assembling short peptide prediction model includes a molecular dynamics simulation module, a self-assembly index evaluation module and a self-assembling short peptide output module; the molecular dynamics simulation module is preset with a configurable force field type parameter; the self-assembly index evaluation module is preset with a configurable simulation number threshold parameter; the self-assembly short peptide output module is preset with a configurable output structure morphology parameter; Receiving a first polypeptide molecule set, a first simulation number threshold, and a first output structure type input by a user; the first polypeptide molecule set includes a plurality of first polypeptide molecules; the molecular structure types of the first polypeptide molecules include one-dimensional structure, two-dimensional structure, and three-dimensional structure; the first output structure type includes one-dimensional structure, two-dimensional structure, and three-dimensional structure; The force field type parameter of the molecular dynamics simulation module of the self-assembly short peptide prediction model is set to the PACE force field, the simulation number threshold parameter of the self-assembly index evaluation module is set to the corresponding first simulation number threshold, and the output structure morphology parameter of the self-assembly short peptide output module is set to the corresponding first output structure type; The self-assembling short peptide prediction model predicts the structure of the self-assembling short peptide chain based on the first polypeptide molecule set to obtain a corresponding first short peptide prediction report.

2. The method for processing the self-assembly short peptide prediction model according to claim 1, wherein The self-assembling short peptide prediction model includes an initialization module, a feature encoding module, a self-training module, a first number N of machine learning models M i , a self-assembly molecule screening module, the molecular dynamics simulation module, the self-assembly index evaluation module and the self-assembly short peptide output module, the first number N is an integer not less than 2, 1≤model index i≤first number N; The initialization module is respectively connected to the input end of the self-assembly short peptide prediction model, the feature encoding module, and each of the machine learning models M i , the self-assembly molecular screening module and the molecular dynamics simulation module are connected; each of the machine learning models M i are respectively connected to the self-training module, the feature encoding module and the self-assembly molecular screening module; the self-assembly molecular screening module is connected to the molecular dynamics simulation module; the molecular dynamics simulation module is connected to the self-assembly index evaluation module; the self-assembly index evaluation module is respectively connected to the self-training module and the self-assembly short peptide output module; the self-assembly short peptide output module is connected to the output end of the self-assembly short peptide prediction model; The machine learning model M i It includes a first main control unit and a first prediction model; the first main control unit is respectively connected to the initialization module, the self-training module, the feature encoding module, the first prediction model and the self-assembly molecular screening module; the model structure of the first prediction model includes a linear regression model structure, a kernel function model structure, a support vector machine model structure, a decision tree model structure, a random forest model structure, a gradient boosting model structure, an XGBoost model structure, a LightGBM model structure and a deep learning model structure; every two of the machine learning models M i The first main control unit is the same as each of the two machine learning models M i The first prediction models are not forced to be the same; each of the two machine learning models M i The initial model parameters of the first prediction model and the model parameters obtained after training are different; The self-training module is pre-set with a self-training data set; the self-training data set includes a plurality of first self-training data; the first self-training data includes a first training polypeptide molecule and a first label indicator sequence; An evaluation record list is preset on the self-assembly index evaluation module; the evaluation record list includes multiple first evaluation records; the first evaluation record includes a first original molecule set field, a first candidate molecule set field, a first polymer set field and a first evaluation index set field.

3. The method for processing the self-assembly short peptide prediction model according to claim 2, wherein: The self-assembly short peptide prediction model is used to predict the input polypeptide molecule set {x j } performs self-assembly short peptide structure prediction and outputs the corresponding short peptide prediction report; the polypeptide molecule set {x j } includes multiple polypeptide molecules x j , the molecular index j is an integer greater than 0; the polypeptide molecule x j The structural types include one-dimensional structure, two-dimensional structure and three-dimensional structure; the short peptide prediction report includes multiple first short peptides; the structural types of the first short peptides correspond to the output structural morphological parameters, including one-dimensional structure, two-dimensional structure and three-dimensional structure.

4. The method for processing the self-assembly short peptide prediction model according to claim 3, wherein: The initialization module is used to receive the input polypeptide molecule set {x j }; and each of the polypeptide molecules x in the collection j Input the feature encoding module to encode and obtain the corresponding first encoding tensor e j ; and all the first encoding tensors e obtained by j The first coding sequence corresponding to {e j }; and the polypeptide molecule set {x j } as the corresponding first original molecule set X init and the first candidate molecule set X 1,opt ; and the first original molecule set X init , the first candidate molecule set X 1,opt and the first coding sequence {e j } input to the molecular dynamics simulation module; and the first original molecule set X init To each of the said machine learning models M i The first main control unit sends the first original molecule set X init and the first coding sequence {e j }Sending to the self-assembly molecular screening module; The feature encoding module is used to perform feature encoding on the input one-dimensional, two-dimensional or three-dimensional molecules and output the corresponding encoding tensor; The molecular dynamics simulation module is used to receive the first original molecule set X sent by the initialization module init , the first candidate molecule set X 1,opt and the first coding sequence {e j }, for the first original molecule set X init and initialize the count value of the preset simulation counter to 0; and based on the preset molecular dynamics simulation input file format, the first coding sequence {e j }Convert the simulation input file to obtain the corresponding current simulation input file; call the preset molecular dynamics simulation interface to perform molecular dynamics simulation according to the force field type parameters and the current simulation input file, and continuously time this simulation; When the simulation timing reaches a preset first short-duration threshold, it is determined whether a clustering effect occurs during the current simulation process. If no clustering effect occurs, the simulation is terminated early. If a clustering effect occurs, the simulation is continued and the simulation is terminated normally when the simulation timing reaches a preset first long-duration threshold. When the simulation is terminated normally, the count value of the simulation counter is used as the corresponding first count value C1. The simulation trajectory file output by the simulation is used as the corresponding first trajectory file. Cluster structure identification is performed on the first trajectory file, and each identified cluster structure is used as the corresponding first aggregate p. k , the aggregate index k is an integer greater than 0; and all the first aggregates p obtained this time k The first count value C1, the first trajectory file, the first original molecule set X init , the first candidate molecule set X 1,opt and the first polymer set P1 to the self-assembly index evaluation module; The self-assembly index evaluation module is used to receive the first count value C1, the first trajectory file, the first original molecule set X init , the first candidate molecule set X 1,opt and the first aggregate set P1, each of the first aggregates p in the first aggregate set P1 is tracked based on the first trajectory file. k Perform multiple self-assembly index calculations to obtain the corresponding first polymer index sequence E k ; and all the first polymer index sequences E obtained k The first evaluation index set GE1 is formed; and the first original molecule set X init , the first candidate molecule set X 1,opt , the first polymer set P1 and the first evaluation index set GE1 are used as the corresponding first original molecule set field, the first candidate molecule set field, the first polymer set field and the first evaluation index set field to form a corresponding first evaluation record and add it to the local evaluation record list; and judge whether the first count value C1 is less than the simulation number threshold parameter; if the first count value C1 is greater than or equal to the simulation number threshold parameter, the evaluation record list is extracted as the corresponding first evaluation record list and sent to the self-assembly short peptide output module, and the local evaluation record list is cleared; if the first count value C1 is less than the simulation number threshold parameter, each of the first polymers p k and the corresponding first polymer index sequence E k As a set of corresponding first training polypeptide molecules and first label indicator sequences, the obtained first training polypeptide molecules and the first label indicator sequences form a corresponding first self-training data set, and all the obtained first self-training data forms a corresponding first batch data set added to the self-training data set of the self-training module; the first evaluation index set GE1 includes multiple first polymer indicator sequences E k ; The first polymer index sequence E k Including multiple first aggregate indicators e k,h , the index h is an integer greater than 0; the first aggregate index e k,h The index types include aggregation strength index, water solubility index and structural order index; The self-training module is configured to extract all the first self-training data from the current self-training data set to form a corresponding first self-training data set each time a first batch of data sets is added to the self-training data set; And according to the first self-training data set, each of the machine learning models M i The first prediction model performs a round of model training; and all the machine learning models M i At the end of this round of training, each of the machine learning models M i The first main control unit sends a corresponding single-round training end flag; The machine learning model M i The first main control unit is used to receive the first original molecule set X sent by the initialization module init Save it when The machine learning model M i The first main control unit is further configured to, upon receiving the single-round training end flag sent by the self-training module, convert the first original molecule set X init Each of the polypeptide molecules x j Input the feature encoding module to encode and obtain the corresponding second encoding tensor e i,j ; and all the second encoding tensors e obtained by i,j The corresponding second coding sequence {e i,j }; and the second coding sequence {e i,j }Send to the first prediction model, and receive the first prediction indicator set GE sent back by the first prediction model p,i ; and the second coding sequence {e i,j } and the first set of predictive indicators GE p,i Form a corresponding first learning data D i Sending to the self-assembly molecular screening module; The machine learning model M i The first prediction model is used to receive the second coding sequence {e i,j }, according to each of the second encoding tensors e in the sequence i,j Perform multi-class self-assembly index prediction to obtain the corresponding first prediction index sequence E i,j , and all the first prediction indicator sequences E are obtained by i,j The first prediction indicator set GE p,i Send back to the first main control unit; the first prediction indicator set GE p,i Including a plurality of the first prediction indicator sequences E i,j ; The first prediction indicator sequence E i,j Including multiple first prediction indicators e i,j,h The first prediction indicator e i,j,h The index types include aggregation strength index, water solubility index and structural order index; The self-assembly molecular screening module is used to receive the first original molecular set X sent by the initialization module init and the first coding sequence {e j } when saving both; The self-assembly molecular screening module is also used to receive all the learning models M i The first learning data D sent i Afterwards, all the first learning data D obtained this time i Combining a corresponding first learning data set; and integrating the first learning data set with the first original molecule set X init Each of the polypeptide molecules x j Corresponding to all the first prediction indicators e i,j,h Extracted to form a corresponding first molecule observation vector b j , and all the first molecular observation vectors b are obtained by j The first observation tensor B is formed; and the first observation tensor B is substituted into the preset acquisition function equation for high probability evaluation to obtain a plurality of first evaluation data v j The first evaluation vector V is composed of the first evaluation data v j The polypeptide molecule x is not less than the preset qualification assessment threshold j Denote as the corresponding first candidate molecule x opt,w , the molecular index w is an integer greater than 0; and all the first candidate molecules x are obtained opt,w The corresponding second candidate molecule set X 2,opt ; and the first coding sequence {e j } and each of the first candidate molecules x opt,w The corresponding first encoding tensor e j As the corresponding third encoding tensor e w , and all the third encoded tensors e are obtained by w The corresponding third coding sequence {e w }; and the second candidate molecule set X 2,opt and the third coding sequence {e w } is sent to the molecular dynamics simulation module; the first evaluation vector V includes a plurality of the first evaluation data v j , the first evaluation data v j With the first molecular observation vector b j One-to-one correspondence; The molecular dynamics simulation module is further configured to receive the second candidate molecule set X sent by the self-assembly molecular screening module. 2,opt and the third coding sequence {e w }, the count value of the simulation times counter is increased by 1; and the third coding sequence {e w }Convert the simulation input file to obtain the corresponding current simulation input file; call the molecular dynamics simulation interface to perform molecular dynamics simulation according to the force field type parameters and the current simulation input file, and continuously time the simulation; When the simulation timing reaches the first short-duration threshold, whether a clustering effect occurs in the current simulation process is identified; if no clustering effect occurs, the simulation is terminated early; if a clustering effect occurs, the simulation is continued and the simulation is terminated normally when the simulation timing reaches the first long-duration threshold; when the simulation is terminated normally, the count value of the simulation times counter is used as the corresponding second count value C2; ​​and the simulation trajectory file output by the simulation is used as the corresponding second trajectory file; and the cluster structure is identified on the second trajectory file and each identified cluster structure is used as the corresponding second aggregate p r , the aggregate index r is an integer greater than 0; and all the second aggregates p obtained this time r The corresponding second polymer set P2 is formed; the second count value C2, the second trajectory file, the first original molecule set X init , the second candidate molecule set X 2,opt and the second polymer set P2 to the self-assembly index evaluation module; The self-assembly index evaluation module is further configured to receive the second count value C2, the second trajectory file, and the first original molecule set X init , the second candidate molecule set X 2,opt and the second aggregate set P2, each second aggregate p in the second aggregate set P2 is generated based on the second trajectory file. r Calculate multiple types of self-assembly indices to obtain the corresponding second polymer index sequence E r ; and all the second polymer index sequences E obtained r The corresponding second evaluation index set GE2 is formed; and the first original molecule set X init , the second candidate molecule set X 2,opt , the second polymer set P2 and the second evaluation index set GE2 are used as the corresponding first original molecule set field, the first candidate molecule set field, the first polymer set field and the first evaluation index set field to form a corresponding first evaluation record and add it to the local evaluation record list; and judge whether the second count value C2 is less than the simulation number threshold parameter; if the second count value C2 is greater than or equal to the simulation number threshold parameter, the evaluation record list is extracted as the corresponding first evaluation record list and sent to the self-assembly short peptide output module, and the local evaluation record list is cleared; if the second count value C2 is less than the simulation number threshold parameter, each second polymer p r and the corresponding second polymer index sequence E r As a set of corresponding first training polypeptide molecules and first label indicator sequences, the obtained first training polypeptide molecules and the first label indicator sequences constitute a corresponding first self-training data, and all the obtained first self-training data constitute a corresponding first batch data set added to the self-training data set of the self-training module; the second evaluation index set GE2 includes multiple second polymer indicator sequences E r The second polymer index sequence E r Including multiple second aggregate indicators e r,h , the index h is an integer greater than 0; the second aggregate index e r,h The index types include aggregation strength index, water solubility index and structural order index; The self-assembling short peptide output module is used to poll each of the first evaluation records in the first evaluation record list in sequence when receiving the first evaluation record list; and when polling, the first evaluation record currently polled is used as the corresponding current evaluation record, and the first polymer set field and the first evaluation index set field of the current evaluation record are extracted as the corresponding current polymer set P and current evaluation index set GE; and each polymer index sequence E in the current evaluation index set GE is traversed; and when traversing, the currently traversed polymer index sequence E is used as the corresponding current polymer index sequence, and when all polymer indicators in the current polymer index sequence exceed the preset index threshold of their respective corresponding types, the current polymer is traversed. The combined index sequence is recorded as the corresponding first qualified index sequence, and the polymer p corresponding to the current polymer index sequence in the current polymer set P is recorded as the corresponding first qualified polymer, and when the molecular structure type of the first qualified polymer does not match the output structural morphological parameters, the first qualified polymer is converted into a one-dimensional, two-dimensional or three-dimensional structure corresponding to the structural morphological parameters and the conversion result is used as the new first qualified polymer, and it is identified whether the first qualified polymer is a short peptide. If so, the current first qualified polymer is used as a corresponding first short peptide; and when all the first evaluation records in the first evaluation record list are polled, the corresponding short peptide prediction report is composed of all the obtained first short peptides and output.

5. The method for processing the self-assembly short peptide prediction model according to claim 4, wherein: The first self-training data set is used to train each of the machine learning models M i The first prediction model performs a round of model training, specifically including: Step 51: the self-training module extracts the first self-training data from the first self-training data set as the corresponding current training data; and uses the currently trained machine learning model M i As the corresponding current machine learning model; Step 52: The self-training module extracts the first training polypeptide molecule and the first label indicator sequence from the current training data as the corresponding current training molecular structure and current label indicator sequence; Step 53: The self-training module inputs the current training molecular structure into the first main control unit of the current machine learning model; and the first main control unit inputs the current training molecular structure into the feature encoding module for encoding to obtain the corresponding first training encoding tensor e tr Send to the first prediction model of the current machine learning model; and the first prediction model encodes the first training tensor e tr Perform multi-class self-assembly index prediction to obtain the corresponding first training prediction index sequence E tr Send back to the first main control unit; and the first main control unit sends the first training prediction indicator sequence E tr Sending back to the self-training module; Step 54: the self-training module converts the first training prediction indicator sequence E tr Substituting the current label indicator sequence into a preset first loss function to calculate a corresponding first loss value; In step 55, the self-training module determines whether the first loss value satisfies a preset first loss value range. If the first loss value satisfies the first loss value range, the self-training module determines whether the current training data is the last first self-training data in the first self-training data set. If the current training data is the last first self-training data set, the process proceeds to step 56. If the current training data is not the last first self-training data set, the self-training module extracts the next first self-training data in the first self-training data set as the new current training data and returns to step 52 to continue training. If the first loss value does not satisfy the first loss value range, the model parameters of the first prediction model of the current machine learning model are modulated, and when the modulation is completed, the process returns to step 53 to continue training. Step 56: The self-training module confirms the machine learning model M corresponding to the current machine learning model i This round of training is over.

6. A device for executing the processing method of the self-assembling short peptide prediction model according to any one of claims 1 to 5, characterized in that: The device includes: a model building module, a data receiving module, a model configuration module and a model application module; The model construction module is used to construct a self-assembling short peptide prediction model; the self-assembling short peptide prediction model includes a molecular dynamics simulation module, a self-assembly index evaluation module and a self-assembling short peptide output module; the molecular dynamics simulation module is preset with a configurable force field type parameter; the self-assembly index evaluation module is preset with a configurable simulation number threshold parameter; the self-assembly short peptide output module is preset with a configurable output structure morphology parameter; The data receiving module is used to receive a first polypeptide molecule set, a first simulation number threshold, and a first output structure type input by a user; the first polypeptide molecule set includes a plurality of first polypeptide molecules; the molecular structure types of the first polypeptide molecules include one-dimensional structure, two-dimensional structure, and three-dimensional structure; the first output structure type includes one-dimensional structure, two-dimensional structure, and three-dimensional structure; The model configuration module is used to set the force field type parameter of the molecular dynamics simulation module of the self-assembly short peptide prediction model to the PACE force field, and set the simulation number threshold parameter of the self-assembly index evaluation module to the corresponding first simulation number threshold, and set the output structure morphology parameter of the self-assembly short peptide output module to the corresponding first output structure type; The model application module is used to predict the structure of a self-assembled short peptide chain according to the first polypeptide molecule set using the self-assembled short peptide prediction model to obtain a corresponding first short peptide prediction report.

7. An electronic device, characterized in that: include: memory, processors, and transceivers; The processor is configured to be coupled to the memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 5; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Constructing method of predicting model, predicting method of peptide synthesis difficulty and device

    CN111383721A

  • Method or system for predicting secreted protein sequence based on transcriptome sequencing

    CN115424667A