Polypeptide self-assembly nanofiber design system and method based on deep learning

CN122224292BActive Publication Date: 2026-08-21JILIN JIANZHU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610669760.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-08-21
Estimated Expiration
2046-05-15

AI Technical Summary

Technical Problem

[0006]本发明解决了现有技术无法从功能需求出发直接设计多肽纳米纤维序列,缺乏针对多肽纳米纤维自组装特性及药物释放性能的协同优化能力,研发周期长、实验成本高的问题

Benefits of technology

本发明通过构建用户输入模块、多肽序列生成模块、结构预测模块、自组装行为预测模块、药物释放预测模块、多目标优化模块和输出模块的完整闭环系统,从功能需求输入到最优序列输出全程无需人工干预,大幅提高设计效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122224292B_ABST
    Figure CN122224292B_ABST
Patent Text Reader

Abstract

The application discloses a polypeptide self-assembly nanofiber design system and method based on deep learning, and particularly relates to an intelligent polypeptide nanofiber design and optimization system based on artificial intelligence.The system solves the problems that the prior art cannot directly design a polypeptide nanofiber sequence from a functional requirement, lacks a synergistic optimization capability for polypeptide nanofiber self-assembly characteristics and drug release performance, and has a long research and development cycle and a high experimental cost.The design system comprises an input module, a polypeptide sequence generation module, a structure prediction module, a self-assembly behavior prediction module, a drug release prediction module and a multi-objective optimization module.The system predicts a polypeptide structure and extracts geometric features through deep learning, and iteratively generates an optimized polypeptide sequence meeting multiple performance targets in combination with a multi-objective optimization algorithm.The application is suitable for fields of intelligent drug delivery systems, tissue engineering and regenerative medicine, and research and development of functional polypeptide nanomaterials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of biomaterials design and artificial intelligence technology, specifically to an artificial intelligence-based intelligent design and optimization system for polypeptide nanofibers. Background Technology

[0002] Peptide nanofibers are one-dimensional nanostructure materials formed by the self-assembly of short peptide sequences. They have excellent biocompatibility, biodegradability and structural designability, and show broad application prospects in drug delivery, tissue engineering and wound healing.

[0003] The design of traditional peptide nanofibers mainly relies on empirical rules and trial-and-error experiments, which has the following problems: First, the design cycle is long, often taking months or even years from concept to experimental verification; second, the experimental cost is high, requiring a lot of synthesis and characterization work; third, the design space is limited, making it difficult to systematically explore possible sequence combinations; fourth, multi-objective optimization is difficult, making it difficult to simultaneously meet multiple performance requirements such as drug loading, release rate, and stability.

[0004] In recent years, artificial intelligence (AI) technology has made significant progress in the field of biomolecular design. Deep learning models can learn complex mapping relationships from large amounts of sequence-function data, and their prediction accuracy has approached experimental levels. However, existing AI methods are mainly applied to protein design, and intelligent design systems for the specific application scenario of peptide nanofibers have not yet been reported.

[0005] In summary, the core problems faced by existing technologies are: they cannot directly design peptide nanofiber sequences based on functional requirements; traditional trial-and-error methods rely on empirical rules and are inefficient; and existing AI methods are only applicable to protein design and lack the ability to synergistically optimize the self-assembly characteristics and drug release performance of peptide nanofibers. This results in long R&D cycles and high experimental costs, making it difficult to meet the rapid R&D needs of novel drug delivery materials. Summary of the Invention

[0006] This invention solves the problems of existing technologies that cannot directly design peptide nanofiber sequences based on functional requirements, lack synergistic optimization capabilities for peptide nanofiber self-assembly characteristics and drug release performance, and have long research and development cycles and high experimental costs.

[0007] A deep learning-based peptide self-assembly nanofiber design system, comprising the following modules: The input module is used to input the target functional properties, constraints, and drug parameter configurations of the peptide self-assembled nanofibers to be designed. The peptide sequence generation module is used to generate a set of candidate peptide sequences based on the target functional properties and constraints, using a sequence generation strategy. The structure prediction module is used to perform secondary structure prediction and tertiary structure prediction for each candidate polypeptide sequence in the candidate polypeptide sequence set, and is also used to extract geometric features related to self-assembly based on the prediction results and generate structural feature data. The self-assembly behavior prediction module is used to predict the self-assembly ability of peptides and nanofiber morphology parameters based on the structural feature data, and generate self-assembly performance indicators. The drug release prediction module is used to predict the drug loading rate and release kinetic parameters of the polypeptide nanofibers under the drug parameter configuration based on the structural feature data, and generate drug release performance indicators. The multi-objective optimization module is used to iteratively optimize the candidate peptide sequence set using a multi-objective optimization algorithm with the goal of maximizing the self-assembly performance index and the drug release performance index, and output an optimized peptide sequence set that satisfies multiple performance objectives.

[0008] In a further preferred embodiment, the target functional properties include hydrophobicity, charge, and polarity; the constraints include sequence length range, forbidden mode detection, and essential amino acid verification.

[0009] A further preferred approach is to include probability sampling generation, beam search optimization, constraint verification, and functional motif generation in the sequence generation strategy.

[0010] In a further preferred embodiment, the structure prediction module also includes a Transformer encoding unit, which is used to perform global context encoding on the polypeptide sequence, extract high-dimensional feature representations, and use them as input features for secondary structure prediction and tertiary structure prediction, respectively.

[0011] In a further preferred embodiment, the secondary structure prediction employs a deep convolutional neural network, using the high-dimensional feature representation output by the Transformer encoding unit as input, and outputting the probability of the secondary structure type for each residue position; the secondary structure type includes polypeptides. -spiral, - Content of folds and random curls.

[0012] In a further preferred embodiment, the tertiary structure prediction employs a graph neural network, based on the high-dimensional feature representation output by the Transformer encoding unit, using amino acids as nodes and inter-residue interactions as edges to predict the three-dimensional coordinates of atoms.

[0013] In a further preferred embodiment, the self-assembly performance indicators include the critical concentration for peptide self-assembly, nanofiber diameter, fiber length distribution, and morphological stability.

[0014] A further preferred option is that the multi-objective optimization algorithm is an improved non-dominated sorting genetic algorithm NSGA-II.

[0015] In a further preferred embodiment, the multiple performance objectives include: maximizing self-assembly capability, maximizing drug loading rate, achieving target release rate, minimizing sequence length, and maximizing synthesis feasibility.

[0016] A deep learning-based design method for peptide self-assembly nanofibers, wherein the design method is as follows: S1. Based on the target functional properties and constraints of the self-assembled peptide nanofibers to be designed, a sequence generation strategy is used to generate a set of candidate peptide sequences. S2, perform secondary and tertiary structure prediction on each candidate polypeptide sequence in the candidate polypeptide sequence set, and generate secondary structure prediction results and tertiary structure prediction results; S3. Based on the prediction results of the secondary structure and the tertiary structure, extract the geometric features related to self-assembly and generate structural feature data. S4. Based on the structural feature data, predict the self-assembly ability of the peptide and the morphological parameters of the nanofiber, and generate self-assembly performance indicators. S5. Based on the structural feature data, predict the drug loading rate and release kinetic parameters of the polypeptide nanofibers under the drug parameter configuration, and generate drug release performance indicators. S6, with the optimization objectives of maximizing the self-assembly performance index and the drug release performance index, the candidate peptide sequence set is iteratively optimized using a multi-objective optimization algorithm to output an optimized peptide sequence set that satisfies multiple performance objectives.

[0017] The beneficial effects of this invention compared to the prior art are as follows: This invention constructs a complete closed-loop system comprising a user input module, a peptide sequence generation module, a structure prediction module, a self-assembly behavior prediction module, a drug release prediction module, a multi-objective optimization module, and an output module. From functional requirement input to optimal sequence output, no manual intervention is required, significantly improving design efficiency.

[0018] This invention employs a non-dominated sorting genetic algorithm through a multi-objective optimization module, using self-assembly capability, drug loading rate, release rate, sequence length, and synthesis feasibility as optimization objectives. It can simultaneously optimize multiple competing performance indicators and provide a Pareto optimal solution set for users to choose from.

[0019] This invention employs a deep learning model trained on large-scale experimental data through a structure prediction module, a self-assembly behavior prediction module, and a drug release prediction module, achieving a prediction accuracy of over 85%, thus providing a reliable quantitative basis for peptide sequence design.

[0020] This invention uses a results visualization module to display the Pareto front, convergence curve, and performance distribution of candidate sequences during the optimization process, providing detailed performance predictions of candidate sequences and visualization of the optimization process, which facilitates user understanding and decision-making.

[0021] This invention supports online model updates through a model training module, continuously improving prediction performance as new experimental data accumulates, enabling the system to self-evolve.

[0022] The design system and method described in this invention are applicable to fields such as intelligent drug delivery systems, tissue engineering and regenerative medicine, and the development of functional peptide nanomaterials. Attached Figure Description

[0023] Figure 1 This is the overall architecture diagram of the deep learning-based peptide self-assembly nanofiber design system described in Implementation Method 1. Figure 2 This is a flowchart of the polypeptide sequence generation process described in Implementation Method 2; Figure 3 It is the structural prediction diagram described in Implementation Method 5; Figure 4 It is the drug release curve diagram described in Implementation Method 8; Figure 5 It is the convergence curve of the optimization process described in Implementation Method Nine; Figure 6 It is the Pareto front diagram for multi-objective optimization as described in Implementation Method 10; Figure 7 The image shows the prediction results of the candidate sequence (RGD-KLVFFA-GG) of the wound healing polypeptide nanofiber described in Implementation Method Thirteen; (a) the three-dimensional spatial coordinates and prediction score phantom distribution of the candidate sequence RGD-KLVFFA-GG; (b) the comparison of cell compatibility scores of different sequences; and (c) the comprehensive performance prediction parameter table of the candidate sequence. Detailed Implementation

[0024] Various embodiments of the present invention will now be clearly and completely described with reference to the accompanying drawings. The embodiments described with reference to the drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0025] Implementation Method 1: This implementation method provides a deep learning-based peptide self-assembly nanofiber design system, such as... Figure 1 As shown, the design system includes the following modules: The input module is used to input the target functional properties, constraints, and drug parameter configurations of the peptide self-assembled nanofibers to be designed. The peptide sequence generation module is used to generate a set of candidate peptide sequences based on the target functional properties and constraints, using a sequence generation strategy. The structure prediction module is used to perform secondary structure prediction and tertiary structure prediction for each candidate polypeptide sequence in the candidate polypeptide sequence set, and is also used to extract geometric features related to self-assembly based on the prediction results and generate structural feature data. The self-assembly behavior prediction module is used to predict the self-assembly ability of peptides and nanofiber morphology parameters based on the structural feature data, and generate self-assembly performance indicators. The drug release prediction module is used to predict the drug loading rate and release kinetic parameters of the polypeptide nanofibers under the drug parameter configuration based on the structural feature data, and generate drug release performance indicators. The multi-objective optimization module is used to iteratively optimize the candidate peptide sequence set using a multi-objective optimization algorithm with the goal of maximizing the self-assembly performance index and the drug release performance index, and output an optimized peptide sequence set that satisfies multiple performance objectives.

[0026] Implementation Method Two: This implementation method further defines Implementation Method One, and the design system further includes: The results visualization module is connected to the multi-objective optimization module. The results visualization module is used to display the Pareto front, convergence curve and performance distribution of candidate sequences during the optimization process. The knowledge base module stores the sequence feature database and the structure performance mapping library, and supports similar sequence retrieval. The API interface module provides a RESTful API, including a sequence design interface and a structure prediction interface.

[0027] Implementation Method 3: This implementation method further defines Implementation Method 1 and provides examples to illustrate the target functional attributes and constraints.

[0028] The target functional properties include hydrophobicity, charge, and polarity; the constraints include sequence length range, forbidden mode detection, and essential amino acid verification.

[0029] Implementation Method Four: This implementation method further defines Implementation Method One and provides an example to illustrate the sequence generation strategy.

[0030] The sequence generation strategy includes probability sampling generation, beam search optimization, constraint verification, and functional motif generation; This implementation introduces an amino acid function preference matrix, which calculates the selection probability of each amino acid based on the target attribute, while also considering sequence constraints to avoid generating overly similar candidate sequences.

[0031] Specifically, such as Figure 2 As shown, the data processing flow of the polypeptide sequence generation module is as follows: First, input the target functional attributes and constraints. Calculate the amino acid functional preference matrix based on the target functional attributes to obtain the selection probability of each amino acid at each position, forming an amino acid selection probability matrix.

[0032] Then, a generation strategy is selected: one or more of the following methods are used to generate initial candidate sequences based on the amino acid selection probability matrix: probability sampling generation, beam search optimization generation, or functional motif generation.

[0033] After generating candidate sequences, constraint verification is performed: check whether the sequence length of the candidate sequence is within the preset range, whether it contains prohibited patterns, whether it meets the requirements of essential amino acids, and perform diversity checks to avoid generating overly similar sequences.

[0034] If the constraint verification passes, the score of the candidate sequence is calculated, which includes the self-assembly tendency score and the functional matching score, and the candidate sequence that meets the conditions is output; if the constraint verification fails, the sequence is regenerated until the obtained candidate sequence meets the constraint conditions or reaches the preset iteration limit.

[0035] Implementation Method 5: This implementation method further defines Implementation Method 1 and provides an example of the structure prediction module.

[0036] The structure prediction module includes a Transformer coding unit, such as... Figure 3 As shown, the Transformer encoding unit serves as a shared feature extraction layer, used to perform global context encoding on the input polypeptide sequence and extract high-dimensional feature representations of the sequence rich in long-range dependency information.

[0037] The network hierarchy of the Transformer coding unit includes an input embedding layer, a position coding layer, and a main body consisting of six stacked encoder layers. Each encoder layer contains a multi-head self-attention sub-layer, a feedforward fully connected sub-layer, residual connections, and layer normalization, and is finally connected to the output layer.

[0038] The input embedding layer will input the polypeptide amino acid sequence (length is...) Each amino acid in the matrix is ​​mapped to a 512-dimensional continuous vector representation, with an embedding matrix of dimension |V|×512, where |V| is the size of the amino acid vocabulary (containing 20 standard amino acids and special markers). <cls> 、 <pad> 、 <unk>(A total of 23 terms); the embedding layer weights are initialized using a uniform distribution using Xavier.

[0039] The position coding layer uses sinusoidal position coding with a maximum coding length. =128 amino acids, positional encoding dimension equals total attention dimension =512 (This dimension is determined by the number of attention heads) =8 with key / value / query dimensions for each attention head =64, that is = × =512). For positions in the sequence and dimensions The location code calculation formula is: Even-numbered dimensions:

[0040] Odd-numbered dimensions:

[0041] The positional encoding and the input embedding are added element-wise and then input into the encoder in the multi-head self-attention sublayer.

[0042] The Transformer encoding unit is configured with six encoder layers, each containing a multi-head self-attention sublayer and a feedforward fully connected sublayer. In the multi-head self-attention mechanism, the number of attention heads... =8, key / value / query dimensions for each attention head: / / Both are 64, total attention dimension = × =512. Each self-attention sublayer uses scaled dot product attention computation:

[0043] in, For query matrix, the dimension is ; The key matrix has dimensions of . ; It is a value matrix with dimension . .

[0044] After each attention head is calculated, The outputs of each size are spliced ​​together, and then the output projection matrix is ​​used. ∈ Map back In dimensional space, the Dropout ratio for each self-attention sublayer is... =0.1.

[0045] The feedforward fully connected sublayer contains two linear transformations, with the GELU activation function used in between. Let the input be... (dimension) =512), then the calculation formula for the feedforward network is:

[0046] in, This is the first-layer weight matrix, with dimension . , =2048 is the intermediate layer dimension of the feedforward network; The first layer bias has a dimension of ; This is the second-layer weight matrix, with dimension . , This is the second layer bias, with dimension . The dropout ratio of the feedforward sublayer is... =0.1.

[0047] Residual Connectivity and Layer Normalization: Each sub-layer (including multi-head self-attention sub-layers and feedforward fully connected sub-layers) adopts a structure of residual connectivity followed by layer normalization. The specific calculation rules are as follows: The sub-layer input first undergoes feature transformation of the corresponding sub-layer, then is added to the original input through residual connectivity, and finally, layer normalization is performed on the added result to obtain the final output of the sub-layer. Its mathematical expression is:

[0048] in, This is the input feature vector of the current sub-layer. This is the sublayer transformation function, corresponding to either multi-head self-attention transformation or feedforward fully connected transformation. This is a residual connection operation used to alleviate the vanishing gradient problem during deep neural network training. This is a layer normalization function that standardizes the dimension of the feature vectors to accelerate model training convergence and improve training stability.

[0049] The layer normalization function is:

[0050] in, The input vector is the layer-normalized vector, which is the result of summing the residuals. , For input vectors The mean across the feature dimensions, For input vectors Standard deviation in the feature dimension This is an element-wise multiplication operation. and The learnable scaling and offset parameters of the model, the dimensions of the input vector Consistent It is a very small constant, with a value of 1 × 10⁻⁶. -5 This is used to prevent numerical calculation errors caused by a denominator of zero.

[0051] The aforementioned residual connections and layer normalization structures are uniformly reused in the multi-head self-attention sublayer and the feedforward fully connected sublayer, ensuring that the Transformer coding unit can stably extract the global context features and long-range dependency information of the peptide sequence.

[0052] Output layer: The final output dimension of the Transformer encoding unit is... The high-dimensional feature matrix of ×512 contains rich long-range dependence of amino acid sequences and global context information. This feature matrix will be supplied as shared input features to the downstream secondary structure prediction unit (deep convolutional neural network) and tertiary structure prediction unit (graph neural network), respectively, thereby replacing the traditional single one-hot encoding or simple physicochemical feature splicing and significantly improving the accuracy of subsequent structure prediction.

[0053] Implementation Method Six: This implementation method further defines Implementation Method One and provides examples to illustrate the secondary and tertiary predictions.

[0054] The structure prediction module also includes the following units: The secondary structure prediction unit, based on a deep convolutional neural network, takes the high-dimensional feature representation output by the Transformer coding unit as input and outputs the probability of the secondary structure type of residues at each position; the secondary structure type includes polypeptides. -spiral, - Content of folds and random curls; The secondary structure prediction unit uses the following dataset: The main training set uses the CB513 public dataset, which contains 513 rigorously selected non-redundant protein chains (sequence identity ≤25%). Secondary structure annotations for each chain were extracted from PDB crystal structure files using the DSSP program, and the annotation categories were uniformly mapped to three classes: H (…). - spiral), E ( - Folded) and C (random curl, including corners and loops). The main training set contains 410 training samples, 51 validation samples, and 52 test samples (8:1:1 ratio).

[0055] Expand the training set: Screen for protein structures from the Protein Data Bank (PDB) that meet the following criteria: resolution of 2.5 angstroms; Factor ≤ 0.25; protein chain length between 20 and 500 amino acid residues; determined by X-ray crystallography. After redundancy removal screening with sequence identity ≤ 30% using the CD-HIT program, approximately 11,500 protein chains were obtained as an expanded training set; the expanded training set contains 9,200 proteins in the training set, 1,150 in the validation set, and 1,150 in the test set (8:1:1 ratio).

[0056] The above dataset was preprocessed as follows: each amino acid residue in all sequences was labeled with three types of secondary structures (H / E / C) by the DSSP program; the amino acid sequences were uniformly converted into single-letter code representations (a total of 20 standard amino acid characters); zero-padding was performed on the ends of sequences with less than 128 residues, and sequences with more than 128 residues were randomly truncated into fragments of 128 consecutive residues.

[0057] The deep convolutional neural network adopts the following network hierarchy structure: Input layer: Receives the output of the Transformer encoding unit in dimension 1. A high-dimensional feature matrix of ×512 is considered as having 512 feature channels and a length of... One-dimensional signal input. Sequence length. If the number of amino acids is less than 128, fill it with zero to bring it up to 128; if the number of amino acids exceeds 128, cut off the first 128 amino acids.

[0058] Initial convolutional layer: one-dimensional convolution, number of input channels =512, Number of output channels C=256, Convolution kernel size =7, step size =1, padding method is "same" to maintain the sequence length. Batch normalization (momentum parameter) is sequentially connected after the convolutional layer. =0.9, =1×10 -5 ) and the ReLU activation function.

[0059] Residual Convolutional Blocks: A total of 4 residual convolutional blocks are stacked. Each residual block contains two sub-convolutional layers. Each sub-convolutional layer is a one-dimensional convolution (with 256 input and output channels and a kernel size of...). =3, stride 1, padding "same") → batch normalization → ReLU activation; residual connections add the input of the residual block to the output of the second sub-convolutional layer element-wise, then activate via ReLU. The dropout ratio of each convolutional layer within each residual convolutional block is... =0.2.

[0060] Multi-scale dilated convolutional layers: Two dilated convolutional layers operate in parallel on the output feature map of the fourth residual block, with dilation rates of [missing information]. =2 and =4, kernel size =3, both input and output channels are 256, with adaptive padding to maintain the sequence length; each of the two dilated convolutional layers is followed by batch normalization and ReLU activation, and then the two outputs are concatenated along the channel dimension to obtain a dimension of ×512 multi-scale fusion feature map.

[0061] Global Feature Convergence Layer: Global average pooling and global max pooling are performed along the sequence dimensions of the fused feature map to obtain two global feature vectors, each with a dimension of 512. These are then concatenated into a 1024-dimensional vector; simultaneously, the feature map is preserved. ×512-dimensional positional features.

[0062] Output layer: Maps positional features to a dimension of 3 through a two-layer fully connected network (512→256, ReLU→256→3). The ×3 output layer outputs a three-dimensional vector for each amino acid residue position. After normalization using the Softmax function, this vector represents the probability distribution of the secondary structure type of that residue, corresponding to... -helix (H) - Folding (E) and random curling (C). The Dropout ratio of fully connected layers is =0.3.

[0063] After the deep convolutional neural network model is trained, its performance is measured on an independent test set using the following evaluation metrics: Tri-state accuracy Q3 (Q3 = number of correctly predicted residues / total number of residues × 100%), fragment overlap score SOV (measures the overall degree of overlap between the predicted secondary structure fragment and the actual fragment, with tolerance for the deviation of the structural fragment boundary).

[0064] On the test set of this implementation, Q3 reaches 87.3% and SOV reaches 82.1%, which is better than the traditional method PSIPRED (Q3 is about 80%) and the baseline CNN model based on a single one-hot encoding (Q3 is about 83%).

[0065] The tertiary structure prediction unit is used to predict the three-dimensional coordinates of atoms based on a graph neural network, using the high-dimensional feature representation output by the Transformer coding unit as the basis, with amino acids as nodes and inter-residue interactions as edges. The three-level structure prediction unit uses the following dataset: Peptide and small protein structures meeting the following criteria were screened from the PDB database: peptide / protein chain length ≤ 50 amino acid residues; resolution ≤ 2.5 Å; determined by X-ray crystallography or cryo-electron microscopy. After redundancy removal using the CD-HIT program with sequence consistency ≤ 40%, 8200 peptide 3D structure data were obtained. The dataset was randomly divided in an 8:1:1 ratio into a training set of approximately 6560 structures, a validation set of approximately 820 structures, and a test set of approximately 820 structures.

[0066] The above dataset was preprocessed as follows: the three-dimensional coordinates (unit: Å) of the main chain atoms (N, Ca, C, O) of each polypeptide structure were extracted as training labels; the coordinates were centered (the centroids of all atoms were shifted to the origin); and alpha carbon virtual atoms were filled into sequences with a length of less than 50 residues.

[0067] The graph neural network adopts the following network hierarchy structure: Node Feature Construction Layer: Graph nodes are defined as each amino acid residue in the polypeptide sequence; the initial feature vector of each node is composed of two concatenated parts: (a) Physicochemical properties (10 dimensions): Kyte-Doolittle hydrophobicity, net charge (pH 7.4), polarity, molecular weight, side chain volume, isoelectric point, number of hydrogen bond donors, number of hydrogen bond acceptors, aromaticity, and Krigbaum flexibility index. The characteristic values ​​of 20 standard amino acids are shown in Table 1. Table 1

[0068] The above-mentioned physicochemical descriptor values ​​can be calculated based on amino acid physicochemical property databases such as AAindex and feature extraction tools such as protlearn. The values ​​from different sources may have slight differences, but they do not affect the implementation of the technical solution of this invention. (b) Transformer Encoding Features (512-dimensional): The feature vector (512-dimensional) corresponding to the amino acid residue position is extracted from the high-dimensional feature matrix output by the Transformer encoding unit. Initial feature dimension of the total node. =522.

[0069] Edge Feature Construction Layer: Graph edges are defined as the interaction relationships between any two amino acid residues, constructing a fully connected graph (i.e., there is an edge between any two nodes). The initial feature vector of each edge contains three components: the Gaussian radial basis function expansion of the Ca-Ca distance (with a step size of 0.5 Å, ranging from 0 to 20 Å, a total of 40 dimensions), the One-Hot encoding of the sequence interval distance (maximum interval 128, a total of 129 dimensions), and the normalized value of the BLOSUM62 substitution matrix score (1 dimension). After concatenation, these components are mapped to a 128-dimensional initial feature vector of the edge through a learnable edge embedding layer (Linear(170→128) + ReLU).

[0070] Message Passing Layers: There are 4 stacked message passing layers. Each message passing layer performs the following three steps: Step A: Edge Feature Update: For each edge (connecting nodes) and nodes ), will node Features ,node Features With current edge features The concatenation results in a (522+522+128=1172) dimensional vector, which is then input into a two-layer MLP (1172→384→128, with ReLU activation in the intermediate layer) to generate updated edge features. .

[0071] Step B: Message Aggregation: Variant Computation Nodes Employing Multi-Head Attention Mechanism The message vector. Compute node With each of its neighboring nodes Attention weights between First, the nodes are transformed using a linear transformation. Features, Nodes Features and edge features Map these to the 64-dimensional query, key, and value spaces respectively, and then compute:

[0072] in, To query the linear transformation matrix, the nodes are... The features are mapped to a 64-dimensional query space; The key linear transformation matrix will transform the nodes. The features are mapped to a 64-dimensional query space; The linear transformation matrix for edge features maps the edge features to a 64-dimensional bond space; This is a scaling factor, equal to the dimension of the query / key vector, used to stabilize the gradient; For nodes All neighboring nodes Normalization yields attention weights. .

[0073] After calculating the attention weights, the value vectors of each neighbor node are summed according to the attention weights to obtain the node. Aggregated message vector The attention dropout ratio for message aggregation is 0.1.

[0074] Step C: Node Feature Update: Update the node Current features With aggregated message vectors The input is a gated recurrent unit (GRU) for updating; the hidden dimension of the GRU is 256. Updated node features. This serves as the input for the next layer of message passing. The hidden dimension of nodes in each layer of message passing is 256, and the hidden dimension of edges is 128.

[0075] Coordinate Prediction Header: The node features (256-dimensional) output from the 4th layer message passing are mapped to the 3D coordinate offsets of the main chain atoms of each amino acid residue through a three-layer fully connected network (256→128, ReLU→128→64, ReLU→64→12). Specifically, this involves 12 (x, y, z) coordinates for the four main chain atoms N, Ca, C, and O. An independent fully connected branch (256→64→8) is appended after the coordinate prediction head, outputting the sine and cosine values ​​(8-dimensional) of the four torsion angles (φ, ψ, χ1, χ2) of each residue's side chain, used to reconstruct the complete side chain spatial conformation.

[0076] Geometric Constraint Layer: Based on the coordinate prediction head, a differentiable geometric constraint layer is added. This layer imposes bond length constraints (Ca-N: 1.47±0.03, Ca-C: 1.53±0.03Å, CN: 1.33±0.03Å, C=O: 1.23±0.03Å) and bond angle constraints on the predicted atomic coordinates, ensuring the physical plausibility of the predicted 3D structure. This constraint is presented as a soft constraint and serves as the regularization term of the loss function.

[0077] After the model is trained, its performance is measured on an independent test set using the following evaluation metrics: The root mean square deviation (RMSD) is the root mean square deviation between the predicted and actual coordinates of the Ca atom, in angstroms, and the template modeling score (TM-score) is measured (ranging from 0 to 1, comprehensively evaluating the accuracy of the overall folding topology of the peptide; a TM-score > 0.5 is considered a correct fold). On the test set of this system, the average Ca RMSD is 2.8 angstroms, and the average TM-score is 0.76.

[0078] The Transformer encoding unit (sharing the underlying feature extraction layer) is jointly pre-trained end-to-end with the second-level prediction head (deep convolutional neural network) and the third-level prediction head (graph neural network). The training configuration is as follows: (1) Loss function design: Joint multi-task loss function It consists of a weighted sum of three components: Component 1: Secondary Structure Prediction Loss We employ a weighted cross-entropy loss method to calculate the cross-entropy between secondary structure tri-classification prediction and DSSP annotation for each amino acid residue position. Category weight settings: - Spiral (H) weight =1.2, -Folding (E) weight =1.2, random curl (C) weight =1.0, to alleviate the class imbalance problem caused by the large proportion of random curls in the training set. The calculation formula is:

[0079] in For sequence length, For position Category The true label (one-hot) For position Category The predicted probability.

[0080] Component 2: Predicted Loss of Three-Level Structure The frame alignment point error loss is used as the main loss to calculate the distance error between the predicted and actual atomic coordinates in the alignment reference frame. The FAPE loss calculation process is as follows: for each amino acid residue position, using the local coordinate system defined by the actual Ca, C, and N atoms as a reference, the deviation of the predicted atomic coordinates relative to this reference frame is calculated. The Clamp function is used to limit the single-point error to within 10 angstroms to enhance training stability. Simultaneously, a main chain bond length and bond angle constraint loss is added. As a regularization term:

[0081] in The predicted bond lengths of each chemical bond, The actual bond lengths of each chemical bond. For the predicted bond angles of each bond angle, These are the actual key angles for each key angle. and These are the constraint weights for each bond and bond corner.

[0082] Total loss:

[0083] in =0.5 is the balancing weight of the third-level structure loss relative to the second-level structure loss. =0.1 is the geometric constraint regularization weight, and this ratio is the optimal value determined by grid search.

[0084] (2) Optimizer configuration: The AdamW optimizer is used to estimate the exponential decay rate using the first moment. =0.9, second moment estimate of exponential decay rate =0.999, numerical stability constant =1×10 -8 The weight decay coefficient is 0.01. Weight decay only applies to the weight matrix parameters, not to the bias terms and layer normalization parameters.

[0085] (3) Learning rate scheduling strategy: A cosine annealing learning rate scheduling strategy is adopted. Initial learning rate =1×10 -4 The linear warm-up steps are 2000 (during the warm-up phase, the learning rate linearly increases from 0 to...). Minimum learning rate =1×10 -6 The learning rate decays cosine after each training epoch:

[0086] in For the current round, This refers to the total number of training rounds.

[0087] (4) Batch size and training epochs: Batch size = 64 peptide sequences. When the sequence lengths within a batch are inconsistent, the longest sequence in the batch is used for zero-padding to a uniform length. Simultaneously, a padding mask matrix is ​​constructed to ignore the contribution of padding positions during self-attention calculation. Maximum training epochs. =200 rounds. Early stopping strategy adopted: monitoring the total loss on the validation set. The patience value is set to 15 rounds, meaning that if the validation set loss does not decrease for 15 consecutive rounds (the minimum difference between the current validation loss and the historical best validation loss is less than 1 × 10), then the validation set loss will be considered valid. -4 If the criterion is "no decline observed", then training is terminated and the model parameter checkpoint with the minimum validation loss is restored.

[0088] (5) Data augmentation strategy: Online data augmentation is performed on the training set sequence at the beginning of each training round. Augmentation methods include: random sequence pruning (randomly truncating continuous sub-fragments of 32 to 128 residues in length from the complete protein chain; if the original sequence has less than 32 residues, it is retained as is); amino acid conservative substitution (randomly selecting replacement amino acids from the set of amino acids with a BLOSUM62 substitution matrix score ≥ 0 with an independent probability of 0.05 for each amino acid residue); input feature noise (adding Gaussian noise with a mean of 0 and a standard deviation of 0.02 to the input embedding of the Transformer coding unit to enhance the robustness of the model).

[0089] (6) Pre-training hardware environment and time consumption: Pre-training was performed in data parallelism on 4 NVIDIA A100 (80GB VRAM) GPUs, with a local batch size of 16 sequences per GPU (total equivalent batch size of 64). The training time per round was approximately 12 minutes (training set of approximately 12,000 sequences), and the total training time for 200 rounds was approximately 40 hours.

[0090] The structural feature extraction unit is used to extract geometric features related to self-assembly from the secondary structure prediction results and the tertiary structure prediction results. It serves as a core bridge connecting structural prediction and performance prediction. Its specific composition and calculation method are as follows: Secondary structure content characteristics: -Spiral ratio, - Folding ratio, random curling ratio, etc.; Main chain dihedral distribution characteristics: Angle and Mean, standard deviation, and higher-order statistics of angles, etc.; Molecular surface characteristics: total surface area, hydrophobic surface area, hydrophilic surface area, and charge distribution calculated based on the Shrake-Rupley algorithm; Global characteristics of the sequence: sequence length, average hydrophobicity, net charge, polarity index, etc. Three-dimensional structural features: radius of gyration, molecular volume calculated based on the Monte Carlo method, shape factor, and number of hydrogen bonds calculated based on the DSSP algorithm, etc. Other features include self-assembly tendency score, number of functional motifs, and synthesis feasibility score. These various original features and their cross-combinations are calculated and concatenated to form a high-dimensional initial feature vector. This vector is then reduced to 256 dimensions using a trained PCA model and standardized using Z-scores to obtain the final 256-dimensional structural feature vector.

[0091] Implementation Method Seven: This implementation method is a further limitation of Implementation Method One, and provides an example of the self-assembly behavior prediction module.

[0092] The self-assembly behavior prediction module includes the following units: Critical aggregation concentration prediction unit, used to predict the concentration threshold at which peptides begin self-assembly; Nanofiber diameter prediction unit, used to predict the diameter range of peptide nanofibers; A fiber length prediction unit is used to predict the length distribution of polypeptide nanofibers. A morphological stability prediction unit is used to predict the stability of peptide nanofibers under different environmental conditions.

[0093] The self-assembly behavior prediction module uses the following dataset: Data source: Data from approximately 450 peptide self-assembly experimental studies published between 2015 and 2025 were manually compiled. Data inclusion criteria included: clearly defined peptide sequences, clearly defined experimental conditions (pH, temperature, ionic strength), and complete self-assembly characterization data (TEM or AFM images, CD spectra).

[0094] Data Content: Each experimental record includes the following fields: peptide amino acid sequence, critical aggregation concentration (CAC) (mM), nanofiber diameter measured by transmission electron microscopy (TEM) (nm), fiber length distribution measured by dynamic light scattering (DLS) or atomic force microscopy (AFM) (D10 / D50 / D90, nm), and morphological stability score (a comprehensive score of TEM morphological consistency at 24-hour and 72-hour time points, ranging from 0 to 1, with higher scores indicating greater stability). Approximately 3600 experimental records were initially collected. After data cleaning (removing outliers: records with negative CAC or greater than 100 mM, and records with fiber diameters less than 1 nm or greater than 200 nm), approximately 2850 valid records were retained.

[0095] The dataset was randomly divided into a training set of approximately 2280 records, a validation set of approximately 285 records, and a test set of approximately 285 records in an 8:1:1 ratio.

[0096] Data augmentation: The following data augmentation operations were performed on each polypeptide sequence in the training set (generating 3 augmented sequences for each original sequence): random truncation (removing continuous sub-fragments of the original sequence with a minimum length of ≥8 residues); conserved amino acid substitution (based on the BLOSUM62 matrix, performing substitutions of the same type of amino acid at a random ratio of 10% for positions with substitution scores ≥0). The total size of the augmented training set is approximately 9120 sequences.

[0097] The self-assembly behavior prediction module receives the high-dimensional sequence feature representation output by the Transformer encoding unit and the structural feature data output by the structure prediction module as joint input for supervised multi-task learning. The training configuration is as follows: Parameter freezing strategy: During the training of the self-assembly behavior prediction module, the weight parameters of the Transformer encoding unit are frozen (gradient updates are not performed), and only the parameters of the self-assembly behavior prediction module itself are updated.

[0098] Loss function design: Total loss across multiple tasks It consists of the sum of the losses from four sub-tasks.

[0099] The critical aggregation concentration prediction uses mean squared error (MSE) loss.

[0100] in, For the first The predicted critical aggregation concentration value for each sample. For the first The true critical aggregation concentration value of each sample; The nanofiber diameter prediction uses mean square error (MSE) loss. Fiber length distribution prediction uses quantile loss. Quantile losses are calculated and summed for D10 (10th quantile), D50 (50th quantile, median), and D90 (90th quantile).

[0101] in, For the first The true value of fiber length for each sample. The model predicts the first Fiber length values ​​for each sample; The morphological stability score prediction uses mean squared error (MSE) loss, with a total loss of:

[0102] No additional weighting coefficients were applied to the losses of each subtask because each subtask is physically equally important and the output scales have been standardized to similar ranges.

[0103] Optimizer and learning rate: The Adam optimizer is used. =0.9, =0.999, =1×10 -8 Initial learning rate =5×10 -5 The learning rate employs a stepped decay strategy: every 50 training epochs, the learning rate is decayed to 0.5 times the current value.

[0104] Batch size and training epochs: Batch size = 32 peptide sequences. Maximum training epochs. =150 rounds, early stop patience value=10 rounds. The early stop monitoring metric is the total multi-task loss on the validation set. .

[0105] Training time: For a single NVIDIA A100 GPU, the training time for one round is approximately 3 minutes, and the total training time is approximately 7.5 hours.

[0106] After the model is trained, its performance is measured on an independent test set using the following evaluation metrics: The coefficient of determination (R²) and mean absolute percentage error (MAPE) for each subtask. On the test set of this embodiment, the critical aggregation concentration (CAC) prediction had R² = 0.89 and MAPE = 12.3%; the nanofiber diameter prediction had R² = 0.85 and MAPE = 9.7%; the fiber length distribution (D50) prediction had R² = 0.81 and MAPE = 15.2%; and the morphological stability score prediction had R² = 0.88 and MAPE = 6.8%.

[0107] This implementation method is based on a published peptide self-assembly dataset and employs a multi-task learning framework to predict multiple related attributes simultaneously, thereby improving prediction accuracy and efficiency.

[0108] Implementation Method 8: This implementation method is a further limitation of Implementation Method 1, and provides an example of the drug release prediction module.

[0109] The drug release prediction module includes the following units: Diffusion coefficient estimation unit, used to estimate the diffusion coefficient of different drug types in peptide nanofibers; The release model fitting unit is used to fit the drug release time curve and calculate the release rate constant and half-life. The release equation is:

[0110] in, The release rate constant, To release the index, For time, For the drug in time The cumulative release percentage at that time.

[0111] The mechanism determination unit is used to determine the type of release mechanism, including diffusion-controlled, degradation-controlled, and pH-responsive mechanisms. The drug release was achieved using the following dataset: Data source: Drug release experimental data were manually compiled from approximately 180 research papers on drug release experiments loaded with peptide nanofibers published between 2015 and 2025. Data inclusion criteria were: clearly defined peptide nanofiber sequence, clearly defined drug type, clearly defined release experimental conditions (37°C, pH 7.4 PBS buffer), and complete release curve data.

[0112] Data Content: Each experimental record includes the following fields: peptide amino acid sequence, loaded drug type (One-Hot encoded, covering 8 commonly used anticancer drugs including doxorubicin (DOX), paclitaxel (PTX), cisplatin (CDDP), curcumin (CUR), and camptothecin (CPT), totaling 8 dimensions), drug loading rate (%w / w), time-cumulative release rate data points of the release time curve (time points include 0.5h, 1h, 2h, 4h, 8h, 12h, 24h, 48h, 72h, 96h, 120h, and 168h, a total of 12 time points, each time point corresponding to a cumulative release percentage value), and the release rate constant k (h) obtained by fitting the Korsmeyer-Peppas model. -n The dataset includes the release index n and release mechanism classification labels (Fick diffusion-controlled, dissolution-controlled, swelling-controlled, pH-responsive, and synergistic, totaling 5 categories). Approximately 1250 experimental records were initially collected. After data cleaning (removing abnormal records with negative drug loading rates or greater than 80% w / w, and records with missing release curve data for more than 3 time points), 980 valid records were retained. The dataset was randomly divided in an 8:1:1 ratio into a training set of 784 records, a validation set of 98 records, and a test set of 98 records.

[0113] The drug release prediction module receives the high-dimensional sequence feature representation output by the Transformer encoding unit and the structural feature data output by the structure prediction module as joint input for supervised learning. The training configuration is as follows: Parameter freezing strategy: During the training of the drug release prediction module, the weight parameters of the Transformer encoding units are frozen, and only the parameters of the drug release prediction module itself are updated.

[0114] Loss function design: Total loss across multiple tasks It consists of the weighted sum of the following components.

[0115] Drug loading rate prediction loss : Use Huber loss (smoothed L1 loss), Huber loss threshold =0.5 (i.e., use the MSE smoothing segment when the absolute value of the prediction error is less than 0.5, otherwise use the MAE linear segment), to balance the fine fitting of small errors with the robustness of large errors:

[0116]

[0117] Diffusion coefficient predicts loss Mean squared error (MSE) loss is used (because the data has been logarithmically transformed after preprocessing to stabilize the variance); Release rate constant Predicting losses Mean squared logarithmic error (MSLE) is used.

[0118] Release Index Predicting losses : Use Smooth L1 loss (smoothing L1, parameter =0.1); Release mechanism classification loss Cross-entropy loss was used to classify the release mechanism types (5 categories).

[0119] Total loss from multitasking:

[0120] The weighting coefficients of each loss component were determined by grid search in a set of independent validation experiments (search range: {0.1, 0.3, 0.5, 1.0, 2.0}).

[0121] Optimizer and learning rate: The AdamW optimizer is used. =0.9, =0.999, =1×10 -8 The weight decay factor is 0.005. The initial learning rate is... =3×10 -5 The ReduceLROnPlateau learning rate scheduling strategy is adopted: the total multi-task loss on the validation set is monitored. The mode is set to "min", and the validation loss is not reduced for 5 consecutive rounds (the reduction threshold is 1×10). -4 When the learning rate is reduced to 0.3 times its current value (decay factor), the learning rate will be reduced to 0.3 times its current value. =0.3), minimum learning rate set to =1×10 -7 .

[0122] Batch size and training epochs: Batch size = 16 sequences (because each sequence needs to be combined with parameters of different drug loading types for joint forward propagation, the computational cost per sample is relatively large). Maximum training epochs. =100 rounds, early stop patience value=8 rounds. The early stop monitoring metric is the total multi-task loss on the validation set. .

[0123] Training time: For a single NVIDIA A100 GPU, the training time per round is approximately 2 minutes, and the total training time is approximately 3.3 hours.

[0124] After the model is trained, its performance is measured on an independent test set using the following evaluation metrics: The coefficient of determination (R²) and mean absolute error (MAE) for each subtask. On the test set of this embodiment, the predicted drug loading rate had R² = 0.87 and MAE = 2.1% w / w; the predicted release rate constant (k) had R² = 0.83 and MAE (log scale) = 0.12; the predicted release index (n) had R² = 0.78 and MAE = 0.06; and the accuracy of the release mechanism classification was 84.6%.

[0125] Figure 4 To illustrate the predicted drug release curves of the generated peptide nanofibers, the horizontal axis represents time (hours), and the vertical axis represents the cumulative release rate (%). The figure shows multiple release curves predicted by the system, including zero-order release, first-order release, the Higuchi model, the KP model, and others, to compare the fitting effects of different release kinetics. The release curve predicted by this system exhibits typical drug sustained-release characteristics: an initial burst release effect of approximately 12%, followed by a steady-release phase. The release rate constant k = 0.0235 h was obtained by fitting the curve using the Korsmeyer-Peppas model. -n The release index n = 0.52 indicates that the drug release mechanism is a synergistic effect of diffusion and dissolution (when 0.45 ≤ n < 0.89). The goodness of fit R² = 0.96 indicates that the model and experimental data are in good agreement. The slope changes and plateau characteristics of the curves can intuitively reflect the initial burst release effect, the sustained release rate in the middle and late stages, and the overall release cycle, providing a direct basis for evaluating whether candidate peptide sequences meet the target release rate.

[0126] This module takes into account the influence of multiple factors such as peptide sequence, nanofiber morphology, drug type, and release environment.

[0127] Implementation Method Nine: This implementation method further defines Implementation Method One and provides an example of the multi-objective optimization algorithm.

[0128] The multi-objective optimization algorithm is an improved non-dominated sorting genetic algorithm NSGA-II; The improved non-dominated sorting genetic algorithm NSGA-II introduces a confidence penalty mechanism based on prediction uncertainty, which is implemented as follows: Mathematical expression for confidence penalty mechanism:

[0129] in, Based on the fitness value, The penalty coefficient is set to 0.5. The confidence level of the model prediction was obtained by calculating the standard deviation of the prediction results through Monte Carlo Dropout sampling 10 times. The multi-objective fitness evaluation function vector is constructed as follows: f1 = - Self-assembly capability score; f2 = -drug loading rate; f3 = |Release Rate - Target Release Rate| / Target Release Rate; f4 = sequence length; f5 = -Synthetic feasibility score; Key algorithm parameters: population size 100, maximum number of iterations 200 generations, crossover rate 0.9 (single-point crossover), mutation rate 0.1 (random replacement of one amino acid), using fast non-dominated sorting algorithm and standard crowding distance calculation method; Optimization process: After initializing the population, perform structure prediction, self-assembly prediction, and drug release prediction for each sequence and calculate the confidence level; apply the confidence level penalty mechanism to calculate the final fitness value; generate a new population through non-dominated sorting and selection, crossover and mutation, and repeat the iteration until the non-dominated solution on the Pareto front is output.

[0130] Compared to the standard NSGA-II algorithm, the improved non-dominated sorting genetic algorithm NSGA-II described in this embodiment is specifically reflected in the introduction of a confidence penalty mechanism based on prediction uncertainty in the fitness evaluation function: during the evolution process, the prediction confidence of candidate polypeptide sequences is obtained through the structure prediction module, self-assembly behavior prediction module, and drug release prediction module; for any candidate sequence whose prediction confidence of any performance index is lower than a preset threshold, a confidence penalty term is applied when calculating its fitness, reducing its level in the non-dominated sorting or the probability of being selected, thereby guiding the optimization algorithm to prioritize exploring the sequence space with high prediction certainty and avoiding misjudging sequences that are "inaccurately predicted but have artificially high indices" as the optimal solution; For the aforementioned multiple performance objectives, a multi-objective fitness evaluation function vector is constructed, and the specific transformation rules include: For performance metrics that need to be maximized, such as self-assembly capability and drug loading rate, their negative values ​​are used as the minimization target; for performance metrics that need to achieve target values, such as target release rate, the absolute difference between them and the target value is used as the minimization target. For metrics that need to be minimized, such as sequence length, they can be directly used as the minimization target. Figure 5 This is a convergence curve of the improved non-dominated sorting genetic algorithm (NSGA-II) used in this system during the optimization process. The horizontal axis represents the number of evolutionary iterations (generations), and the vertical axis represents the evaluation index value (such as generation distance GD). As can be seen from the graph, the evaluation index value decreases rapidly with the increase of the number of iterations: it drops below the convergence threshold within the first 50 generations, then decreases slowly and eventually stabilizes. This curve shows that the improved NSGA-II, with its confidence penalty mechanism, can converge quickly within a limited number of iterations, effectively avoiding getting trapped in local optima and finding the global optimum, thus verifying the efficiency and stability of the optimization algorithm in this system.

[0131] Implementation Method 10: This implementation method further defines Implementation Method 1 and provides examples to illustrate the multiple performance objectives.

[0132] The multiple performance objectives include: maximizing self-assembly capability, maximizing drug loading rate, achieving target release rate, minimizing sequence length, and maximizing synthesis feasibility; During the optimization process, the confidence level of the prediction model is taken into account. For candidate polypeptide sequences whose prediction model output confidence level is lower than a preset threshold, a confidence penalty term is applied in the fitness evaluation to reduce their fitness value, thereby reducing the retention probability of sequences with high uncertainty in evolutionary selection and guiding the optimization algorithm to prioritize the exploration of high-confidence sequence space. Figure 6 This is a Pareto front and decision analysis diagram after multi-objective optimization. Each scatter point in the diagram represents a candidate polypeptide sequence. The horizontal axis represents the antitumor activity score, and the vertical axis represents the drug loading (% w / w). Both optimization objectives are to maximize the antitumor activity score. The diagram distinguishes three types of solutions using different colors: red dots represent Pareto front solutions (non-dominated solutions), gray dots represent near-front solutions, and blue dots represent dominated solutions (dominated solutions). The results show that there are 9 Pareto front solutions, with antitumor activity scores ranging from 0.15 to 0.78 and drug loading ranging from 2.5% to 21.0%. The boundary line formed by the red dots represents the Pareto front. The set of sequences on this front represents the optimal compromise solution under the current constraints (sequence length, stability), where it is impossible to improve one performance without reducing another. The optimal compromise solution is marked in the diagram, with an antitumor activity score of 0.52 and a drug loading of 14.2%. Users can select the final peptide sequence on this frontier surface based on the focus of the actual application (such as a greater emphasis on anti-tumor activity or higher drug loading).

[0133] Implementation method eleven is a further supplement to implementation methods one through ten, and details the overall training process of each deep learning model in the design system.

[0134] The Transformer encoding unit, the second-level structure prediction unit (deep convolutional neural network), the third-level structure prediction unit (graph neural network), the self-assembly behavior prediction module, and the drug release prediction module involved in the design system are trained and integrated using the following phased joint training strategy: Step 1: Data Preparation. Collect and organize the four datasets described in Implementation Method Sixteen from the PDB database and published literature: protein secondary structure dataset, protein tertiary structure dataset, peptide self-assembly behavior dataset, and drug release dataset. Perform data cleaning, standardization, and training / validation / test set partitioning (8:1:1 ratio) on all datasets.

[0135] Step 2: The first stage is joint pre-training. The Transformer encoding unit, deep convolutional neural network (secondary structure prediction head), and graph neural network (tertiary structure prediction head) are jointly pre-trained end-to-end on protein secondary structure and tertiary structure datasets for 200 epochs (including early stopping). The model checkpoint with the minimum loss on the validation set is saved as the pre-training weights.

[0136] Step 3: Second-stage self-assembly prediction training. Load the weights of the Transformer encoding units pre-trained in the first stage and freeze their parameters. Train the self-assembly behavior prediction module on the peptide self-assembly behavior dataset for 150 epochs (including early stopping). Save the model checkpoint with the minimum loss on the validation set.

[0137] Step 4: Third stage drug release prediction training. Load the weights of the Transformer encoding units pre-trained in the first stage and freeze their parameters. Train the drug release prediction module on the drug release dataset for 100 epochs (including early stopping). Save the model checkpoint with the minimum loss on the validation set.

[0138] Step 5: Model Integration. Integrate the training results from the above four stages: load the Transformer encoding unit weights saved in the first stage (as the underlying layer for shared feature extraction), the self-assembly behavior prediction module weights saved in the second stage, the drug release prediction module weights saved in the third stage, and the weights of the structure prediction modules (deep convolutional neural networks and graph neural networks) to form a complete model inference pipeline.

[0139] Step Six: Model Performance Validation. Evaluate the independent performance of each module on four test sets to ensure each module meets the preset performance thresholds (secondary structure Q3 ≥ 85%, tertiary structure TM-score ≥ 0.7, self-assembly prediction R² ≥ 0.8, drug release prediction R² ≥ 0.8). Simultaneously, perform end-to-end system testing: input a set of target functional attributes and constraints, run the complete multi-objective optimization process, and verify whether the output optimized peptide sequence set meets the multiple performance target requirements.

[0140] Step 7: Online model update. When new experimental data is accumulated, execute the online incremental update process: (1) Incremental data management: New data must have a sequence length of 8-20 amino acids and contain complete experimental characterization data. The maximum number of data updates per time is 1000. (2) Incremental learning strategy: The elastic weight consolidation (EWC) algorithm is used to protect historical knowledge. The training data consists of 100% new data and 10% historical data (hierarchical random sampling). All 6 layers of weights of the Transformer coding unit are fixed, and only the downstream self-assembly and drug release prediction head are fine-tuned. (3) Training parameters: The AdamW optimizer is used, with a learning rate of 5e-5, a batch size of 64, and 10 training rounds. The training is stopped early when the average MAE of the validation set does not decrease for two consecutive rounds. (4) Model evaluation and replacement: The new and old models are compared on a fixed test set. When the overall performance index of the new model is improved by ≥2%, the version is replaced and automatically archived.

[0141] Implementation Method Twelve: This implementation method illustrates the beneficial effects of this application through experimental data on the design of polypeptide nanofibers for antitumor drug delivery.

[0142] This embodiment uses the design system described in this invention to design a polypeptide nanofiber for loading and delivering doxorubicin (DOX), with the following target functional properties: (1) Drug loading rate ≥15%; (2) The release half-life is in the range of 24-48 hours; (3) It is stable in physiological saline for at least 72 hours.

[0143] In the design system described in this invention, the target functional attributes, constraints, and drug parameter configurations are input: hydrophobicity index 0.3-0.7, net charge +1 to +3, target drug loading rate ≥15%, and target release half-life 24-48h. The sequence length is set to 8-20 amino acids, and must include K (lysine) or R (arginine) to provide a positive charge. Sequence patterns such as PPP and GGGG that may affect the structure are prohibited.

[0144] A probabilistic sampling strategy was used to generate 1000 initial candidate sequences. The system calculates the amino acid selection probability based on the target attributes, and tends to select amino acids with moderate hydrophobicity (such as V, I, L) and positive charge (K, R).

[0145] For each candidate sequence, structural prediction is performed, extracting secondary structure content and predicting three-dimensional conformation. Subsequently, self-assembly prediction and drug release prediction modules are run to calculate the performance indicators of each sequence.

[0146] Multi-objective optimization is performed using the NSGA-II algorithm, with a population size of 100 and a maximum iteration count of 200 generations. The optimization objective function is designed as follows: F1 = -drug loading rate F2 = |Release Half-Life - 36| / 36 F3 = - Morphological stability score After optimization, the system outputs 50 non-dominated solutions on the Pareto front. Based on actual needs, the sequence with the best overall performance was selected, resulting in the optimal sequence for peptide nanofibers used in antitumor drug delivery: KLVFFAKLVFFAK (SEQ ID NO: 1). The predicted performance of this sequence is: drug loading rate of 18.5%, release half-life of 38 hours, and 72-hour stability score of 0.92.

[0147] Implementation Method Thirteen: This implementation method illustrates the beneficial effects of this application through experimental data on the design of polypeptide nanofibers for wound healing.

[0148] The design of a peptide nanofiber scaffold for wound healing requires good cell compatibility and a moderate degradation rate.

[0149] Using the deep learning-based peptide self-assembly nanofiber design system described in Embodiments 1-10, different target parameters were set to finally obtain the optimal sequence of peptide nanofibers for wound healing: RGDKLVFFAGG (SEQ ID NO:2), which contains the RGD cell adhesion motif.

[0150] like Figure 7 As shown in (a), the predicted three-dimensional spatial coordinates (X(A), Y(A), Z(A)) of this candidate sequence demonstrate a stable molecular configuration and a well-distributed predicted scoring motif. Figure 7 As shown in (b) of the figure, the system-predicted cell compatibility score is 0.95. Other reference values ​​(0.88 and 0.9) are also given in the figure, indicating that this sequence has excellent cell compatibility. Figure 7 As shown in (c), the detailed predicted parameters for this sequence are: 12 amino acids, a cell compatibility score of 0.95, a degradation half-life of approximately 7 days, a critical concentration for self-assembly of 0.15 mM, a predicted fiber diameter of 8.2 nm, and the functional motif being RGD cell adhesion. These results validate the effectiveness and reliability of this system in the design of wound healing peptide nanofiber scaffolds.< / unk> < / pad> < / cls>

Claims

1. A deep learning-based peptide self-assembly nanofiber design system, characterized in that, The design system includes the following modules: The input module is used to input the target functional properties, constraints, and drug parameter configurations of the peptide self-assembled nanofibers to be designed. The peptide sequence generation module is used to generate a set of candidate peptide sequences based on the target functional properties and constraints, using a sequence generation strategy. The structure prediction module is used to perform secondary structure prediction and tertiary structure prediction for each candidate polypeptide sequence in the candidate polypeptide sequence set, and is also used to extract geometric features related to self-assembly based on the prediction results to generate structural feature data. The self-assembly behavior prediction module is used to predict the self-assembly ability of peptides and nanofiber morphology parameters based on the structural feature data, and generate self-assembly performance indicators. The drug release prediction module is used to predict the drug loading rate and release kinetic parameters of the polypeptide self-assembled nanofibers under the drug parameter configuration based on the structural feature data, and generate drug release performance indicators. The multi-objective optimization module is used to iteratively optimize the candidate peptide sequence set using a multi-objective optimization algorithm with the optimization objectives of maximizing the self-assembly performance index and the drug release performance index, and output an optimized peptide sequence set that satisfies multiple performance objectives.

2. The deep learning-based peptide self-assembly nanofiber design system according to claim 1, characterized in that, The target functional properties include hydrophobicity, charge, and polarity; the constraints include sequence length range, forbidden mode detection, and essential amino acid verification.

3. The deep learning-based peptide self-assembly nanofiber design system according to claim 1, characterized in that, The sequence generation strategy includes probability sampling generation, beam search optimization, constraint verification, and functional motif generation.

4. The deep learning-based peptide self-assembly nanofiber design system according to claim 1, characterized in that, The structure prediction module also includes a Transformer encoding unit, which is used to perform global context encoding on the polypeptide sequence, extract high-dimensional feature representations, and use them as input features for secondary structure prediction and tertiary structure prediction, respectively.

5. The deep learning-based peptide self-assembly nanofiber design system according to claim 4, characterized in that, The secondary structure prediction employs a deep convolutional neural network, using the high-dimensional feature representation output by the Transformer encoding unit as input, and outputting the probability of the secondary structure type for each residue position; the secondary structure type includes polypeptides. -spiral, - Content of folds and random curls.

6. The deep learning-based peptide self-assembly nanofiber design system according to claim 4, characterized in that, The tertiary structure prediction employs a graph neural network, based on the high-dimensional feature representation output by the Transformer encoding unit, using amino acids as nodes and inter-residue interactions as edges to predict the three-dimensional coordinates of atoms.

7. The deep learning-based peptide self-assembly nanofiber design system according to claim 1, characterized in that, The self-assembly performance indicators include the critical concentration for peptide self-assembly, nanofiber diameter, fiber length distribution, and morphological stability.

8. The deep learning-based peptide self-assembly nanofiber design system according to claim 1, characterized in that, The multi-objective optimization algorithm is an improved non-dominated sorting genetic algorithm NSGA-II.

9. The deep learning-based peptide self-assembly nanofiber design system according to claim 1, characterized in that, The multiple performance objectives include: maximizing self-assembly capability, maximizing drug loading rate, achieving target release rate, minimizing sequence length, and maximizing synthesis feasibility.

10. A deep learning-based design method for peptide self-assembly nanofibers, wherein the design method is as follows: S1. Based on the target functional properties and constraints of the self-assembled peptide nanofibers to be designed, a sequence generation strategy is used to generate a set of candidate peptide sequences. S2, perform secondary and tertiary structure prediction on each candidate polypeptide sequence in the candidate polypeptide sequence set, and generate secondary structure prediction results and tertiary structure prediction results; S3. Based on the prediction results of the secondary structure and the tertiary structure, extract the geometric features related to self-assembly and generate structural feature data. S4. Based on the structural feature data, predict the self-assembly ability of the peptide and the morphological parameters of the nanofiber, and generate self-assembly performance indicators. S5. Based on the structural feature data, predict the drug loading rate and release kinetic parameters of the polypeptide self-assembled nanofibers under the drug parameter configuration, and generate drug release performance indicators. S6, with the optimization objectives of maximizing the self-assembly performance index and the drug release performance index, the candidate peptide sequence set is iteratively optimized using a multi-objective optimization algorithm to output an optimized peptide sequence set that satisfies multiple performance objectives.

Citation Information

Patent Citations

  • Nanometer anti-tumor drug design optimization method and system based on Transform model

    CN120089233A

  • Sensitizer intelligent screening method and system for clinical detection

    CN121922290A