6-HB targeted membrane fusion inhibitory peptide prediction method, device, equipment and medium

By employing a two-stage transfer learning framework and a multi-layer graph convolutional network model, and combining sequence general features and spatial conformation features, high-precision 6-HB-targeting membrane fusion inhibitory peptides are generated and screened. This addresses the shortcomings of existing models in recognizing conformational matching characteristics of peptide-target interactions, and enables the efficient development of virus subtype-specific inhibitory peptides.

CN120913657AActive Publication Date: 2025-11-07BEIJING YUEKANGKECHUANG PHARM TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511431984.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2025-11-07
Estimated Expiration
2045-10-09

Smart Images

  • Figure CN120913657A_ABST
    Figure CN120913657A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and discloses a 6-HB targeted membrane fusion inhibition peptide prediction method, device and equipment and a medium, and the method comprises the steps: generating a targeted 6-HB membrane fusion inhibition candidate peptide; inputting the membrane fusion inhibition candidate peptide into a pre-constructed 6-HB targeted membrane fusion inhibition peptide prediction model for classification prediction to obtain a classification result of the membrane fusion inhibition candidate peptide; the membrane fusion inhibitory peptide prediction model is a two-stage transfer learning classification model, and in the first stage, anti-virus peptide and non-anti-virus peptide dichotomy prediction is carried out based on sequence general characteristics; in the second stage, spatial conformation features are fused, and binary classification prediction of membrane fusion inhibition peptide and non-membrane fusion inhibition antiviral peptide is carried out; and screening a classification result to obtain the 6-HB targeted membrane fusion inhibition candidate peptide. The crossing from coarse-grained antiviral activity identification to fine-grained 6-HB inhibitory activity classification is realized, and the prediction accuracy of the enveloped virus type / subtype specific 6-HB inhibitory peptide is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of deep learning, in particular to a 6-HB targeting membrane fusion inhibitory peptide prediction method, device, equipment and medium. BACKGROUND

[0002] Enveloped viruses refer to viral types with external lipid membranes, such as influenza viruses, HIV, coronaviruses, etc. The infection mechanism of enveloped viruses is highly dependent on the membrane fusion process mediated by viral envelope proteins. The core structural basis of this process is the "six helix bundle" (Six-Helix Bundle, 6-HB) fusion core formed by the activated virus fusion protein. This structure is composed of a stable coiled helix structure formed by the reverse parallel combination of the HR1 (a heptad repeat sequence domain in the envelope protein) core and the HR2 (also a heptad repeat sequence domain in the envelope protein) peptide segment. This process releases energy and drives viral-cell membrane fusion. Fusion inhibitory peptides targeting 6-HB (such as enfuvirtide) competitively bind to the HR1 hydrophobic groove by simulating the HR2 domain, blocking the formation of natural 6-HB, and becoming an important strategy for antiviral drug design.

[0003] However, traditional antiviral peptide recognition models (such as sequence similarity-based, physicochemical feature-based, or shallow machine learning models) mainly rely on linear sequence information, while the inhibitory activity of 6-HB is highly dependent on the three-dimensional spatial fit of the peptide segment with the target. The current method cannot effectively represent the conformational matching characteristics (such as spatial topology, local curvature, dihedral angle, etc.) of peptide-target interactions, resulting in high false positive rates and inability to predict subtype-specific inhibitory activity. SUMMARY

[0004] Therefore, the present application provides a 6-HB targeting membrane fusion inhibitory peptide prediction method, device, equipment and medium to efficiently and accurately predict 6-HB targeting membrane fusion inhibitory candidate peptides.

[0005] In a first aspect, the present application provides a 6-HB targeting membrane fusion inhibitory peptide prediction method, which comprises: generating a 6-HB targeting membrane fusion inhibitory candidate peptide; inputting the membrane fusion inhibitory candidate peptide into a pre-constructed 6-HB targeting membrane fusion inhibitory peptide prediction model for classification prediction to obtain the classification result of the membrane fusion inhibitory candidate peptide; wherein the membrane fusion inhibitory peptide recognition model is a two-stage transfer learning classification model, the first stage is based on sequence general features for antiviral peptide and non-antiviral peptide binary classification prediction; the second stage integrates spatial conformation features for membrane fusion inhibitory peptide and non-membrane fusion inhibitory antiviral peptide binary classification prediction; screening the classification result to obtain the 6-HB targeting membrane fusion inhibitory candidate peptide.

[0006] The 6-HB targeting membrane fusion inhibition peptide prediction method provided by the application is an integrated computing framework that combines sequence general characteristics, spatial conformation characteristics, migration learning enhancement, and a generation-prediction-screening closed loop, realizes the leap from coarse-grained antiviral activity recognition to fine-grained 6-HB inhibition activity classification, and drives the directional generation and high-precision screening of functional peptide sequences, thereby accelerating the development of subtype-specific inhibition peptides targeting the fusion core of viruses and improving the prediction accuracy of envelope virus type / subtype-specific 6-HB inhibition peptides.

[0007] In an alternative embodiment, the 6-HB targeting membrane fusion inhibition peptide prediction model is established by the following steps: A first data set is obtained, which includes antiviral peptides and non-antiviral peptides, wherein the antiviral peptides are used as positive data sets and the non-antiviral peptides are used as negative data sets; A first initial multi-layer graph convolution network model is constructed; Based on the first data set, the first initial multi-layer graph convolution network model is trained to obtain a trained general antiviral peptide activity prediction model; A second data set is obtained, which includes 6-HB targeting membrane fusion inhibition peptides and non-6-HB targeting antiviral peptides, wherein the 6-HB targeting membrane fusion inhibition peptides are used as positive data sets and the non-6-HB targeting antiviral peptides are used as negative data sets; A second initial multi-layer graph convolution network model with the same architecture as the first initial multi-layer graph convolution network model is constructed; The general antiviral peptide activity prediction model weight is frozen as a sequence feature extractor, and the extracted sequence general characteristics are used for fusion to the node features of the second initial multi-layer graph convolution network architecture model; The general sequence characteristics extracted by the general antiviral peptide activity prediction model and the precomputed spatial conformation characteristics are spliced based on the second data set, and the second initial multi-layer graph convolution network model is trained to obtain a 6-HB targeting membrane fusion inhibition peptide prediction model.

[0008] In this embodiment, based on the feature transfer strategy, the first-stage general anti-virus peptide activity prediction model is trained first, the weights of the first-stage model are frozen as a feature extractor, the extracted general sequence features and spatial conformation features are spliced to form enhanced feature embedding, and then the 6-HB targeting membrane fusion inhibition peptide prediction model is trained again, so that the second-stage model can more accurately identify the key structural information of the 6-HB targeting membrane fusion inhibition peptide, effectively utilize the general features of the anti-virus peptide, and combine the spatial conformation features of the target peptide to effectively improve the prediction accuracy and efficiency of the 6-HB targeting membrane fusion inhibition peptide. Moreover, through feature transfer and model freezing, the number of parameters that need to be trained is reduced, overfitting is effectively inhibited, the generalization performance of the model is significantly improved, and the training time is shortened, which is suitable for large-scale anti-virus peptide screening and evaluation.

[0009] In an alternative embodiment, based on the first data set, a first initial multi-layer graph convolution network model is trained, comprising: data preprocessing is performed on the first data set; Based on the pre-processed first data set, sequence general features are calculated, including amino acid type encoding features, physicochemical property features and evolutionary semantic features; Based on the pre-processed first data set, a binary adjacency matrix is constructed using a k-neighbor method based on sequence distance; The sequence general features are used as node features, and the reverse sequence distance is used as edge weight to construct a first graph structure; The first graph structure is input into the first initial multi-layer graph convolution network model, and when the preset training indicators are met, a general anti-virus peptide activity prediction model is obtained.

[0010] In this embodiment, the general anti-virus peptide activity prediction model is constructed by combining data preprocessing, sequence feature extraction and graph convolution network technology. The obtained general anti-virus peptide activity prediction model can effectively capture global and local features from the peptide sequence graph. After freezing all learnable parameters during subsequent learning transfer, stable node-level feature embedding representation can be output, thereby effectively enhancing the expression ability of the downstream 6-HB targeting membrane fusion inhibition peptide prediction model for complex relationships in the sequence and improving the prediction accuracy.

[0011] In an alternative embodiment, based on the second data set, the general sequence features extracted by the general anti-virus peptide activity prediction model are spliced with the pre-calculated spatial conformation features to train the second initial multi-layer graph convolution network model, comprising: data preprocessing is performed on the second data set; Based on the pre-processed second data set, spatial conformation features are extracted, including: rotationally and translationally invariant coordinates, Euclidean distance, Normalized inverse Euclidean distance, Discrete curvature, Pseudo dihedral angle, conformation-related AAindx feature, residue solvent accessible surface area-related feature; Based on Euclidean distance, generate a binary adjacency matrix; Concatenate the general sequence features along the feature dimension, input into the general antiviral peptide activity prediction model with frozen weights, to obtain a high-dimensional sequence feature representation; Concatenate the high-dimensional sequence feature representation and the spatial conformation feature along the feature dimension, together as enhanced node features; and construct a second graph structure with the normalized inverse Euclidean distance as the edge weight; Input the second initial multi-layer graph convolution network model after feature migration into the second graph structure, and obtain the 6-HB targeting membrane fusion inhibitory peptide prediction model when the preset training indicators are met.

[0012] In this embodiment, by combining the antiviral peptide general sequence features extracted by the first stage pre-training model and the pre-calculated spatial conformation features, the interaction between the candidate peptide and the target virus 6-HB can be more comprehensively captured from the sequence and conformation levels, thereby effectively improving the recognition ability of the second stage model for 6-HB targeting membrane fusion inhibitory peptides.

[0013] In an alternative embodiment, generating a 6-HB targeting membrane fusion inhibitory candidate peptide comprises: Obtaining antiviral peptide data; Based on the antiviral peptide data, pre-training the pre-constructed model architecture to obtain a general antiviral peptide generation model; Obtaining 6-HB targeting membrane fusion inhibitory peptide data; Based on the 6-HB targeting membrane fusion inhibitory peptide data, fine-tuning the general antiviral peptide generation model to obtain a 6-HB targeting membrane fusion inhibitory peptide generation model; the 6-HB targeting membrane fusion inhibitory peptide generation model is used to generate a membrane fusion inhibitory candidate peptide.

[0014] In this embodiment, a two-stage transfer learning framework is adopted, the first stage pre-trains the model on the antiviral peptide dataset to obtain the general antiviral peptide sequence generation capability; the second stage fine-tunes all parameters on the membrane fusion inhibitory peptide dataset to train and obtain the 6-HB targeting membrane fusion inhibitory peptide generation model. This can effectively improve the generalization ability of the final 6-HB targeting membrane fusion inhibitory peptide generation model, so that it can not only retain the general sequence generation capability when generating 6-HB targeting membrane fusion inhibitory peptide sequences, but also accurately capture key features related to membrane fusion.

[0015] In an alternative embodiment, the 6-HB targeting membrane fusion inhibition peptide generation model comprises: a variable-length sequence preprocessing module for preprocessing the input sequence; preprocessing the input sequence comprises: sorting the input sequence in descending order of actual length; padding the sequence sorted in descending order to a uniform length at its end, and generating a corresponding padding mask to mark the valid data position; an embedding layer connected to the output end of the variable-length sequence preprocessing module, for mapping the preprocessed integer index sequence to a dense vector representation in a high-dimensional space; an LSTM layer connected to the output end of the embedding layer, for capturing long-range dependencies and context information in the preprocessed vector sequence; a fully connected layer, the output of the LSTM layer is finally mapped to a probability distribution in the output space by the fully connected layer; The 6-HB targeting membrane fusion inhibition peptide generation model adopts a dynamic training control strategy in the training phase, which combines gradient norm constraint and continuous validation loss monitoring; The 6-HB targeting membrane fusion inhibition peptide generation model is equipped with a generated sequence filtering unit in the sequence generation phase; the generated sequence filtering unit performs the following three operations: adjusting the output probability distribution of the Softmax function using the temperature parameter to control the generation diversity; dynamically filtering non-amino acid tokens to ensure that each step of generation only samples in the valid amino acid dictionary; responding to the end symbol, the generation process is immediately interrupted once the special symbol representing the end of the sequence is generated.

[0016] The 6-HB targeting membrane fusion inhibition peptide generation model constructed in this embodiment can process variable-length sequence input, mark valid data positions with padding masks, and reduce invalid calculations; the dynamic training control module improves the generalization performance, training stability and convergence speed of the model through gradient norm constraint and validation loss monitoring; the generated sequence filtering unit adjusts the output probability distribution of the Softmax function through the temperature parameter to control the diversity of the generated sequence; and can effectively filter non-amino acid tokens, respond to the end symbol to interrupt, and ensure the reliability of the generated sequence. The embedding layer maps the preprocessed integer index sequence to a dense vector representation in a high-dimensional space. The LSTM layer captures long-range dependencies and context information in the sequence, and maps it to a probability distribution in the output space through the fully connected layer, thereby generating membrane fusion inhibition peptide sequences with targeting and functional potential.

[0017] In an alternative embodiment, the screening classification result is obtained to obtain the final 6-HB targeting membrane fusion inhibition candidate peptide, comprising: determining the initial candidate peptide in the classification result, the initial candidate peptide being a candidate peptide classified as a 6-HB targeting membrane fusion inhibition peptide; performing test molecular docking on each initial candidate peptide and the target virus 6-HB conformation to obtain a docking score corresponding to each initial candidate peptide; determining whether the docking score meets a preset condition; determining the candidate peptide meeting the preset docking score condition as a final 6-HB targeting membrane fusion inhibition candidate peptide.

[0018] In this embodiment, by performing test docking on each initial candidate peptide and further screening according to the docking score, the success rate of downstream biological verification experiments can be effectively improved.

[0019] In a second aspect, the present application provides a 6-HB targeting membrane fusion inhibition peptide prediction device, which comprises: a sequence generation module for generating a 6-HB targeting membrane fusion inhibition candidate peptide; a prediction module for inputting the membrane fusion inhibition candidate peptide into a pre-constructed 6-HB targeting membrane fusion inhibition peptide prediction model for classification prediction to obtain a classification result of the membrane fusion inhibition candidate peptide; wherein the membrane fusion inhibition peptide recognition model is a two-stage transfer learning classification model, the first stage is based on sequence general features for anti-virus peptide and non-anti-virus peptide binary classification prediction; and the second stage integrates spatial conformation features for membrane fusion inhibition peptide and non-membrane fusion inhibition anti-virus peptide binary classification prediction; a screening module for screening the classification result to obtain a final 6-HB targeting membrane fusion inhibition candidate peptide.

[0020] In a third aspect, the present application provides a computer device, which comprises a memory and a processor, the memory and the processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the 6-HB targeting membrane fusion inhibition peptide prediction method of the first aspect or any of the corresponding embodiments thereof.

[0021] In a fourth aspect, the present application provides a computer readable storage medium, which stores computer instructions, and the computer instructions are used to make the computer execute the 6-HB targeting membrane fusion inhibition peptide prediction method of the first aspect or any of the corresponding embodiments thereof.

[0022] It should be noted that the 6-HB targeting membrane fusion inhibition peptide prediction device, the computer device, and the computer readable storage medium provided by the present application correspond to the 6-HB targeting membrane fusion inhibition peptide prediction method described above. Therefore, the beneficial effects of the 6-HB targeting membrane fusion inhibition peptide prediction device, the computer device, and the computer readable storage medium are described in the corresponding beneficial effects of the 6-HB targeting membrane fusion inhibition peptide prediction method, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the specific embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0024] Figure 1 is a flowchart of the prediction and screening method of 6-HB targeting membrane fusion inhibition peptide according to the embodiment of the present application; Figure 2 is a transfer learning schematic diagram of the 6-HB targeting membrane fusion inhibition peptide generation model and prediction model according to the embodiment of the present application; Figure 3 is a performance comparison schematic diagram of the general antiviral peptide prediction model and the most advanced model of the same task according to the embodiment of the present application; Figure 4 is a 6-HB targeting membrane fusion inhibition peptide prediction model transfer learning ablation experiment result graph according to the embodiment of the present application; Figure 5 is a 6-HB targeting membrane fusion inhibition peptide generation model transfer learning ablation experiment result graph according to the embodiment of the present application; Figure 6 is a 6-HB targeting membrane fusion inhibition peptide generation model architecture schematic diagram according to the embodiment of the present application; Figure 7 is a sequence preprocessing flowchart according to the embodiment of the present application; Figure 8 is a structural block diagram of the 6-HB targeting membrane fusion inhibition peptide prediction device according to the embodiment of the present application; Figure 9 is a hardware structure schematic diagram of the computer device of the embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor also belong to the scope of protection of the present application.

[0026] Currently, the traditional anti-viral peptide research and development technology is facing three core bottlenecks: 1. Ignorance of spatial conformation dependence: The traditional method is difficult to effectively characterize the conformation matching characteristics of the peptide-target interaction, resulting in a high false positive rate and the inability to predict subtype-specific inhibitory activity. 2. Insufficient model generalization due to sample scarcity: The number of experimentally verified membrane fusion inhibitory peptides is extremely limited (especially for specific virus subtypes). Under the condition of small sample, the conventional machine learning classifier is prone to overfitting, and it is difficult to extract discriminative features from sparse data. In addition, the conformational diversity of 6-HB across virus genera further increases the modeling difficulty, and the classification accuracy of the general anti-viral peptide prediction model for membrane fusion inhibition function is generally lower than the acceptable threshold. 3. Loss of function of targeted generation: Existing peptide sequence generation models (such as RNN, VAE or GPT architecture) focus on sequence probability distribution learning, and lack of physical and chemical constraints on target biological functions (such as 6-HB binding energy, steric hindrance effect). The generated results often have high sequence novelty but lack binding activity, and cannot efficiently output functional fusion inhibitory peptides with virus subtype specificity.

[0027] Therefore, according to the embodiments of the present application, a 6-HB targeting membrane fusion inhibitory peptide prediction method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown.

[0028] In this embodiment, a 6-HB targeting membrane fusion inhibitory peptide prediction method is provided, that is, a six-helix bundle (6-HB) targeting membrane fusion inhibitory peptide prediction method, which can be executed by a server, a terminal, a mobile terminal and the like, Figure 1 The flowchart of the 6-HB targeting membrane fusion inhibitory peptide prediction method according to the embodiments of the present application is shown in FIG. 1, which includes the following steps: Figure 1 Step S101, generate 6-HB targeting membrane fusion inhibitory candidate peptides.

[0029] ​6-HB (six-helix bundle) targeting membrane fusion inhibitory peptides are a class of peptide molecules that can specifically interfere with the membrane fusion process mediated by viral envelope proteins (such as HIV gp41 and SARS-CoV-2 S protein). Their mechanism of action lies in targeting the key transitional structure formed by viral fusion proteins during membrane fusion—the six-helix bundle (6-HB). These inhibitory peptides competitively bind to pre-fusion intermediates (such as the HR1 trimer) by mimicking functional regions of viral proteins (such as HR2), thereby preventing endogenous HR1-HR2 interactions, effectively inhibiting the proper formation of 6-HB, and ultimately halting the fusion process between the virus and the host cell membrane, thus inhibiting viral infection.

[0030] In this embodiment, the 6-HB-targeting membrane fusion inhibitory candidate peptide was generated using a pre-constructed 6-HB-targeting membrane fusion inhibitory peptide generation model. The 6-HB-targeting membrane fusion inhibitory peptide generation model was obtained by training an LSTM autoregressive model using a two-stage transfer learning strategy, which will be described in detail below. Of course, the 6-HB-targeting membrane fusion inhibitory candidate peptide in this embodiment can also be extracted from known protein sequences or from sequence libraries; no particular limitation is made here.

[0031] Step S102: Input the membrane fusion inhibition candidate peptide into the pre-constructed 6-HB-targeted membrane fusion inhibition peptide prediction model for classification and prediction to obtain the classification results of the membrane fusion inhibition candidate peptide; wherein, the membrane fusion inhibition peptide identification model is a two-stage transfer learning classification model, which is trained based on sequence general features and spatial conformation features. The first stage is based on sequence general features to perform binary classification prediction of antiviral peptides and non-antiviral peptides; the second stage incorporates spatial conformation features to perform binary classification prediction of membrane fusion inhibition peptides and non-membrane fusion inhibition antiviral peptides.

[0032] In this embodiment, the pre-constructed 6-HB-targeting membrane fusion inhibitory peptide prediction model is a classification model built based on two-stage transfer learning. Specifically, refer to... Figure 2 As shown, the first stage constructs a graph neural network (GNN) based on a k-contact graph of antiviral peptide sequences. This GNN uses residue type encoding extracted through static feature engineering (which can be one-hot, BLOSUM 62, etc., without specific limitations), ESM2 (Evolutionary Scale Modeling 2) feature embeddings, and physicochemically related amino acid indices as node features, and normalized inverse sequence distance as edge weights. The prediction task of the first stage is binary classification of antiviral peptides and non-antiviral peptides. The second stage constructs a graph neural network based on the k-contact graph of antiviral peptide sequences. carbon atom ( GNN of binary contact map of Euclidean distance, wherein the static feature engineering extraction Spatial coordinates, pseudo-dihedral angle, discrete curvature, etc. Spatial geometric features are used as node features, and general sequence feature representations extracted by the first stage pre-training model of parameter freezing are spliced to form enhanced node embedding vectors. The normalized inverse spatial distance is used as the edge weight. The prediction task of the second stage is to classify 6-HB targeting membrane fusion inhibition peptides and non-6-HB targeting antiviral peptides.

[0033] In actual use, the determined 6-HB targeting membrane fusion inhibition candidate peptide is directly input into the 6-HB targeting membrane fusion inhibition peptide prediction model, and it can be predicted whether the candidate peptide is a 6-HB targeting membrane fusion inhibition peptide.

[0034] Step S103, screening the classification result to obtain the final 6-HB targeting membrane fusion inhibition candidate peptide.

[0035] For example, the candidate peptide with a prediction label of 1 (positive) can be selected for protein-peptide flexible docking with the pre-fusion conformation of the influenza A H3N2 subtype 6-HB to obtain a docking score . At least several experimentally verified 6-HB targeting membrane fusion inhibition peptides of influenza A H3N2 are collected from DRAVP (antiviral peptide database) as positive controls, and are docked with the target virus 6-HB under the same method and parameter setting to obtain a score ; then, according to the candidate peptide docking score and the positive control peptide docking score , the statistical quantity comparison relationship is determined, and the candidate peptide with docking affinity better than the statistical standard of the positive control peptide is selected as the high-specificity inhibition peptide candidate of the target virus 6-HB.

[0036] The 6-HB targeting membrane fusion inhibition peptide prediction method provided in this embodiment is an integrated computing framework that combines general sequence features, spatial conformation features, transfer learning enhancement, and a generation-prediction-screening closed loop, realizes the leap from coarse-grained antiviral activity recognition to fine-grained 6-HB targeting prediction, and drives the directional generation and high-precision screening of functional peptide sequences, thereby accelerating the development of subtype-specific inhibition peptides targeting the fusion core of the target virus and improving the prediction accuracy of 6-HB inhibition peptides of envelope virus types / subtypes.

[0037] In some optional embodiments, the 6-HB targeting membrane fusion inhibition peptide prediction model is established by the following steps: Step a1, obtaining a first data set, the first data set comprising antiviral peptides and non-antiviral peptides, wherein the antiviral peptides are used as positive data sets and the non-antiviral peptides are used as negative data sets.

[0038] The following is illustrated by example. The positive data set, i.e. the antiviral peptide, is labeled as 1, and 2762 experimentally verified antiviral peptide sequences can be collected from public databases dbAMP, AVPdb, DRAMP, DBAASP and HIPdb.

[0039] The negative data set, i.e. the non-antiviral peptide, is labeled as 0, and a total of 10089 non-antiviral peptide sequences are screened out, and the sources and standards can be as follows: Screening of antimicrobial but non-antiviral peptides, specifically, peptide sequences with clear antibacterial function but confirmed by literature to have no antiviral activity can be screened from the dbAMP, DBAASP and DRAMP databases. At the same time, during the screening process, all sequences targeting viruses or having antiviral activity are explicitly excluded; Screening of non-antimicrobial peptides, non-antimicrobial peptide sequences can be screened from the UniProt database. Specifically, it can be obtained by filtering by attribute keywords. The exclusion keywords used can include: "toxicity (toxin)", "membrane (membrane)", "secretory (secretory)", "defensive (defensive)", "antibiotic (antibiotic)", "anticancer (anticancer)", "antiviral (antiviral)", "antifungal (antifungal)". That is, only peptide sequences that do not contain any of the above keywords and are not related to antimicrobial function are retained.

[0040] Step a2, constructing a first initial multi-layer graph convolutional network architecture model. The first initial multi-layer graph convolutional network model in this embodiment can adopt a multi-layer graph convolutional network with residual connection.

[0041] Step a3, training the first initial multi-layer graph convolutional network model based on the first data set to obtain a trained general antiviral peptide activity prediction model.

[0042] In some optional embodiments, the above step a3 comprises: Step a31, data preprocessing of the first data set. Specifically, it includes: Removing non-amino acid markers: deleting unknown amino acid markers (such as X) and non-amino acid markers in all peptide sequences; Sequence length filtering: filtering sequences with a length less than 5 or greater than 100 (the filtering length can not be limited to this, which is not specifically limited here); Removing duplicate samples: respectively checking the sequence repeatability in the positive sample data set and the negative sample data set, and deleting the completely identical duplicate sequences to ensure the uniqueness of the samples in the data set; Detecting intersection of positive and negative sample sets: Compare the sequences of the positive sample set and the negative sample set, identify and delete the intersection samples existing in both sets, to ensure the independence of the positive and negative samples of the training data; Removing redundant sequences: To prevent training data leakage from distorting model performance evaluation, sequence redundancy removal is performed within the positive sample set and the negative sample set respectively; among which, the sequence homology analysis tool uses CD-HIT (sequence redundancy removal method is not specifically limited), and the sequence similarity threshold can be set to 80% to 95%; the core parameters are set as follows: the similarity threshold "-c" is valued within the range of [0.6, 0.99], preferably 0.95; the alignment region constraint "-s" is valued within the range of [0.7, 0.9], preferably 0.8; the minimum coverage of representative sequence "-aS" is valued within the range of [0.85, 0.95], preferably 0.9; the minimum coverage of current sequence "-aL" is valued within the range of [0.85, 0.95], preferably 0.9; the length filter "-d" is set to 0; the short sequence special alignment mode "-n" is set to 2. Only one representative sequence is retained in each cluster. The specific definition of the above parameters is not strictly limited here.

[0043] The positive sample set and the negative sample set after cleaning and redundancy removal are merged to form the final complete training data set. The complete data set is stratified sampled according to the label to divide into a training set, a validation set and a test set in the ratio of 6:2:2. The positive and negative sample sets obtained after the first stage of data preparation and preprocessing are used to train the general anti-virus peptide prediction model.

[0044] Step a32, based on the first preprocessed data set, calculate the sequence general features, including amino acid type encoding, physicochemical property features and evolutionary semantic features.

[0045] In this embodiment, a static pre-computation architecture is adopted, that is, an offline pre-computation mode is used to calculate the above-mentioned sequence general features; in this embodiment, a parallel acceleration mechanism is also adopted, that is, task-level parallelism is achieved through a multi-process pool, and computing resources are dynamically allocated. When training the first stage prediction model, only the sequence-related features need to be calculated.

[0046] The sequence general features are described below.

[0047] Regarding amino acid type encoding : a one-to-one mapping from 20 standard amino acids to a 20-dimensional discrete vector space is established; the output feature matrix is , where L is the length of the amino acid sequence, and , (onehot encoding is taken as an example in this embodiment, other encoding methods such as BLOSUM 62 can also be used, without specific limitation).

[0048] Physicochemical property features Calculation: Calculate the AAIndex indicators related to the activity of the antiviral peptide using the predefined AAIndex indicator system, including KYTJ820101 (hydrophobicity), FAUJ880111 (positive charge residues), GRAR740102 (polarity), FAUJ880112 (negative charge residues), ROBB760101 (solvent accessibility), FASG890101 (isoelectric point pi), and other physical and chemical indicators related to the activity of the antiviral peptide; output feature matrix , wherein, is the number of selected AAIndex indicators (in the current embodiment ).

[0049] ESM evolutionary semantic features Calculation: Sequence tokenization: Convert the amino acid sequence into a token index tensor through a predefined vocabulary; Deep representation extraction: Extract sequence representation using the middle layer (preferably the 33rd layer) or the final layer of the ESM model (esm2_t33_650M_UR50D in this embodiment) in the disabled gradient calculation mode; Residue representation pruning: Remove the feature vectors of task-irrelevant markers such as [CLS] and [SEP] at the beginning and end of the sequence; Dimension constraint: Output feature matrix The dimension is i.e. wherein, L is the length of the amino acid of the antiviral peptide, is the hidden layer dimension of the ESM model (1280 in this embodiment).

[0050] Step a33, based on the preprocessed first data set, a k-neighbor method based on sequence distance is used to construct a binary adjacency matrix. In this embodiment, the k-neighbor method is based on the sequence distance between the first and second data sets. is an integer; the binary adjacency matrix does not contain self-loop edges.

[0051] Step a34, the above-mentioned normalized and spliced sequence general characteristics are used as node features, and the reverse sequence distance is used as edge weight to construct the first graph structure.

[0052] Specifically, after constructing the binary adjacency matrix based on the sequence distance-based k-neighbor method, binary conversion is applied: ; wherein, , is the total number of standard amino acids in the peptide chain, indicating the sequence distance between residue i and residue j, is the number of nearest neighbor residues, which is 20 in this embodiment ; concatenate the amino acid type encoding matrix, the physicochemical property feature matrix and the ESM evolutionary semantic feature matrix along the feature dimension as the node feature; and normalize the reverse sequence distance as the edge weight: ; wherein, is the sequence distance between residue i and residue j; is a small constant 1e-7.

[0053] The k-neighbor graph is constructed based on the sequence distance, and the constructed graph structure data is input into the antiviral peptide activity prediction model for training. The training target is the binary classification of antiviral peptides and non-antiviral peptides.

[0054] In step a35, the first graph structure is input into the first initial multi-layer graph convolution network model, and when the preset training index is met, a general antiviral peptide activity prediction model is obtained.

[0055] The first initial multi-layer graph convolution network architecture in this embodiment includes two graph convolution layers, each of which performs a standardized graph convolution operation: ; wherein, is the adjacency matrix with self-loop added; is the degree matrix of , satisfying ; is the trainable weight matrix of the lth layer; the graph convolution layer hidden dimension is preferably 64; and the inter-layer feature reuse mechanism is: ; wherein, is the input feature matrix of the lth layer, is the trainable weight matrix of the lth layer, , and represents a graph convolution operation with an activation function. In the model training phase, the input features are first standardized and concatenated. Specifically, the mean and standard deviation of the physicochemical property feature matrix on the training set are calculated, and the z-score standardization is performed on the training set, the validation set and the test set. The evolutionary semantic feature matrix is sequentially subjected to LayerNorm standardization and linear transformation dimension reduction processing; the amino acid type encoding matrix, the standardized physicochemical property feature matrix and the dimension-reduced evolutionary semantic feature matrix are concatenated along the feature dimension to generate the node feature matrix; and the output dimension of the linear transformation dimension reduction is 16 to 128. In this embodiment, the linear projection output dimension is 32. In this embodiment, the feature standardization and concatenation are realized in the model training phase.

[0056]

[0057] ​In the model training stage, two-level output is adopted, first the node-level features are aggregated into a graph-level feature vector through global average pooling operation, and then a fully connected layer is used to map the graph-level features to binary classification logits output; in the feature migration application scenario, all trainable parameters of the model are frozen, and the node-level feature vector output by the last layer of graph convolution is used as the sequence feature embedding representation.

[0058] The model training strategy is as follows: Optimizer configuration: optimizer type AdamW, initial learning rate 1e-3; Early stopping strategy: monitor the validation loss, and trigger early stopping when the validation loss does not improve for 20 consecutive epochs (epochs); Performance detection index: monitor the comprehensive performance index of the validation set during training, including ROC-AUC, accuracy (Accuracy), recall (Recall), F1 score, Matthew correlation coefficient (MCC), specificity (Specificity), and area under the PR curve (AUCPR).

[0059] To alleviate the impact of class imbalance on model performance, a phased processing strategy is adopted in this embodiment: in the model training stage, when the sample class imbalance in the training set (in this embodiment, the ratio of the number of negative samples to the number of positive samples in the first stage ), a class weight is applied to the minority class (positive samples) in the loss function (this weight is only used for training gradient calculation and does not change the inference logic) to strengthen the model's learning of the minority class (positive samples). After the model training is completed, first, the temperature scaling is applied to the model output probability on the validation set to improve its confidence accuracy. Second, based on the calibrated probability, the optimal decision threshold that maximizes the F1 score (or other specified indicators) is searched on the validation set. Finally, the trained model parameters, the calibrated temperature parameter T, and the optimal threshold are saved.

[0060] Specifically, temperature scaling and probability threshold calibration are implemented as follows: the original logits value z is scaled by a trainable temperature parameter T: ; The scaled logits value is input into the Sigmoid function to calculate the calibrated probability: ; The trainable temperature parameter T is optimized by gradient descent method (L-BFGS optimizer), and the optimization goal is to minimize the cross-entropy loss of the validation set. Then the decision threshold search is performed: based on the calibrated probability, the decision threshold that maximizes the F1 score is searched on the validation set and record the calibration performance indicators. This process is only for performance monitoring, all weight parameters are frozen, only temperature parameters T are optimized. Save the trained model parameters and the temperature parameters T and optimal threshold values obtained by calibration In the inference phase, calculate for the test sample: ; outputs the positive class label if and only if .

[0061] where is to evaluate the classification performance under different thresholds to determine the optimal decision threshold The recorded performance indicators include at least one of the validation set F1 score (which is the harmonic mean of precision and recall), accuracy, recall, Matthew correlation coefficient (MCC), specificity, etc.

[0062] In this embodiment, the general anti-virus peptide activity prediction model is constructed by integrating data preprocessing, sequence feature extraction and graph convolution network technology. The general anti-virus peptide activity prediction model obtained can effectively capture the global and local features in the graph convolution network. When learning migration is performed subsequently, after freezing all learnable parameters, stable node-level feature embedding representation can be output, thereby effectively enhancing the expression ability of the downstream 6-HB targeting membrane fusion inhibiting peptide prediction model for complex relationships in the sequence and improving the prediction accuracy.

[0063] In this embodiment, the performance of the first-stage general anti-virus peptide prediction model is compared with that of the state-of-the-art model (SOTA) as follows Figure 3The performance of the model of the present embodiment is compared with that of two comparative models. The comparative models include: a model developed by Jiahui Guan et al. (Briefings in Bioinformatics, 2024, DOI: 10.1093 / bib / bbae208), i.e., the first comparative model, and a model developed by Ruifen Cao et al. (Briefings in Bioinformatics, 2023, DOI: 10.1093 / bib / bbad353.), i.e., the second comparative model. To ensure fairness, all comparative models use the code library disclosed by the original author, strictly follow the official recommended hyperparameters reported in the literature, and are retrained on the same training set as the model of the present embodiment and evaluated on the same reserved test set. The performance indicators include: AUC (Area Under the Curve, referring to the area under the Receiver Operating Characteristic Curve (ROC curve)), F1 score (which is the harmonic mean of precision and recall), accuracy (Accuracy), recall (Recall), Matthews Correlation Coefficient (MCC), specificity (Specificity), AUCPR (Area Under the Precision-Recall Curve, referring to the area under the Precision-Recall Curve).

[0064] Step a4, obtaining a second data set, the second data set comprising 6-HB targeting membrane fusion inhibition peptides and non-6-HB targeting antiviral peptides, wherein the 6-HB targeting membrane fusion inhibition peptides are taken as positive data sets, and the non-6-HB targeting antiviral peptides are taken as negative data sets.

[0065] The following is illustrated by example. The positive data set, i.e., the 6-HB targeting membrane fusion inhibition peptide, is labeled as 1, and can be manually screened based on the antiviral peptide target, mechanism or source annotation of the dbAMP and AVPdb databases, and manually screened based on the antiviral peptide target, mechanism or source annotation of the dbAMP and AVPdb databases. The positive sample 540 is finally obtained.

[0066] The negative data set, i.e., the non-6-HB targeting antiviral peptide, is labeled as 0, and can be manually screened based on the same database, and manually screened based on the antiviral peptide target, mechanism or source annotation of the dbAMP and AVPdb databases. The negative sample 490 is finally obtained.

[0067] Step a5, constructing a second initial multi-layer graph convolutional network model with the same architecture as the first initial multi-layer graph convolutional network model. The second initial multi-layer graph convolutional network model in the present embodiment still adopts a multi-layer graph convolutional network with residual connection.

[0068] Step a6, the general anti-viral peptide activity prediction model weight is frozen as a sequence general feature extractor, and the extracted sequence general features are fused to the node features of the second initial multi-layer graph convolution network model. That is, all trainable parameters of the general anti-viral peptide activity prediction model are frozen as a fixed feature extractor, and the extracted sequence features are input into the second initial multi-layer graph convolution network.

[0069] Specifically, the pre-trained general anti-viral peptide activity prediction model can be loaded and all trainable parameters are frozen to extract a node-level sequence embedding feature matrix. Then, spatial conformation feature extraction is performed. In this embodiment, the spatial conformation feature extraction system can be used to calculate conformation-related AAindex features, Rotation and translation invariant coordinates, Euclidean distance, Normalized inverse Euclidean distance, Discrete curvature, Pseudo dihedral angle, and solvent accessible surface area (SASA) related features. Further, graph structure construction is performed based on Euclidean distance matrix, a fixed distance threshold is set to generate a binary adjacency matrix. The node-level sequence representation extracted by the pre-trained model is concatenated with the normalized conformation-related AAindex features, Pseudo dihedral angle, and normalized Discrete curvature, Rotation and translation invariant coordinates, and normalized SASA related features along the feature dimension are concatenated to construct a joint node feature matrix; and Normalized inverse Euclidean distance is used as the edge weight to construct A spatial proximity graph.

[0070] Step a7, based on the second data set, the general sequence features extracted by the general anti-viral peptide activity prediction model are concatenated with the pre-computed spatial conformation features, and the second initial multi-layer graph convolution network model is trained to obtain a 6-HB targeting membrane fusion inhibitory peptide prediction model.

[0071] Specifically, the constructed graph structure data is input into the second initial multi-layer graph convolution network architecture to train a 6-HB targeting membrane fusion inhibitory peptide prediction model. The training target is to classify 6-HB targeting membrane fusion inhibitory peptides and non-6-HB targeting anti-viral peptides.

[0072] In some optional embodiments, the above step a7 includes: Step a71, data preprocessing is performed on the second data set; The data preprocessing of the second data set is completely consistent with the first stage data preparation, including: removing non-amino acid markers, sequence length filtering, removing duplicate samples by label set, detecting positive and negative sample set intersection, removing redundant sequences by label set, etc. The second data set after data preprocessing is stratified random sampling according to the label, and in this embodiment, it is divided into training set, validation set and test set in the ratio of 6:2:2. The positive and negative sample sets obtained after the second stage data preparation and preprocessing are used to train the 6-HB targeting membrane fusion inhibitory peptide prediction model.

[0073] Step a72, based on the preprocessed second data set, extract sequence features according to the foregoing method, and extract space conformation features at the same time, which include: rotationally and translationally invariant coordinates, Euclidean distance, normalized inverse Euclidean distance, discrete curvature, pseudo-dihedral angle, conformation-related AAindex feature, and residue solvent accessible surface area-related feature.

[0074] In this embodiment, before extracting the three-dimensional conformation features of the sequence, the input peptide sequence is also subjected to structure modeling to obtain a pdb file containing structure information. Further, the structure is repaired, that is, the PDBFixer can be used to complete the main chain and side chain of the input peptide, add hydrogen atoms in the physiological state (pH=7.4), and map non-standard residues to the most similar standard residues, that is, to complete the missing atoms and optimize the protonation state of the three-dimensional structure of the input peptide.

[0075] In this embodiment, a static pre-computation architecture is adopted, that is, an offline pre-computation mode is adopted, and sequence general features, that is, amino acid type encoding, physicochemical property features, ESM evolutionary semantic features, and space conformation features, that is, conformation-related AAindex features, rotationally and translationally invariant coordinates, Euclidean distance, normalized inverse Euclidean distance, discrete curvature, pseudo-dihedral angle, and residue solvent accessible surface area-related feature; in this embodiment, a parallel acceleration mechanism is also adopted, that is, task-level parallelism is realized through a multi-process pool, and computing resources are dynamically allocated.

[0076] The extraction of three-dimensional conformation features is as follows.

[0077] Regarding the space geometry invariant coordinate transformation: that is, by establishing a local coordinate system of the principal axis of inertia to eliminate the spatial rotation and translation variance of the atomic coordinates of the input peptide, in this embodiment, the space geometry invariant coordinate transformation unit and the space geometry feature calculation unit use the Calculate the spatial distribution covariance matrix C and the eigenvector matrix V as follows: Based on all input peptides Coordinate calculation of spatial distribution covariance matrix C: ; in, For the i-th The spatial coordinate vector of an atom; All of the target peptide The mean vector of spatial coordinates, i.e., the centroid coordinates; N is... total.

[0078] Perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvector matrix. and eigenvalue diagonal matrix The decomposition satisfies the following relation: ; in, It is a diagonal matrix composed of eigenvalues; eigenvector matrix This indicates the direction of the principal axis of inertia of the peptide chain.

[0079] By moving the center of mass to the origin and projecting it onto the principal inertial axis coordinate system, rotation and translation invariance are achieved, eliminating the influence of the spatial pose of the peptide conformation on spatial geometric features. ; in, For the i-th The spatial coordinate vector of an atom after a rotation- and translation-invariant transformation. .

[0080] When the peptide chain length in the input dataset varies greatly, such as with a coefficient of variation (CV) ≥ 0.3, the following standardized scaling normalization can be used to scale the coordinates to a distribution with a mean of 0 and a standard deviation of 1, so that peptide chains of different lengths have comparable spatial characteristics: ; Where σ represents all The standard deviation vector on each coordinate axis; ε is a very small constant (e.g., 1e-8) to prevent division by zero.

[0081] Standard deviation vector ( , , Calculated independently by coordinate axis components: ; ; ; in, , , is the i-th standard residue of the peptide chain x, y, z coordinate components in the principal axis inertial rotated post coordinate system; , , is the mean of x, y, z coordinate components in the principal axis inertial rotated post coordinate system; N is the total number of standard residues of the peptide chain.

[0082] about Euclidean distance and normalized inverse Euclidean distance are calculated as: Euclidean distance between residues is calculated using geometric invariance coordinates (GIC): ; Min-Max normalization is performed on Euclidean distance: ; Normalized inverse Euclidean distance is calculated as: ; about discrete curvature is calculated as: ; where , , is the mean of x, y, z coordinate components in the principal axis inertial rotated post coordinate system; N is the total number of standard residues of the peptide chain. .

[0083] The calculated discrete curvature vector is aligned with the length of the peptide chain to obtain the discrete curvature feature vector of the peptide chain : ; about pseudo dihedral angle is calculated as: ; ; where , , , .

[0084] The calculated pseudo dihedral angle vector is aligned with the length of the peptide chain to obtain the pseudo dihedral angle feature vector of the peptide chain​​​​ Aligning with the peptide chain length N, we obtain the pseudo-dihedral feature vector of the peptide chain. ) : ; in, .

[0085] Regarding the Solvent Accessible Surface Area (SASA) characteristic calculate: In this embodiment, a spherical probe algorithm with a probe radius of 1.4 Å is used to calculate the solvent-accessible surface area of ​​residues. Based on the solvent-accessible surface area of ​​residues, residue-level SASA feature vectors are calculated. This includes: absolute total SASA value, relative total SASA value, polar surface area ratio, side chain surface area ratio, and exposure status marker (when the SASA of the residue is >40). .

[0086] Regarding the calculation of conformation-related AAIndex features: In this embodiment, a predefined AAAindex index system is used to calculate the combination of AAAindex indices related to the conformation of the 6-HB targeting peptide, including but not limited to: ZHOH04010 (interface tendency), RADA880106 (transmembrane helix tendency), CHAM820101 (α-helix tendency), CHOC760101 (β-turn tendency), FAUJ880103 (residue volume), ROSG850101 (side chain spatial complexity), CHAM830107 (van der Waals volume), and CHAM820102 (β-sheet tendency); the output feature matrix is ​​shown. ,in, The number of selected AAIndex indicators (in this example) ).

[0087] Step a73, based on Euclidean distance is used to generate a binary adjacency matrix.

[0088] In this embodiment, based on Euclidean distance matrix, with a fixed distance threshold. Generate a binary adjacency matrix: ; in, , The total number of standard amino acids in the peptide chain. Based on the distribution of non-bonded interaction distances in the crystal structure of the 6-HB complex, the preferred method is... In this embodiment =8.0.

[0089] Step a74, input the general sequence characteristics into the general antiviral peptide activity prediction model with frozen weights to obtain a high-dimensional sequence characteristic representation.

[0090] Step a75, splice the high-dimensional sequence characteristic representation and the pre-computed spatial conformation characteristics along the feature dimension to serve as enhanced node features; and the normalized inverse Euclidean distance as the edge weight to construct a second graph structure.

[0091] In this embodiment, the node-level sequence representation extracted by the pre-trained model is spliced with the standardized conformation-related AAindex features, pseudo-dihedral angle, standardized discrete curvature, the rotationally and translationally invariant coordinates, and the standardized SASA-related features along the feature dimension are spliced to construct a joint node feature matrix; in addition, the normalized inverse Euclidean distance is taken as the edge weight: ; wherein, is the Euclidean distance between residues i and j in the rotationally and translationally invariant coordinate system; is a very small constant 1e-7.

[0092] Step a76, input the second graph structure into the second initial multi-layer graph convolutional network model after feature migration, and obtain a 6-HB targeting membrane fusion inhibiting peptide prediction model when the preset training indicators are met.

[0093] The second initial multi-layer graph convolutional network architecture in this embodiment includes two graph convolutional layers, each of which performs a standardized graph convolutional operation: ; wherein, is an adjacency matrix with self-loops added; is the degree matrix of , satisfying ; is the trainable weight matrix of the i-th layer; In this embodiment, the hidden dimension of the graph convolutional layer is preferably 64; the feature reuse mechanism between layers is: ; wherein, is the input feature matrix of the i-th layer, is the trainable weight matrix of the i-th layer, denotes a graph convolutional operation with an activation function.

[0094] After aggregating node-level features into a graph-level feature vector by global average pooling, a fully connected layer is used to map the graph-level features to binary classification logits output.

[0095] In the model training phase, the mean and standard deviation of the residue solvent accessible surface area related feature matrix, the discrete curvature feature matrix and the conformation related AAindex feature matrix on the training set are calculated before the pre-extracted sequence representation is spliced with the spatial conformation feature, and the z-score standard is applied to the training set, the validation set and the test set respectively.

[0096] The training strategy in this embodiment is as follows: Optimizer configuration: optimizer type AdamW, initial learning rate 1e-3; Early stopping strategy: monitor the validation loss, and trigger early stopping when the validation loss does not improve for 10 consecutive epochs (epochs); Performance detection index: monitor the comprehensive performance index of the validation set during the training process, including ROC-AUC, accuracy, recall, F1 score, Matthew correlation coefficient, specificity and area under the PR curve.

[0097] In order to alleviate the influence of class imbalance on model performance, a phased processing strategy is adopted, and the implementation method is the same as the first phase.

[0098] In this embodiment, the second phase classification model, i.e. the 6-HB targeting membrane fusion inhibitory peptide prediction model , the positive and negative samples are close to balance, and no class weighting is performed on the loss function. In order to optimize the decision performance, temperature scaling and decision threshold calibration are implemented as in the first phase.

[0099] In this embodiment, by integrating the sequence and conformation features of antiviral peptides, the key information related to 6-HB targeting can be more comprehensively captured, and the effective model can improve the classification performance of 6-HB targeting membrane fusion inhibitory peptides.

[0100] Based on the feature transfer strategy, the sequence features and spatial conformation features of antiviral peptides are combined, which significantly improves multiple performance indicators of the model on the target task test set, as shown in Figure 4 In addition, by freezing the pre-trained model, the number of trainable parameters for fine-tuning is greatly reduced, the risk of overfitting is effectively reduced, and the model training time is shortened.

[0101] In some optional embodiments, the above step S101, i.e. generating 6-HB targeting membrane fusion inhibitory candidate peptides, comprises: Step S1011, obtaining antiviral peptide data. In this embodiment, the pre-training data set contains 2051 antiviral peptides, which are randomly divided into training set, validation set and test set in the ratio of 6:2:2.

[0102] Step S1012, pre-training the model architecture based on the antiviral peptide data to obtain a general antiviral peptide generation model.

[0103] Step S1013, obtaining 6-HB targeting membrane fusion inhibition peptide data.

[0104] In this embodiment, the fine-tuning data set contains 371 experimentally verified membrane fusion inhibition peptides, which are randomly divided into training set, validation set and test set in the ratio of 6:2:2.

[0105] Step S1014, fine-tuning the general antiviral peptide generation model based on the 6-HB targeting membrane fusion inhibition peptide data to obtain a 6-HB targeting membrane fusion inhibition peptide generation model; the 6-HB targeting membrane fusion inhibition peptide generation model is used to generate membrane fusion inhibition candidate peptides.

[0106] In this embodiment, a two-stage transfer learning framework is adopted, the first stage pre-trains the model on the broad-spectrum antiviral peptide data set to enable it to have the ability to generate general antiviral peptide sequences; the second stage performs end-to-end fine-tuning on the 6-HB targeting membrane fusion inhibition peptide data set to give the model the ability to generate 6-HB targeting. This strategy not only enables the model to have both general generation and specific targeting generation capabilities, but also effectively suppresses overfitting and significantly improves the generalization performance of the model, as shown in Figure 5 .

[0107] In some optional embodiments, as shown in Figure 6 , the 6-HB targeting membrane fusion inhibition peptide generation model comprises: a variable-length sequence preprocessing module for preprocessing the input sequence; preprocessing the input sequence includes: sorting the input sequence in descending order of actual length; padding the sequence sorted in descending order at the end to a uniform length, and generating a corresponding padding mask to mark the valid data position, and skipping the redundant calculation of the padding symbol by the LSTM through a dynamic masking mechanism; an embedding layer connected to the output end of the variable-length sequence preprocessing module, for mapping the preprocessed integer index sequence to a dense vector representation in a high-dimensional space; an LSTM layer connected to the output end of the embedding layer, for capturing long-range dependencies and context information in the preprocessed vector sequence; a fully connected layer, the output of the LSTM layer is finally mapped to a probability distribution in the output space by the fully connected layer; The 6-HB targeting membrane fusion inhibition peptide generation model adopts a dynamic training control strategy in the training stage, which combines gradient norm constraint and continuous validation loss monitoring to improve sequence stability and model generalization ability; The sequence generation stage is equipped with a generated sequence filtering unit; the generated sequence filtering unit performs the following three operations: adjusting the output probability distribution of the Softmax function using the temperature parameter to control the generation diversity; dynamically filtering non-amino acid tokens, ensuring that each step of generation only samples in the valid amino acid dictionary; in response to the end symbol, the generation process is immediately interrupted once a special symbol representing the end of the sequence is generated.

[0108] In particular, with regard to the variable-length sequence preprocessing unit, refer to Figure 7 as shown in the figure: Add control symbols at the beginning and end of the sequence: add start symbols at the beginning and end of the sequence, respectively <sos>and end markers <eos>; Vocabulary encoding system construction: A 24-dimensional vocabulary encoding system was constructed, including 20 standard amino acids and 4 control symbols (filler <pad>, start symbol <sos>, end of character <eos>, unknown symbol <unk>); Ranking and padding: sequences ranked in descending order by number of amino acids; short sequences padded at the end <pad>to a uniform length; Variable-length compression metadata generation: skip padding position calculation by sequence compression algorithm (pack_padded_sequence), only keep the hidden state output of the valid sequence segment in the LSTM layer processing.

[0109] Regarding the above-mentioned generated sequence filtering unit, in the embodiment, the temperature parameter is used to adjust the output probability distribution of the Softmax function; when the model outputs the preset sequence terminator ( <eos>) immediately terminate the generation of the current sequence; and also reject non-amino acid special tokens in the vocabulary <pad> , <sos> , <eos> , <unk>As valid amino acid outputs; and also automatically discard the generated candidate sequences with valid peptide sequence length less than 5. In this embodiment, the embedding layer: maps the sequence of 24-type token encodings into a 32-dimensional distributed representation; the LSTM layer: 2 stacked layers, 48 hidden units (scalable to 32-64), 0.5 dropout rate between layers; the fully connected classification layer: output dimension 24, corresponding to the vocabulary probability distribution; the loss function is the cross-entropy with padding mask (set ignore_index=0).

[0110] Optimizer configuration: optimizer type AdamW, initial learning rate 1e-3, weight decay coefficient 0.01; Gradient clipping mechanism: gradient norm threshold 1.0 (L2 constraint), to prevent gradient explosion; Early stopping strategy: the pre-training early stopping strategy is set to terminate training when the validation loss does not improve for 20 consecutive training epochs; the corresponding setting in the fine-tuning stage is 10 epochs. In this embodiment, the pre-training stage triggers early stopping at the 190th epoch, and the optimal model has a validation loss of 2.5288 and a validation confusion of 12.5379. The fine-tuning stage triggers early stopping at the 104th epoch, and the optimal model has a validation loss of 1.8965 and a validation confusion of 6.6623.

[0111] Gradient masking mechanism: exclude padding positions from gradient calculation by setting the ignore_index parameter of the cross-entropy loss to the padding index 0.

[0112] The 6-HB targeting membrane fusion inhibitory peptide generation model uses an unconditional autoregressive mode in the inference stage, i.e., sequence generation is performed by the start symbol ( <sos>) are triggered without relying on externally provided conditional information; the next token prediction at each time step only depends on the model hidden state before that time step; and the model does not accept any external input of protein structure information or functional labels as generation conditions during the generation process.

[0113] The specific configuration of model inference generation is as follows: the sequence generation technology adopts a self-recurrent inference mechanism. Among them, the start symbol <sos>(Index 1) Trigger sequence generation; temperature sampling: diversity regulation coefficient τ = 0.7; termination condition: output <eos>(Index 2) or forced termination at maximum length of 50 residues. Then, sequence filtering is performed, i.e. special symbols are masked: <pad>(Index 0), <sos>(Index 1), <eos>(Index 2), <unk>(Index 3), only standard amino acid symbols (Index 4-23) are reserved; Length constraint: valid sequence length >= 5 residues; automatically discard invalid sequences with length < 5; batch output: valid sequences are stored in FASTA format; default: generate 100 candidate peptides at a time (range adjustable).

[0114] In this embodiment, the effective token loss and perplexity are used as the model performance evaluation indexes: Effective token loss: ; wherein, is the cross-entropy loss of the i-th token; is an indicator function, when the token non-padding <pad>where T is 1 if the word is in the vocabulary, otherwise 0, B is the batch size, i.e., the total number of tokens in a batch. The valid token loss on the validation set in this example is 2.4988.

[0115] Perplexity: a general indicator to evaluate the quality of sequence generation, the lower the value, the stronger the predictive ability of the model. The perplexity on the validation set in this example is 12.1680.

[0116] The 6-HB targeting membrane fusion inhibition peptide generation model constructed in this example can process variable-length sequence input and reduce redundant calculations through a dynamic masking mechanism; the sequence stability and model generalization ability are improved through gradient clipping and continuous validation loss monitoring. In the inference stage, the diversity of the output distribution is regulated through the temperature parameter; and the reliability and effectiveness of the generated sequence are ensured by integrating the dynamic filtering of non-amino acid tokens and the response end symbol response mechanism.

[0117] In some optional embodiments, the step S103 of screening the classification result to obtain the final 6-HB targeting membrane fusion inhibition candidate peptide comprises: Step S1031, determining the initial candidate peptide in the classification result, the initial candidate peptide being a candidate peptide classified as a membrane fusion inhibition peptide.

[0118] Step S1032, performing test molecular docking between each initial candidate peptide and the 6-HB conformation of the target virus to obtain a docking score corresponding to each initial candidate peptide.

[0119] Step S1033, determining whether the docking score meets a preset condition.

[0120] Step S1034, determining the initial candidate peptide corresponding to the docking score that meets the preset docking score condition as a high-specificity candidate peptide targeting 6-HB.

[0121] In the above process, an LSTM autoregressive model under a two-stage transfer learning framework is first used to infer and generate 100 (the specific number is not limited) 6-HB targeting membrane fusion inhibition candidate peptides. The three-dimensional structure of the predicted candidate peptide is predicted, and the sequence features of the candidate peptide are calculated, including amino acid type encoding, physicochemical property features, evolutionary semantic features, and spatial conformation features, including conformation-related AAindex features, rotation and translation invariant coordinates, Euclidean distance, normalized inverse Euclidean distance, discrete curvature, pseudo-dihedral angle, and residue solvent accessible surface area related features. The 6-HB targeting membrane fusion inhibition peptide prediction model is used to classify and predict the 6-HB targeting property of the candidate peptide.

[0122] Furthermore, candidate peptides with a predicted tag of 1 (positive) were selected and subjected to protein-peptide flexible docking with the pre-fusion conformation of influenza A H3N2 subtype (or other enveloped viruses or subtypes, as determined according to needs) 6-HB to obtain docking scores. The docking software is HADDOCK (or Rosetta FlexPep Dock or HelixFold3 or other complex prediction software).

[0123] Several experimentally validated membrane fusion inhibitory peptides of influenza A H3N2 were collected from DRAVP (or other sources) as positive controls. These peptides were then docked with the target virus 6-HB under identical methods and parameter settings to obtain scores. ; Calculate the candidate peptides in the docking system Spiral content, selection Spiral content greater than or equal to preset threshold Candidate peptides; in this embodiment Set it to 60% (this step is optional).

[0124] Based on candidate peptide docking score Compared with positive control peptide By comparing the docking score statistics, candidate peptides with docking affinity superior to the statistical criteria of the positive control peptide were selected as high-specificity inhibitory peptide candidates for the target virus 6-HB. mean and standard deviation The comparison relationship is the docking score. Candidate peptides that meet the following criteria are identified as highly specific candidate peptides for the target virus 6-HB: ; in, This is a predefined constant, set to 1 in this embodiment.

[0125] In this embodiment, by performing experimental docking on each initial candidate peptide and further screening based on the docking score, the success rate of downstream biological validation experiments can be effectively improved.

[0126] This embodiment also provides a 6-HB-targeted membrane fusion inhibitory peptide prediction device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0127] The embodiment provides a 6-HB targeting membrane fusion inhibiting peptide prediction device, which comprises the following modules: Figure 8 The device comprises the following modules: A sequence generation module 201 is configured to generate a 6-HB targeting membrane fusion inhibiting candidate peptide. A prediction module 202 is configured to input the membrane fusion inhibiting candidate peptide into a pre-constructed 6-HB targeting membrane fusion inhibiting peptide prediction model to perform classification prediction, so as to obtain a classification result of the membrane fusion inhibiting candidate peptide; wherein the membrane fusion inhibiting peptide recognition model is a two-stage transfer learning classification model, the first stage is based on sequence general features to perform binary classification prediction of antiviral peptides and non-antiviral peptides; and the second stage is based on spatial conformation features to perform binary classification prediction of membrane fusion inhibiting peptides and non-membrane fusion inhibiting antiviral peptides. A screening module 203 is configured to screen the classification result, so as to obtain a final 6-HB targeting membrane fusion inhibiting candidate peptide.

[0128] In some optional embodiments, the device further comprises the following modules: A construction module is configured to obtain a first data set, the first data set comprising antiviral peptides and non-antiviral peptides, wherein the antiviral peptides are used as positive data sets, and the non-antiviral peptides are used as negative data sets; a first initial multi-layer graph convolutional network model is constructed; the first initial multi-layer graph convolutional network model is trained based on the first data set, so as to obtain a trained general antiviral peptide activity prediction model; wherein the following steps are included: data preprocessing is performed on the first data set; sequence general features are calculated based on the preprocessed first data set, the sequence general features comprising amino acid type coding features, physicochemical property features and evolutionary semantic features; a binary adjacency matrix is constructed by using a k-neighbor method based on sequence distance based on the preprocessed first data set; the sequence general features are used as node features, and the reverse sequence distance is used as edge weight, so as to construct a first graph structure; the first graph structure is input into the first initial multi-layer graph convolutional network model, so as to obtain the general antiviral peptide activity prediction model when a preset training index is met.

[0129] The construction module is also configured to obtain a second data set, the second data set comprising 6-HB targeting membrane fusion inhibitory peptides and non-6-HB targeting antiviral peptides, wherein the 6-HB targeting membrane fusion inhibitory peptides are taken as positive data sets, and the non-6-HB targeting antiviral peptides are taken as negative data sets; a second initial multi-layer graph convolution network model with the same architecture as the first initial multi-layer graph convolution network model is constructed; the general antiviral peptide activity prediction model weight is frozen as a sequence feature extractor, and the extracted sequence general features are used for fusion to node features of the second initial multi-layer graph convolution network architecture model; the general sequence features extracted by the general antiviral peptide activity prediction model are spliced with the pre-calculated spatial conformation features based on the second data set, the second initial multi-layer graph convolution network model is trained, and a 6-HB targeting membrane fusion inhibitory peptide prediction model is obtained. Wherein, it comprises: data preprocessing of the second data set; based on the preprocessed second data set, spatial conformation features are extracted, the spatial conformation features comprising: rotationally translationally invariant coordinates, Euclidean distance, normalized inverse Euclidean distance, discrete curvature, pseudo-dihedral angle, conformation-related AAindx feature, and residue solvent accessible surface area-related feature; based on Euclidean distance, a binary adjacency matrix is generated; the general sequence features are spliced along the feature dimension and input to the general antiviral peptide activity prediction model with frozen weights to obtain high-dimensional sequence feature representation; the high-dimensional sequence feature representation and the spatial conformation features are spliced along the feature dimension and used as enhanced node features; and the normalized inverse Euclidean distance is taken as an edge weight to construct a second graph structure; the second graph structure is input to the second initial multi-layer graph convolution network model after feature migration, and when a preset training index is met, the 6-HB targeting membrane fusion inhibitory peptide prediction model is obtained.

[0130] 6-HB targeting membrane fusion inhibitory peptide generation model, comprising: a variable-length sequence preprocessing module, configured to preprocess an input sequence; preprocessing the input sequence comprises: sorting the input sequence in descending order of actual length; padding the sequence sorted in descending order to a uniform length at the end thereof, and generating a corresponding padding mask to mark the valid data position; an embedding layer connected to the output end of the variable-length sequence preprocessing module, configured to map the integer index sequence after preprocessing to a dense vector representation in a high-dimensional space; an LSTM layer connected to the output end of the embedding layer, configured to capture long-range dependencies and context information in the vector sequence after preprocessing; a fully connected layer, the output of the LSTM layer is finally mapped to a probability distribution of the output space by the fully connected layer; the 6-HB targeting membrane fusion inhibitory peptide generation model adopts a dynamic training control strategy in the training stage, which combines gradient norm constraint and continuous validation loss monitoring; the 6-HB targeting membrane fusion inhibitory peptide generation model is equipped with a generated sequence filtering unit in the sequence generation stage; the generated sequence filtering unit performs the following three operations: using a temperature parameter to adjust the output probability distribution of the Softmax function to control the generation diversity; dynamically filtering non-amino acid tokens to ensure that each step of generation only samples in the valid amino acid dictionary; response terminator, immediately interrupt the generation process as soon as a special symbol representing the end of the sequence is generated.

[0131] In some optional embodiments, the sequence generation module 201 comprises: a sequence generation unit configured to: obtain antiviral peptide data; pretrain a pre-constructed model architecture based on the antiviral peptide data to obtain a general antiviral peptide generation model; obtain 6-HB targeting membrane fusion inhibitory peptide data; fine-tune the general antiviral peptide generation model based on the 6-HB targeting membrane fusion inhibitory peptide data to obtain a 6-HB targeting membrane fusion inhibitory peptide generation model; and use the 6-HB targeting membrane fusion inhibitory peptide generation model to generate membrane fusion inhibitory candidate peptides.

[0132] In some optional embodiments, the screening module 203 comprises: a screening unit configured to: determine initial candidate peptides in the classification result, the initial candidate peptides being membrane fusion inhibitory candidate peptides classified as 6-HB targeting membrane fusion inhibitory peptides; perform trial molecular docking of each initial candidate peptide with a target virus 6-HB conformation to obtain a docking score corresponding to each initial candidate peptide; determine whether the docking score meets a preset condition; and determine a candidate peptide meeting the preset docking score condition as a final 6-HB targeting membrane fusion inhibitory candidate peptide.

[0133] The 6-HB targeting membrane fusion inhibitory peptide prediction device in this embodiment is presented in the form of a functional unit. Here, the unit refers to an ASIC circuit, a processor and a memory executing one or more software or fixed programs, and / or other devices that can provide the above functions.

[0134] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0135] This invention also provides a computer device having the above-described features. Figure 8 The device shown is a 6-HB-targeted membrane fusion inhibitory peptide prediction device.

[0136] Please see Figure 9 , Figure 9 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 9 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 9 Take a processor 10 as an example.

[0137] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0138] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0139] The memory 20 can include a program storage area and a data storage area. The program storage area can store an operating system, application programs required for at least one function, and the like. The data storage area can store data created according to the use of the computer device, and the like. In addition, the memory 20 can include a high-speed random access memory, and can further include a non-transitory memory such as at least one of a magnetic disk storage device, a flash memory device, or other non-transitory solid state memory device. In some alternative embodiments, the memory 20 can optionally include a memory disposed remotely from the processor 10, which can be connected to the computer device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0140] The memory 20 can include a volatile memory such as a random access memory, and can further include a non-volatile memory such as a flash memory, a hard disk, or a solid state disk, and a combination thereof.

[0141] The computer device further includes a communication interface 30 for communication of the computer device with other devices or communication networks.

[0142] The embodiments of the present application also provide a computer readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or as software code to be recorded in a storage medium, or originally stored in a remote storage medium or a non-transitory machine readable storage medium to be downloaded through a network and stored in a local storage medium, so that the method described herein can be processed by such software using a general purpose computer, a special purpose processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid state disk, and the like. Further, the storage medium can include a combination of the above-mentioned storage media. It can be understood that the computer, the processor, the microprocessor controller, or the programmable hardware includes a storage component that can store or receive software or computer code, which, when accessed and executed by the computer, the processor, or the hardware, implements the method shown in the above embodiments.

[0143] Although the embodiments of the present application have been described with reference to the accompanying drawings, various modifications and changes can be suggested to one skilled in the art, and it is intended that the present application encompass such modifications and changes as fall within the scope of the appended claims.< / pad> < / unk> < / eos> < / sos> < / pad> < / eos> < / sos> < / sos> < / unk> < / eos> < / sos> < / pad> < / eos> < / pad> < / unk> < / eos> < / sos> < / pad> < / eos> < / sos>

Claims

1. A 6-HB targeting membrane fusion inhibiting peptide prediction method, characterized in that, The method comprises: generating 6-HB targeting membrane fusion inhibition candidate peptides; inputting the membrane fusion inhibition candidate peptides into a pre-constructed 6-HB targeting membrane fusion inhibition peptide prediction model for classification prediction to obtain the classification results of the membrane fusion inhibition candidate peptides; wherein the membrane fusion inhibition peptide recognition model is a two-stage transfer learning classification model, the first stage is based on sequence general features, and the antiviral peptide and the non-antiviral peptide are classified and predicted; the second stage integrates spatial conformation features, and the membrane fusion inhibition peptide and the non-membrane fusion inhibition antiviral peptide are classified and predicted; screening the classification results to obtain the final 6-HB targeting membrane fusion inhibition candidate peptides.

2. The method of claim 1, wherein, The 6-HB targeting membrane fusion inhibition peptide prediction model is established by the following steps: obtain a first data set, the first data set includes antiviral peptides and non-antiviral peptides, wherein the antiviral peptides are used as positive data sets, and the non-antiviral peptides are used as negative data sets; construct a first initial multi-layer graph convolution network model; based on the first data set, the first initial multi-layer graph convolution network model is trained to obtain a trained general antiviral peptide activity prediction model; obtain a second data set, the second data set includes 6-HB targeting membrane fusion inhibition peptides and non-6-HB targeting antiviral peptides, wherein the 6-HB targeting membrane fusion inhibition peptides are used as positive data sets, and the non-6-HB targeting antiviral peptides are used as negative data sets; construct a second initial multi-layer graph convolution network model with the same architecture as the first initial multi-layer graph convolution network model; freeze the weight of the general antiviral peptide activity prediction model as a sequence general feature extractor, and the extracted sequence general features are used for fusion to the node features of the second initial multi-layer graph convolution network model; based on the second data set, the general sequence features extracted by the general antiviral peptide activity prediction model are spliced with the pre-calculated spatial conformation features, and the second initial multi-layer graph convolution network model is trained to obtain the 6-HB targeting membrane fusion inhibition peptide prediction model.

3. The method of claim 2, wherein, The training of the first initial multi-layer graph convolution network model based on the first data set comprises: data preprocessing is performed on the first data set; based on the preprocessed first data set, sequence general features are calculated, which include amino acid type encoding features, physicochemical property features, and evolutionary semantic features; based on the preprocessed first data set, a binary adjacency matrix is constructed by using a k-neighbor method based on sequence distance; the sequence general features are used as node features, and the reverse sequence distance is used as edge weight to construct a first graph structure; input the first graph structure into the first initial multi-layer graph convolution network model, and obtain a general antiviral peptide activity prediction model when the preset training indicators are met.

4. The method of claim 3, wherein, The training of the second initial multi-layer graph convolution network model based on the second data set, which splices the general sequence features extracted by the general antiviral peptide activity prediction model with the pre-calculated spatial conformation features, comprises: data preprocessing is performed on the second data set; Based on the pre-processed second data set, spatial conformation features are extracted, including: rotationally translationally invariant coordinates, Euclidean distance, normalized inverse Euclidean distance, discrete curvature, pseudo-dihedral angle, conformation-related AAindex features, residue solvent-accessible surface area-related features; based on the Euclidean distance, generating a binary adjacency matrix; concatenate the general sequence features along the feature dimension, input into the general anti-virus peptide activity prediction model after the weight is frozen, to obtain a high-dimensional sequence feature representation; concatenate the high-dimensional sequence feature representation and the spatial conformation feature along the feature dimension, together as enhanced node features; and construct a second graph structure by taking the normalized inverse Euclidean distance as an edge weight; input the second graph structure into the second initial multi-layer graph convolution network model after feature migration, and obtain the 6-HB targeting membrane fusion inhibition peptide prediction model when a preset training index is met.

5. The method of claim 1, wherein, The method for generating a 6-HB targeting membrane fusion inhibition candidate peptide comprises: obtaining anti-virus peptide data; pre-training a previously constructed model architecture based on the anti-virus peptide data, to obtain a general anti-virus peptide generation model; obtaining 6-HB targeting membrane fusion inhibition peptide data; fine-tuning the general anti-virus peptide generation model based on the 6-HB targeting membrane fusion inhibition peptide data, to obtain a 6-HB targeting membrane fusion inhibition peptide generation model; the 6-HB targeting membrane fusion inhibition peptide generation model is used to generate the membrane fusion inhibition candidate peptide.

6. The method of claim 1, wherein, The 6-HB targeting membrane fusion inhibition peptide generation model comprises: a variable-length sequence preprocessing module, configured to preprocess an input sequence; the preprocessing of the input sequence comprises: sorting the input sequence in descending order of actual length; padding the sequence sorted in descending order to a uniform length at the end thereof, and generating a corresponding padding mask to mark the valid data positions; an embedding layer, connected to an output end of the variable-length sequence preprocessing module, configured to map the integer index sequence after preprocessing to a dense vector representation in a high-dimensional space; an LSTM layer, connected to an output end of the embedding layer, configured to capture long-range dependencies and context information in the vector sequence after preprocessing; a fully connected layer, configured to map the output of the LSTM layer to a probability distribution in an output space; the 6-HB targeting membrane fusion inhibition peptide generation model adopts a dynamic training control strategy in the training phase, which combines gradient norm constraint and continuous validation loss monitoring; the 6-HB targeting membrane fusion inhibition peptide generation model is equipped with a generated sequence filtering unit in the sequence generation phase; the generated sequence filtering unit performs the following three operations: using a temperature parameter to adjust the output probability distribution of the Softmax function to control the generation diversity; dynamically filtering non-amino acid tokens to ensure that each step of generation only samples in the valid amino acid dictionary; responding to the end symbol, the generation process is immediately interrupted once a special symbol representing the termination of the sequence is generated.

7. The method of claim 1, wherein, The method for screening the classification result to obtain the final 6-HB targeting membrane fusion inhibition candidate peptide comprises: determining an initial candidate peptide in the classification result, the initial candidate peptide being the membrane fusion inhibition candidate peptide classified as a 6-HB targeting membrane fusion inhibition peptide; performing test molecular docking of each initial candidate peptide with a target virus 6-HB conformation to obtain a docking score corresponding to each initial candidate peptide; judging whether the docking score meets a preset condition; The initial candidate peptide satisfying the preset docking score condition is determined as the final 6-HB targeting membrane fusion inhibition candidate peptide.

8. A 6-HB targeting membrane fusion inhibiting peptide prediction apparatus characterized by, The device comprises: a sequence generation module for generating a 6-HB targeting membrane fusion inhibition candidate peptide; a prediction module for inputting the membrane fusion inhibition candidate peptide into a pre-constructed 6-HB targeting membrane fusion inhibition peptide prediction model for classification prediction to obtain a classification result of the membrane fusion inhibition candidate peptide; wherein the membrane fusion inhibition peptide recognition model is a two-stage transfer learning classification model, the first stage is based on sequence general features for binary classification prediction of antiviral peptides and non-antiviral peptides; the second stage integrates spatial conformation features for binary classification prediction of membrane fusion inhibition peptides and non-membrane fusion inhibition antiviral peptides; a screening module for screening the classification result to obtain the final 6-HB targeting membrane fusion inhibition candidate peptide.

9. A computer device, comprising: comprise: a memory and a processor, which are in communication connection with each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the 6-HB targeting membrane fusion inhibition peptide prediction method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the computer to execute the 6-HB targeting membrane fusion inhibition peptide prediction method of any one of claims 1-7.

Citation Information

Patent Citations

  • Targeted antigen peptide sequence generation and screening method based on active skeleton

    CN118522342A

  • Identification method and system of short antibacterial peptide sequence, terminal and storage medium

    CN119541641A

  • Targeting peptide recognition method and device, electronic equipment and storage medium

    CN119694400A

  • Marine bacteriophage antibacterial lysin prediction method, device, medium and equipment

    CN120260696A

  • Prediction method and prediction system for combination of CD4 + T cell receptor and polypeptide based on deep learning

    CN120412705A