Self-supervised learning pre-training method and system based on protein size prompt

By introducing a self-supervised learning pre-training method with protein size hints in protein representation learning, the problem of the difference in the distribution of pre-training data and downstream task data in protein representation learning is solved, and the model adaptability and prediction performance are improved.

CN119993286APending Publication Date: 2025-05-13FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510067568.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing self-supervised learning methods ignore the distribution differences between pre-training data and downstream task data in protein representation learning, especially the problem of inconsistent protein size distribution, resulting in insufficient adaptability and transfer capabilities of the pre-trained model.

Method used

A self-supervised learning pre-training method based on protein size hints is adopted. By obtaining protein data sets, a pre-training network with fusion protein scale hints is constructed and trained, and the labelless data is used for pre-training, and the inconsistency of upstream and downstream task data is reduced through protein scale hints.

Benefits of technology

It significantly improves the adaptability of the model on different protein sizes, solves the problem of inconsistent data distribution between pre-training and downstream tasks, reduces the dependence on labeled data, and improves the predictive performance of protein properties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993286A_ABST
    Figure CN119993286A_ABST
Patent Text Reader

Abstract

The invention discloses a self-supervised learning pre-training method and a self-supervised learning pre-training system based on protein size prompt. The self-supervised learning pre-training method comprises the following steps: firstly, collecting unlabeled protein data from a protein database; encoding the protein size information into a prompt vector by using a size prompt adapter, and embedding the prompt vector into a protein encoder; then, based on the graph or point cloud representation of the protein, executing a mask prediction task to complete pre-training; and finally, performing fine adjustment on the pre-training model in a downstream task, and optimizing the specific task performance of the pre-training model. According to the method, the problem of inconsistent data distribution between the pre-training and the downstream task is remarkably reduced through protein size prompt, so that the universality of the pre-training model and the performance of the downstream task are improved. The method has high data efficiency, model universality and expansibility, and is suitable for various tasks such as protein function prediction and binding site detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of bioinformatics, and in particular relates to a self-supervised learning pre-training method and system based on protein size prompts. Background Art

[0002] As the basic molecule of life activities, protein is widely involved in many important biological processes such as catalysis, signal transduction, and structural support. The function and structure of protein are not only the key to understanding the essence of life, but also the core content of many fields such as new drug development, disease treatment, and bioengineering. However, although the global protein database has recorded more than 400 million proteins, about 80% of them lack clear functional annotations, which greatly restricts the research and application in related fields.

[0003] Traditional protein function analysis methods usually rely on manually designed feature extraction strategies, such as amino acid sequence analysis, domain annotation, and motif recognition. Although these methods can provide a preliminary understanding of protein function to a certain extent, their performance is limited by the bias in feature selection and the complexity of the biological background. As the number of protein types increases, these traditional knowledge-based methods gradually expose the bottleneck of being unable to handle large-scale data. Especially when dealing with highly complex biological tasks, the applicability and accuracy of traditional methods are increasingly insufficient.

[0004] With the rapid development of deep learning technology, data-driven protein representation learning methods have gradually become the mainstream of research. By constructing complex neural network models, deep learning can automatically learn and extract potential features with higher expressive power from massive protein data. These methods surpass traditional manual feature extraction methods and can capture the complex relationship between protein structure and function at a higher level. Therefore, they have shown great potential in the fields of protein function prediction, drug discovery, disease mechanism analysis, etc. However, these deep learning methods often rely on a large amount of labeled data, and the acquisition of labeled data usually relies on expensive and time-consuming wet experiments, resulting in the bottleneck of data labeling becoming a major obstacle to the widespread application of protein representation learning.

[0005] To solve this problem, self-supervised learning, as an effective unsupervised learning method, has gradually gained attention in the field of protein representation learning. Self-supervised learning uses large-scale unlabeled data for pre-training, so that the model can learn general protein representations without explicit labels, and then fine-tune with a small amount of labeled data to improve the performance of the model in downstream tasks. The advantage of self-supervised learning is that it can greatly reduce the dependence on labeled data, thereby alleviating the challenge of scarce labeled data. Although self-supervised learning methods have shown good prospects in protein representation learning, most existing self-supervised learning methods focus on designing proxy tasks to learn the general characteristics of proteins, ignoring the potential distribution differences between pre-training data and downstream task data. In particular, the factor of protein size distribution differences greatly affects the adaptability and migration ability of the pre-trained model, further restricting its performance in different downstream tasks. Therefore, how to effectively solve the distribution differences between pre-training and downstream tasks, especially the problem of inconsistent protein size distribution, has become a core problem that needs to be broken through in the field of protein representation learning. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a self-supervised learning pre-training method and system based on protein size cues, so that it can be used to pre-train deep networks in protein structure representation learning using a large amount of unlabeled molecular data. After obtaining the prior knowledge of learning a large number of proteins, the deep network's dependence on annotations is alleviated, and the inconsistency of upstream and downstream task data is reduced through protein scale cues, which will improve the protein property prediction performance.

[0007] To achieve the above object, the present invention adopts the following technical solution:

[0008] A self-supervised learning pre-training method based on protein size cues, comprising:

[0009] Step S1, obtaining a protein data set;

[0010] Step S2: constructing and training a pre-trained network based on self-supervised learning according to the protein dataset;

[0011] Step S3: fine-tune the downstream task based on the trained pre-trained network.

[0012] Preferably, in step S1, the protein dataset is converted into a graph modality dataset and a point cloud modality dataset.

[0013] Preferably, in step S2, the pre-trained network includes: an encoder for graph modality that integrates protein scale cues and an encoder for point cloud modality that integrates protein scale cues, extracting multi-scale features of protein structure in a hierarchical manner; wherein the graph modality encoder is used to specifically encode the graph modality dataset; and the point cloud modality encoder is used to specifically encode the graph modality dataset.

[0014] Preferably, in step S2, the pre-trained network to be trained includes:

[0015] Based on the protein dataset, proteins are classified into predefined size intervals according to their residue numbers, and input proteins falling into a specific interval are encoded into corresponding size hint vectors according to the interval sequence number; wherein, the size features of proteins in the same interval are consistent;

[0016] The protein cue vectors are added to the encoders of different modalities and the pre-trained networks are trained to complete the masked amino acid type prediction.

[0017] Preferably, in step S3, the deep network for extracting protein features in the downstream task adopts the trained pre-trained network, and the pre-trained network is fine-tuned by training from scratch using the labeled downstream task.

[0018] The present invention also provides a self-supervised learning pre-training system based on protein size prompts, comprising:

[0019] Acquisition module, used to acquire protein datasets;

[0020] A training module is used to construct and train a pre-trained network based on self-supervised learning according to a protein dataset;

[0021] An optimization module for fine-tuning on downstream tasks based on the pre-trained network after training.

[0022] Preferably, the acquisition module converts the protein dataset into a graph modality dataset and a point cloud modality dataset.

[0023] Preferably, the pre-trained network includes: an encoder for graph modality that integrates protein scale cues and an encoder for point cloud modality that integrates protein scale cues, extracting multi-scale features of protein structure in a hierarchical manner; wherein the graph modality encoder is used to specifically encode the graph modality dataset; and the point cloud modality encoder is used to specifically encode the graph modality dataset.

[0024] Preferably, the training module includes:

[0025] A first processing unit is used for classifying proteins into predefined size intervals according to the number of residues based on the protein data set, and encoding input proteins falling into a specific interval into corresponding size hint vectors according to the interval sequence number; wherein the size features of proteins in the same interval are consistent;

[0026] The second processing unit is used to add the protein prompt vector to the encoders of different modalities and train the pre-trained network to complete the masked amino acid type prediction.

[0027] Preferably, the optimization module uses a pre-trained network after training for extracting protein features in a deep network in a downstream task, and uses labeled downstream tasks to train from scratch to fine-tune the pre-trained network after training.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] The present invention significantly improves the adaptability of the model to different protein sizes by introducing protein size cues, and solves the problem of inconsistent data distribution between pre-training and downstream tasks. The present invention uses unlabeled data for pre-training, reducing the dependence on labeled data, thereby reducing the cost and time of data acquisition. The deep pre-training network of the present invention can extract multimodal features of proteins by fusing graph representation and point cloud representation, thereby improving the integrity and accuracy of the representation. The present invention has a wide range of applications and can be applied to a variety of biological tasks such as protein function prediction and binding site detection, providing a new technical means for research in related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0031] Figure 1 A schematic diagram of a process flow of a self-supervised learning pre-training method according to an embodiment of the present invention;

[0032] Figure 2 This is a general flow chart of self-supervised learning in an embodiment of the present invention;

[0033] Figure 3 A flowchart of protein size feature learning according to an embodiment of the present invention;

[0034] Figure 4 This is a flowchart of protein feature encoding according to an embodiment of the present invention. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0036] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0037] Embodiment 1:

[0038] like Figure 1 As shown, an embodiment of the present invention provides a self-supervised learning pre-training method based on protein size prompts, comprising:

[0039] S1. Collect unannotated protein datasets from public protein databases.

[0040] First, unlabeled data containing protein sequences and three-dimensional structures are downloaded from public protein databases. These datasets do not contain functional annotations or experimental verification information of proteins. Therefore, in order to achieve self-supervised learning, the proteins in the dataset must be representative and cover a variety of protein types and structures. These protein data contain not only their amino acid sequences, but also their three-dimensional structural information, which can be used for subsequent representation learning of graph modalities and point cloud modalities. This lays the foundation for the construction of pre-trained network architecture in subsequent steps.

[0041] Then, the collected protein datasets are converted into graph modality datasets and point cloud modality datasets.

[0042] Graph-modal dataset: The three-dimensional structure of a protein is converted into a graph structure by treating each residue as a node and setting the edges between amino acids according to the distance threshold. The characteristics of each node (residue) are determined by the amino acid type of the node. These graph-modal data will help the network capture the local and global structural features of the protein.

[0043] Point cloud modality dataset: The spatial distribution of proteins is represented by point clouds. Each residue position corresponds to a point in the point cloud. The point cloud data reflects the three-dimensional geometric properties and chemical properties of proteins. Through the point cloud modality, the network can learn the spatial conformation of proteins and their spatial positional relationships.

[0044] Through these preprocessing steps, it is ensured that the image modality and point cloud modality data can be used as network input for further training.

[0045] S2. Build a pre-trained network architecture

[0046] Based on the collected protein dataset, a pre-trained network architecture for self-supervised learning is constructed. Figure 2 As shown, the self-supervised learning pre-training method based on protein size cues is similar to traditional self-supervised learning. The deep network encoder is first pre-trained on an unlabeled dataset, and then the pre-trained deep network encoder weights are applied to downstream tasks and fine-tuned on labeled downstream task datasets. The only difference is the introduction of a protein scale adapter. The first step (S201) is to encode the protein size, and the second step (S202) is to consider how to incorporate this size information into encoders for different modalities. The generated protein size cue vector is embedded into the graph modality encoder and the point cloud modality encoder. In this way, the network can consider the size changes of proteins during feature extraction, thereby effectively dealing with the impact of protein size differences.

[0047] S201. Encoding of protein size information

[0048] like Figure 3 As shown in Figure 1, in this step, the size information of the protein (i.e., the size of the protein) is introduced into the pre-trained network by means of a protein size hint. Based on the number of residues in the protein, the protein size is divided into multiple intervals, and a unique hint vector is generated for each interval. The hint vector is embedded as hint information in the subsequent model to help the network consider the size information of the protein when processing the protein.

[0049] S202 encoder construction

[0050] Graph encoder construction: Figure 4 As shown in the figure, a general graph convolutional network is used in this example, which is mainly used to process the graph modal data of proteins. The graph representation of proteins consists of nodes (residues) and edges (chemical bonds). The graph convolutional network extracts the structural features of proteins layer by layer by performing convolution operations on these nodes and edges. Usually, the graph convolutional network has many layers. In each layer of graph convolution, the representation of the node is updated by the information of its neighboring nodes, so that the model can capture local and global structural information. After each layer of graph convolutional network, the protein scale vector is used as a prompt and the features output by each layer of the network are fused and input into the next layer of graph convolutional network.

[0051] Point cloud modality encoder construction: Figure 4As shown in the figure, the common point cloud network is used in this example, which is mainly used to process the three-dimensional spatial data of proteins. Point cloud data represents the three-dimensional coordinate information and residue chemical feature information of each residue in the protein. The point cloud deep network is a hierarchical point cloud processing framework. Its main feature is to capture the local and global geometric features of the point cloud by using the layer-by-layer point cloud abstraction layer sampling and upsampling modules. In the input point cloud data, each point not only contains three-dimensional coordinates and residue features, but also adds a protein size hint vector. In the sampling stage, representative key points are selected through multi-scale feature sampling. The local geometric features of the point cloud are learned using a sub-network in the local neighborhood. In the global feature aggregation stage, the point cloud network aggregates all local features into a global representation. After each abstract layer and upsampling layer, the calculated protein size vector is fused with the output feature of the layer and then input into the next layer of the network. The network can more effectively process protein point cloud data of different sizes and improve the ability to understand three-dimensional spatial information.

[0052] S3. Pre-trained network training

[0053] The preprocessed protein data is input into the network for pre-training. At this time, a prediction head is added after the encoder to predict the masked amino acid type. This is used as an unlabeled pre-training proxy task to allow the encoder to learn the feature extraction ability related to protein scale on a large-scale unlabeled dataset, and then the pre-trained encoder is used for downstream tasks.

[0054] S4. Fine-tuning and downstream task application

[0055] Based on the trained network weights, fine-tune the model and apply it to downstream tasks. In downstream tasks, the pre-trained network is applied to tasks such as protein function prediction and protein binding site detection through transfer learning. Fine-tune the model using a small amount of annotated data so that the model can better adapt to the needs of specific tasks and optimize its prediction performance. The fine-tuning process includes the following steps: fine-tune each encoder of the network to better handle the specific structural features of proteins; train for specific tasks to ensure that the network can achieve the best performance in downstream tasks.

[0056] Embodiment 2:

[0057] The present invention also provides a self-supervised learning pre-training system based on protein size prompts, comprising:

[0058] Acquisition module, used to acquire protein datasets;

[0059] A training module is used to construct and train a pre-trained network based on self-supervised learning according to a protein dataset;

[0060] An optimization module for fine-tuning on downstream tasks based on the pre-trained network after training.

[0061] As an implementation manner of the embodiment of the present invention, the acquisition module converts the protein dataset into a graph modality dataset and a point cloud modality dataset.

[0062] As an implementation method of an embodiment of the present invention, the pre-trained network includes: an encoder for a graph modality that integrates protein scale cues and an encoder for a point cloud modality that integrates protein scale cues, extracting multi-scale features of the protein structure in a hierarchical manner; wherein the graph modality encoder is used to specifically encode a graph modality dataset; and the point cloud modality encoder is used to specifically encode a graph modality dataset.

[0063] As an implementation of an embodiment of the present invention, the training module includes:

[0064] A first processing unit is used for classifying proteins into predefined size intervals according to the number of residues based on the protein data set, and encoding input proteins falling into a specific interval into corresponding size hint vectors according to the interval sequence number; wherein the size features of proteins in the same interval are consistent;

[0065] The second processing unit is used to add the protein prompt vector to the encoders of different modalities and train the pre-trained network to complete the masked amino acid type prediction.

[0066] As an implementation of an embodiment of the present invention, the optimization module is used to extract protein features in the downstream task using a pre-trained network after training, and the labeled downstream task is trained from scratch to fine-tune the pre-trained network after training.

[0067] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. A self-supervised learning pre-training method based on protein size cues, characterized in that: include: Step S1, obtaining a protein data set; Step S2: constructing and training a pre-trained network based on self-supervised learning according to the protein dataset; Step S3: fine-tune the downstream task based on the trained pre-trained network.

2. The protein size-based self-supervised learning pre-training method according to claim 1, characterized in that: In step S1, the protein dataset is converted into a graph modality dataset and a point cloud modality dataset.

3. The protein size-based self-supervised learning pre-training method according to claim 2, characterized in that: In step S2, the pre-trained network includes: an encoder for the graph modality that integrates protein scale cues and an encoder for the point cloud modality that integrates protein scale cues, extracting multi-scale features of the protein structure in a hierarchical manner; wherein the graph modality encoder is used to specifically encode the graph modality dataset; and the point cloud modality encoder is used to specifically encode the graph modality dataset.

4. The protein size-based self-supervised learning pre-training method according to claim 3, characterized in that: In step S2, the pre-trained network to be trained includes: Based on the protein dataset, proteins are classified into predefined size intervals according to their residue numbers, and input proteins falling into a specific interval are encoded into corresponding size hint vectors according to the interval sequence number; wherein, the size features of proteins in the same interval are consistent; The protein cue vectors are added to the encoders of different modalities and the pre-trained networks are trained to complete the masked amino acid type prediction.

5. The protein size-based self-supervised learning pre-training method according to claim 4, characterized in that: In step S3, the deep network for extracting protein features in the downstream task uses the trained pre-trained network, and is trained from scratch using the labeled downstream task to fine-tune the trained pre-trained network.

6. A self-supervised learning pre-training system based on protein size cues, characterized in that: include: Acquisition module, used to acquire protein datasets; A training module is used to construct and train a pre-trained network based on self-supervised learning according to a protein dataset; An optimization module for fine-tuning on downstream tasks based on the pre-trained network after training.

7. The protein size-cued self-supervised learning pre-training system according to claim 6, characterized in that: The acquisition module converts the protein dataset into a graph modality dataset and a point cloud modality dataset.

8. The protein size-cued self-supervised learning pre-training system according to claim 7, characterized in that: The pre-trained network includes: an encoder for graph modality that integrates protein scale cues and an encoder for point cloud modality that integrates protein scale cues, which extract multi-scale features of protein structure in a hierarchical manner; wherein the graph modality encoder is used to specifically encode the graph modality dataset; and the point cloud modality encoder is used to specifically encode the graph modality dataset.

9. The protein size-cued self-supervised learning pre-training system according to claim 8, characterized in that: The training modules include: A first processing unit is used for classifying proteins into predefined size intervals according to the number of residues based on the protein data set, and encoding input proteins falling into a specific interval into corresponding size hint vectors according to the interval sequence number; wherein the size features of proteins in the same interval are consistent; The second processing unit is used to add the protein prompt vector to the encoders of different modalities and train the pre-trained network to complete the masked amino acid type prediction.

10. The protein size-cued self-supervised learning pre-training system according to claim 9, characterized in that: The optimization module uses the trained pre-trained network to extract protein features in the deep network of the downstream task, and uses the labeled downstream task to train from scratch to fine-tune the trained pre-trained network.